workstudioinsightslet's talk ↗
We turn ideas into digital products.
  • Home
  • Work
  • Studio
  • Insights
  • Let's talk

Find us

  • Instagram
  • LinkedIn
When AWS Sneezes, the Internet Catches a Cold
When AWS Sneezes, the Internet Catches a Cold

When AWS Sneezes, the Internet Catches a Cold

What the October 2025 outage revealed about our quiet dependency on the cloud

October 20, 2025

It started like any other Monday: coffee, Slack pings, a quick deploy. Then the graphs went weird.

You know that look — latency spikes that make no sense, alerts stacking on top of each other like falling dominoes. Someone on the team joked, “US-East-1’s haunted again,” and everyone laughed. Because that’s what you do when you’re scared something might actually be wrong.

Five minutes later, it was.

DynamoDB was failing. Lambda was stalling. EC2 was rolling dice. Reddit was down. So was Snapchat. So was Ring. And Alexa had gone eerily silent, which somehow made it all worse.

When AWS sneezed, the internet caught a cold.

The Day the Cloud Came Back to Earth

Amazon’s postmortem would later reveal the cause: a race condition in their internal DNS automation system.

Two processes tried to update the same configuration. One lagged. The other thought it was cleaning up old data and, in doing so, wiped out live IP addresses. Within seconds, a core part of AWS forgot where its own services lived.

That’s it. That’s all it took. A small, invisible slip — and a huge chunk of the modern web blinked offline.

The outage rippled outward, hitting banking apps, IoT devices, and global retail systems. CyberCube, an insurance analytics firm, estimated the insured losses between $38 million and $581 million. A half-billion-dollar bug, courtesy of the most sophisticated cloud platform on earth.

Our Beautiful, Fragile Internet

I’ve always loved the metaphor of “the cloud.” It feels limitless. Weightless. Safe.

But outages like this remind us it’s not a cloud — it’s a data center in Virginia. With racks, cables, human engineers, and scripts that sometimes collide. The cloud isn’t floating; it’s very much on the ground.

We tell ourselves the web is distributed, but it’s not. It’s concentrated in a few massive providers — AWS, Azure, Google Cloud — each one a single point of failure wearing a high-availability badge.

And that’s fine, until it isn’t.

The Developer’s Reality Check

Every developer I know did the same thing that morning: Checked status.aws.amazon.com. Refreshed Twitter (sorry, X). Then opened architecture diagrams in mild panic.

Multi-AZ suddenly didn’t feel like much of a safety net. Because if the problem’s DNS-level, availability zones don’t matter — they’re all connected through the same vulnerable systems.

You start thinking about multi-region setups. Or cross-cloud deployments. You even Google “GCP free tier.” But by the time you get to pricing, AWS is back up, and the sense of urgency quietly fades.

Until the next time.

What Broke, and What Broke Us

AWS fixed the race condition, rebuilt automation safeguards, and promised better internal isolation. They always do, and they always mean it.

But for everyone else, the bigger story isn’t how AWS failed — it’s how much we did.

Whole startups went offline. Entire workflows froze. For many companies, the outage revealed how few contingencies actually existed beyond “hope Amazon figures it out soon.”

We’ve spent years building on someone else’s resilience. It’s worked brilliantly — until it doesn’t.

Decentralization, or Just Diversification?

Whenever a big outage happens, people start whispering about decentralization. Web3. Distributed compute. Peer-to-peer redundancy.

It’s a compelling vision — no single company with the power to knock half the internet offline. But decentralization comes with trade-offs: governance, latency, cost, complexity. It’s not a silver bullet.

Maybe the answer isn’t pure decentralization — maybe it’s diversification. A mindset that says no provider, no region, no system should hold all the keys.

True resilience doesn’t mean avoiding failure. It means expecting it — and designing your system so failure doesn’t take you with it.

The Morning After the Outage

When services started coming back online, I refreshed our dashboards. Everything looked fine again. The graphs leveled out. Customers returned.

But I couldn’t shake the thought that we’d just survived on luck and other people’s engineering.

We talk about “trusting the cloud” as if trust is a feature you can deploy. But it’s not. Trust without redundancy is just faith. And faith doesn’t scale.

The next time I spin up a project, I’ll probably still click “us-east-1.” Out of habit. Out of convenience. But I’ll do it with a little more humility — and maybe a quiet plan B.

Because the cloud isn’t a metaphor. It’s a system built by humans, run by humans, and occasionally, broken by humans.

And when it sneezes, we all feel it.

Ollie Darby

About Ollie Darby

Ollie Darby, is the co-founder of Tradible, a visionary leader in the realm of digital collectibles. With a robust background as a skilled full-stack software engineer, Ollie's expertise spans across a spectrum of technologies.

When AWS Sneezes, the Internet Catches a Cold

What the October 2025 outage revealed about our quiet dependency on the cloud

October 20, 2025

It started like any other Monday: coffee, Slack pings, a quick deploy. Then the graphs went weird.

You know that look — latency spikes that make no sense, alerts stacking on top of each other like falling dominoes. Someone on the team joked, “US-East-1’s haunted again,” and everyone laughed. Because that’s what you do when you’re scared something might actually be wrong.

Five minutes later, it was.

DynamoDB was failing. Lambda was stalling. EC2 was rolling dice. Reddit was down. So was Snapchat. So was Ring. And Alexa had gone eerily silent, which somehow made it all worse.

When AWS sneezed, the internet caught a cold.

The Day the Cloud Came Back to Earth

Amazon’s postmortem would later reveal the cause: a race condition in their internal DNS automation system.

Two processes tried to update the same configuration. One lagged. The other thought it was cleaning up old data and, in doing so, wiped out live IP addresses. Within seconds, a core part of AWS forgot where its own services lived.

That’s it. That’s all it took. A small, invisible slip — and a huge chunk of the modern web blinked offline.

The outage rippled outward, hitting banking apps, IoT devices, and global retail systems. CyberCube, an insurance analytics firm, estimated the insured losses between $38 million and $581 million. A half-billion-dollar bug, courtesy of the most sophisticated cloud platform on earth.

Our Beautiful, Fragile Internet

I’ve always loved the metaphor of “the cloud.” It feels limitless. Weightless. Safe.

But outages like this remind us it’s not a cloud — it’s a data center in Virginia. With racks, cables, human engineers, and scripts that sometimes collide. The cloud isn’t floating; it’s very much on the ground.

We tell ourselves the web is distributed, but it’s not. It’s concentrated in a few massive providers — AWS, Azure, Google Cloud — each one a single point of failure wearing a high-availability badge.

And that’s fine, until it isn’t.

The Developer’s Reality Check

Every developer I know did the same thing that morning: Checked status.aws.amazon.com. Refreshed Twitter (sorry, X). Then opened architecture diagrams in mild panic.

Multi-AZ suddenly didn’t feel like much of a safety net. Because if the problem’s DNS-level, availability zones don’t matter — they’re all connected through the same vulnerable systems.

You start thinking about multi-region setups. Or cross-cloud deployments. You even Google “GCP free tier.” But by the time you get to pricing, AWS is back up, and the sense of urgency quietly fades.

Until the next time.

What Broke, and What Broke Us

AWS fixed the race condition, rebuilt automation safeguards, and promised better internal isolation. They always do, and they always mean it.

But for everyone else, the bigger story isn’t how AWS failed — it’s how much we did.

Whole startups went offline. Entire workflows froze. For many companies, the outage revealed how few contingencies actually existed beyond “hope Amazon figures it out soon.”

We’ve spent years building on someone else’s resilience. It’s worked brilliantly — until it doesn’t.

Decentralization, or Just Diversification?

Whenever a big outage happens, people start whispering about decentralization. Web3. Distributed compute. Peer-to-peer redundancy.

It’s a compelling vision — no single company with the power to knock half the internet offline. But decentralization comes with trade-offs: governance, latency, cost, complexity. It’s not a silver bullet.

Maybe the answer isn’t pure decentralization — maybe it’s diversification. A mindset that says no provider, no region, no system should hold all the keys.

True resilience doesn’t mean avoiding failure. It means expecting it — and designing your system so failure doesn’t take you with it.

The Morning After the Outage

When services started coming back online, I refreshed our dashboards. Everything looked fine again. The graphs leveled out. Customers returned.

But I couldn’t shake the thought that we’d just survived on luck and other people’s engineering.

We talk about “trusting the cloud” as if trust is a feature you can deploy. But it’s not. Trust without redundancy is just faith. And faith doesn’t scale.

The next time I spin up a project, I’ll probably still click “us-east-1.” Out of habit. Out of convenience. But I’ll do it with a little more humility — and maybe a quiet plan B.

Because the cloud isn’t a metaphor. It’s a system built by humans, run by humans, and occasionally, broken by humans.

And when it sneezes, we all feel it.

Ollie Darby

About Ollie Darby

Ollie Darby, is the co-founder of Tradible, a visionary leader in the realm of digital collectibles. With a robust background as a skilled full-stack software engineer, Ollie's expertise spans across a spectrum of technologies.

What did you think of this post?

AmazingGoodMehBad
Thanks for rating this post—join the conversation by commenting below.

Subscribe to Circle Square Blog

Get the latest insights on AI, design, and technology

StarStarStarStarStar

"This might be the best value you
can get from an AI subscription."

- Jay S.

MailEvery Content
AI&I PodcastAI&I Podcast
CoraCora
SparkleSparkle
SpiralSpiral

Join 100,000+ leaders, builders, and innovators

Community members

Email address

Email

Already have an account? Sign in

What is included in a subscription?

Daily insights from AI pioneers + early access to powerful AI tools

SparkleSpiralAI&I PodcastEveryCora
PencilFront-row access to the future of AI
CheckIn-depth reviews of new models on release day
CheckPlaybooks and guides for putting AI to work
CheckPrompts and use cases for builders
CheckIn-depth reviews of new models on release day
CheckPlaybooks and guides for putting AI to work
CheckPrompts and use cases for builders
SparksBundle of AI software
Sparkle

Sparkle:Organize your Mac with AI

Cora

Cora:The most human way to do email

Spiral

Spiral:Repurpose your content endlessly

Sparkle
Sparkle:Organize your Mac with AI
Cora
Cora:The most human way to do email
Spiral
Spiral:Repurpose your content endlessly

Related Essays

When Microsoft Misfired

When Microsoft Misfired

What the October 2025 Microsoft Azure outage taught us about cloud fragility

Oct 27, 2025

Ollie DarbyOllie Darby

Say no more It's a match ↓

  • General questions
    hello@circlesquare.studio
  • New business enquiries
    Contact
  • Instagram
  • LinkedIn
  • 3rd Floor, 86-90 Paul Street London England EC2A 4NE United Kingdom

    GPS
    N 51° 31' 32.32" W 0° 5' 01.15"
  • Psst Psst. Wanna get Spammm?

Venture — coming soon. We co-build with founders.

Tell us what you're building. Give us the real version — what stage you're at, what's broken, what you need. The more honest you are, the faster we can tell you whether we can help.

01.
Brief us on what you need:
02.
Budget (£)
03.
Your Company
04.
Your Name
05.
Your Email
06.
Project Details
07.
Project Delivery Date
08.
How did you find us?

FAQ

Eight things worth knowing before we work together.

  • Mostly early-stage startups and founders building something for the first time — or rebuilding something that didn't work the first time. We've been through it ourselves, so we understand the constraints.

  • We're flexible. Sometimes it's a full brand and website. Sometimes it's one component that's holding everything else back. Tell us what you need and we'll be straight with you.

  • It depends on the scope, but most projects start with a brief discovery session — paid, focused, and designed to get us both clear on what actually needs doing.

  • By project, not by the hour. We scope it properly upfront so there are no surprises on either side.

  • Yes. If you've got something that's working, we build on it. If it isn't working, we'll tell you.

  • Fully. We've worked with founders across the UK and Europe. If you're local and want to meet, great — but it's not a requirement.

  • Front-end: Next.js, React, TypeScript. We pick the right tool for the project rather than defaulting to one stack.

  • Occasionally we co-build with early-stage founders — equity arrangements, sprint-based partnerships, or acting as a technical co-founder for the first stretch. It's selective. If it sounds relevant, mention it when you get in touch.

Testimonials

The kind words are on their way. We're just getting to work first.

Coming soon

Good things to say — soon.

Coming soon

Watch this space

01 — 05