Amrit DePaulo
← Writing

Oops, the Internet Broke Again, Thanks, AWS (I mean DNS)!

Amrit DePauloOctober 2025

Quick Take

On Monday, October 20, 2025, Amazon Web Services (AWS) suffered a major disruption centered in its US-EAST-1 (Northern Virginia) region that cascaded across the internet (hey, that's where I keep all my stuff). The root cause, confirmed in AWS's post-event summary, was a latent race condition in DynamoDB's automated DNS management: one automation applied a stale DNS plan over a newer one, cleanup automation then deleted it, and the DNS record for DynamoDB's regional endpoint ended up empty. Dozens of dependent AWS services cascaded into elevated error rates and downtime for scores of consumer and enterprise apps worldwide over roughly 15 hours.

The ripples were pretty ripply. Platforms including Snapchat, Duolingo, Roblox, Venmo, and Slack, UK banks (Lloyds Banking Group, Halifax, Bank of Scotland), and airlines (United and Delta) were impacted, and some government services went offline or degraded.

In short, this is a giant reminder that when a "cloud backbone" player sneezes, much of the internet catches a pretty nasty cold, although it generally doesn't last all that long.

Scope of What Happened

The incident began early Monday in US-EAST-1, per AWS's status page and independent monitoring. It wasn't just one tiny micro-service; AWS confirmed "multiple AWS services... increased error rates and latencies." The impact spanned consumer apps (games, social, fintech), enterprise SaaS, financial institutions, and some government services. In the UK alone, hundreds of thousands of user reports were logged; globally the numbers were in the millions.

There was no evidence of a malicious cyberattack. AWS attributed it to an internal fault in DNS automation. The outage underscores the single-point-of-failure risk inherent in centralized cloud infrastructures, and experts continue to point to this as a systemic weakness.

Who Was Impacted (and Who Could Have Been)

Impact observed:

  • Consumer apps like Snapchat, Roblox, Duolingo, and Fortnite (this pissed off my kids)
  • Fintech and crypto services like Venmo and Coinbase reported issues (this hurt some wallets)
  • E-commerce and Amazon's own services: the Amazon retail site and Ring doorbell cameras experienced disruptions (I might actually have to get in my car and buy something with my hands, physically <shudder>)
  • Financial institutions and government in the UK: Lloyds, Halifax, and Bank of Scotland had access issues and the HMRC website behavior altered (ok, that might leave a mark)
  • Airlines including Delta and United faced disruptions, with passengers unable to check in, view reservations, or change seat assignments (this may just be "normal," I experience these issues while traveling pretty much every single time)

Could have been impacted (and likely were) but less visible:

  • Enterprise internal systems running on AWS (customer-data platforms and internal apps), which means revenue and productivity losses behind the scenes
  • Critical infrastructure providers that depend on AWS: telecom backends, logistics, healthcare digital records
  • Supply-chain services and vendors that use AWS for order fulfillment and SaaS. Your business may feel the pain even if your in-house systems were fine.

How Organizations (Including Government) Should Respond

Incident review and root cause analysis. Public cloud providers like AWS publish post-event summaries for major incidents. Organizations should review their own dependency maps: what services sit on AWS (or any single cloud) and what the failure modes are.

Communication and transparency. For impacted services, internal or external, be ready to notify stakeholders and customers quickly. Government entities should ensure that public services can communicate status clearly through fallback websites and social media.

Assess regulatory and contractual implications. Outages may trigger SLA claims and regulatory reporting requirements, especially in finance and government. Review contracts with cloud providers and determine what happens when the backbone goes down.

Re-evaluate governance of critical digital infrastructure. Governments need to ask whether too many vital services depend on too few providers. As one commentator put it, internet users are at the mercy of too few providers. Consider regulatory frameworks for digital resilience: cloud diversification, backup routes, and vendor lock-in risk.

Run tabletop scenarios and live drills. If a major cloud region fails, what happens to your business or public service? Spoiler alert: it's not just IT. It is also procurement, communications, finance, and customer service.

After-action accountability. Who in the org is responsible for cloud resilience? Is it just "IT"? Or does it span business continuity, risk, procurement, and legal? Document lessons learned and feed this into next year's risk register.

What Enterprises Should Be Doing to Mitigate This Type of Single Point of Failure

This isn't the first time, and it won't be the last, that we experience these types of large technology failures, but the point stands: big singular dependencies have bitten many.

For example, while not AWS, the CrowdStrike update fiasco in 2024 knocked millions of machines offline. Or think of the 2021 Facebook and Instagram outage, where one routing change cascaded globally (this was actually a good thing and the world was better for it, just kidding).

Here's a simple (to talk about, but really hard and expensive to actually implement) enterprise resilience checklist:

Multi-Cloud / Hybrid / "Not-Just-One" Strategy

  • Architect systems so that a failure in a single cloud region or provider doesn't bring your business to its knees.
  • Use multiple regions, and multiple providers if you can afford this level of resiliency (AWS and GCP and/or Azure, or not Azure), or fall back to private cloud or on-premises, but this is all very expensive.
  • Design for cross-region failover, data replication, and automated recovery.
  • Beware of cloud monoculture. Just because it's easy and cheap doesn't mean it's low risk. Experts have warned of the danger of global reliance on a small number of tech companies for years.

Resilient Architecture and Chaos Engineering

  • Adopt chaos testing: intentionally inject failures (region down, network partition, cloud API unresponsive) and verify your systems survive or gracefully degrade.
  • Design for graceful degradation. If service X is down, can you still offer limited functionality? Show cached data? Fall back to read-only?
  • Ensure your blast radius is minimal, and remember that a large monolithic service hosted in one region equals BASR (big ass scary risk).

Vendor and Third-Party Risk Management

  • Inventory your cloud-service dependencies. Are key services of yours, or your vendors, hosted in the affected region or provider?
  • Review vendor contracts. Do they have regional redundancy? What happens if their cloud provider suffers an outage?
  • Include resilience in procurement criteria. Ask potential vendors: "If AWS US-EAST-1 is down for 6 hours, what happens?"

Backup, Data Replication, and Monitoring

  • Regular backups and off-site or multi-region replication of data are table stakes.
  • Don't just back up; practice restore at least annually, quarterly if you can. Determine how long it takes, what the dependencies are, and what can be improved.
  • Monitor and alert on more than "is the service up." Track how long until degradation triggers. In this outage, visibility into error rates and latency was key. Define thresholds that alert you that bad things are coming. There are a ton of tools for this.

Business Continuity and Incident Planning

  • IT downtime isn't just an IT issue: customer service, legal, communications, and finance all get affected. Build your scenario so all functions can answer the question, "What happens if our cloud backbone fails for 4 to 8 hours?"
  • Define roles and responsibilities now, so it is clear who communicates to customers, who activates the backup plan, and who logs the incident.
  • Insure appropriately, and understand that many businesses focus on data breaches and far fewer on a cloud-provider-failure event.

Cybersecurity and Cloud Resilience, the Twin Pillars

  • While this outage wasn't a cyberattack (darn!), the vulnerability is structurally similar: a single point of failure, shared service risk, and cascading dependencies.
  • From a security lens, evaluate your trust surface, and not just on-prem assets but the cloud infrastructure, third-party services, regions and providers, and of course AI (had to toss that in there, because, well, AI).
  • From a resilience lens, assume failure will happen and build for recovery. This is no longer optional.

There is so much more that should be done. Hopefully this event provides an opportunity for companies to take inventory, evaluate their tools and processes, and optimize for the next one, which is most likely just around the digital corner.