All Posts

Cloud & DevOps

AWS and Cloudflare's 2026 Outages Trace to the Same Weak Point

AWS and Cloudflare's 2026 Outages Trace to the Same Weak Point

Bhavika J

Techshorts Editorial Team

Cloudflare's status page logged a cluster of incidents in the first two weeks of August 2026, touching object storage, key-value reads and email filtering. It is the fourth distinct infrastructure failure at a major provider this year that had nothing to do with an attacker or a traffic spike. AWS had three of its own between May and July. Taken together, they show where cloud reliability is actually breaking in 2026: not at the edge, but inside the internal systems providers run to keep the internet online.

A cooling failure took Coinbase and FanDuel offline for most of a day

At 11:50 PM UTC on May 7, 2026, cooling units failed inside a single data center hall serving availability zone use1-az4 in AWS's US-EAST-1 region in Northern Virginia. As temperatures rose past safe operating limits, servers in the affected racks shut down automatically, cutting power to EC2 instances and EBS volumes (CNBC, 2026).

Coinbase was offline for roughly seven hours because its matching engine ran in that single zone by design, a case where being multi-AZ elsewhere in the account did not help (CoinDesk, 2026). FanDuel crashed mid-event during an NBA playoff game and later confirmed the outage was tied to AWS. Recovery took AWS until 1:50 PM PDT the next day to restore cooling capacity, with full service restoration stretching close to 28 hours from the first failure.

A fiber cut, not a hack, knocked out X, Reddit and Zoom in June

On June 22, Cloudflare engineers detected elevated error rates and latency at 13:35 UTC and traced the cause to a fiber cut in Eastern North America by 14:37 UTC. Because Cloudflare sits in front of roughly a quarter of the web as a reverse proxy, the physical cut cascaded into visible outages at X, Reddit, Microsoft Teams, Zoom, Discord and Fortnite, with some services recovering within twenty minutes and others taking longer as Cloudflare rerouted traffic (The National, 2026).

Cloudflare was explicit that this was a physical link failure, not a software bug or an attack. The company's own reverse-proxy share was measured at 24.2% of all websites as of July 28, 2026, more than fifteen times the next closest competitor (W3Techs, via Statista, 2026), which is the underlying reason a single fiber cut produces headlines about "the internet going down."

On July 24, AWS lost connectivity between its US-WEST-2 region in Oregon and the Seattle Metro network path. Outage trackers picked up simultaneous failures at Apple Pay, DoorDash, Reddit, Hulu and PlayStation Network. The disruption ran about 80 minutes from first reports to AWS's resolution notice, and traffic that stayed inside the region kept working while anything crossing the region boundary saw timeouts (Tech Times, 2026).

Incident response firm Pagerly, reviewing the event afterward, noted that multi-region deployment within AWS, running production across at least two regions with automated failover, would have insulated most applications from this specific failure mode (Pagerly, 2026). That is a known mitigation. Most of the affected consumer apps were not running it.

Cloudflare opened August with a cluster of its own

Cloudflare's incident log shows an R2 object storage failure in its Eastern North America region beginning at 14:52 UTC on August 7, with writes to a subset of buckets impacted for roughly two hours before Cloudflare identified the cause and mitigated it (Cloudflare Community, 2026). Five days later, on August 12, an upstream Spamhaus blocklist listing disrupted Cloudflare's Email Security delivery for customers who route filtering through Spamhaus. Cloudflare opened the incident at 15:04 UTC and confirmed all affected IPs were clean by 16:48 UTC, with full resolution logged at 17:23 UTC (Cloudflare Community, 2026).

Neither incident approached the scale of the June fiber cut. What connects them to the rest of the year's pattern is the cause: internal storage and filtering systems, not external pressure.

Why this matters for buyers, not just operators

None of these four incidents was caused by a DDoS attack or an unexpected surge in demand. Each traces to a fault inside infrastructure the provider itself operates: a cooling system, a physical fiber link, a network route, a storage backend. That distinction matters for anyone evaluating a cloud or CDN contract. Uptime SLA percentages describe how often a provider stays up on average. They say nothing about blast radius, how much of a customer's stack goes down when one internal system fails.

The practical questions worth asking a vendor now are narrower than "what is your uptime." They are: which of my workloads share a single availability zone, a single network path, or a single storage backend, and what does the failover plan look like when that specific piece fails. AWS's own guidance already recommends active-active multi-region deployment for exactly the July 24 failure mode. Most of the companies affected in these four incidents were not running it.

What to watch next

Cloudflare has not yet published a formal post-incident review covering the full August cluster; that review, when it lands on the Cloudflare blog, should clarify whether the R2 and Spamhaus incidents share a root cause or are coincidental. AWS has not disclosed whether the May, July, and any later 2026 incidents connect to a common infrastructure issue across its network paths. Both are worth checking before assuming this pattern is behind us.

Sources

  1. CNBC, "AWS data center outage hits trading on FanDuel, Coinbase, recovery to take hours" - https://www.cnbc.com/2026/05/08/aws-outage-data-center-fanduel-coinbase.html
  2. CoinDesk, "Coinbase disruption tied to AWS outage draws criticism amid staff layoffs and Q1 losses" - https://www.coindesk.com/business/2026/05/08/coinbase-disruption-tied-to-aws-outage-draws-criticism-amid-staff-layoffs-and-q1-losses
  3. The National, "Widespread outage hits X, Reddit, Teams, Cloudflare and other services" - https://www.thenationalnews.com/future/technology/2026/06/22/internet-outage-x-cloudflare-reddit/
  4. Tech Times, "AWS Knocks Out Apple Pay, Reddit, Hulu for 80 Minutes in Third Outage Since May" - https://www.techtimes.com/articles/321567/20260725/aws-knocks-out-apple-pay-reddit-hulu-80-minutes-third-outage-since-may.htm
  5. Pagerly, "AWS Outage Incident Response: What July 24 Taught Us" - https://www.pagerly.io/blog/aws-outage-incident-response-2026-07-25
  6. Cloudflare Community, "R2 Availability Issues" - https://community.cloudflare.com/t/r2-availability-issues/946713
  7. Cloudflare Community, "Email Security delivery impacted by Spamhaus listing" - https://community.cloudflare.com/t/email-security-delivery-impacted-by-spamhaus-listing/948172
  8. Statista, citing W3Techs, "Cloudflare, a hidden pillar of the internet" - https://www.statista.com/chart/35487/market-share-of-reverse-proxy-services-cloudflare/