All Posts

Cloud & DevOps

AWS's Repeat Outages Highlight the Platform Maturity Gap

AWS's Repeat Outages Highlight the Platform Maturity Gap

Bhavika J

Techshorts Editorial Team

AWS has had four separate regional reliability incidents since May 2026, the most recent a second connectivity failure between its US-WEST-2 region and the Seattle Metro internet exchange in August (Tech Insider, 2026). The four incidents did not share a root cause: a data-center cooling failure, a third-party network provider fault, and two separate network-hardware problems at the same regional boundary. What they share is where they broke. Internal systems kept running. The failures all landed at the point where AWS''s infrastructure meets the outside internet. That pattern, and what a Kubernetes release and a new platform engineering survey say about the gap between stable and unstable cloud operations this year, is the story.

AWS's Seattle Metro Problem, Twice in Three Weeks

On July 24, 2026, AWS lost connectivity between its US-WEST-2 region in Oregon and the Seattle Metro area, the internet exchange point that carries much of the region's traffic to the wider internet. AWS attributed the fault to networking hardware on that path. Traffic that started and ended inside the region kept working; anything crossing the region boundary, including the AWS Management Console for some customers, saw timeouts (IncidentHub, 2026). Most customers were affected for about 20 minutes, but customers connected through the Westin Building Exchange in Seattle via AWS Direct Connect saw outages of up to 1 hour and 17 minutes before routes fully reconverged (IncidentHub, 2026). Consumer-facing casualties included Apple Pay, DoorDash, Reddit, Hulu, Snapchat, Fortnite and the PlayStation Network, offline for roughly 80 minutes in what Tech Times called AWS's third reliability incident since May (Tech Times, 2026).

A second, shorter US-WEST-2 to Seattle Metro connectivity failure followed in August, briefly touching US-WEST-1 as well (Tech Insider, 2026). AWS has not published a full post-incident report naming which hardware failed or why, on either the July or the August event. That is worth flagging plainly: the account here rests on independent incident-tracking outlets that converge on the same timeline and cause, not on an AWS-issued root-cause analysis.

The two Seattle Metro incidents followed a data-center thermal event in Northern Virginia in May, which triggered a power loss in one availability zone of US-EAST-1 and knocked Coinbase offline for roughly seven hours (OutageCost, 2026), and a June network disruption tied to third-party provider Zayo that also affected Cloudflare (Tech Insider, 2026). Four incidents, four different failure points, in seventeen weeks. AWS still holds 28% of the global cloud infrastructure market as of the first quarter of 2026, more than Microsoft Azure (21%) and Google Cloud (14%) combined against everyone else (Synergy Research Group, as reported by CRN, 2026). Repeat failures at that scale are why concentration risk is back in the conversation among buyers who single-source a cloud provider without a cross-region or cross-provider fallback.

Kubernetes Closes a Six-Year Gap on Pod Resizing

Kubernetes v1.35, released December 17, 2025 under the working name "Timbernetes," graduated In-Place Pod Resize to stable, more than six years after the feature entered alpha in Kubernetes v1.27 (Kubernetes Blog, 2025). The change lets a running pod's CPU and memory requests and limits be adjusted without restarting the container in most cases. Before this, resizing a pod meant killing it and starting a replacement, which is disruptive for stateful workloads and one of the standing complaints about running databases or long-lived services on Kubernetes.

The release paired that with the Vertical Pod Autoscaler's InPlaceOrRecreate update mode graduating to beta, meaning VPA can now resize a live pod instead of always evicting and recreating it (The New Stack, 2025). Kubernetes 1.35 shipped 60 enhancements in total, with 17 reaching stable status (Kubernetes Blog, 2025). Adoption has moved through the managed-platform layer quickly: Amazon EKS and EKS Distro added 1.35 support in January 2026 (AWS, 2026), and K3s shipped a matching release the same month (K3s, 2026). The latest patch, 1.35.6, landed June 9, 2026. For platform teams, this closes one of the longer-running operational gaps in the project and removes a real reason some stateful workloads stayed off Kubernetes.

The Survey That Explains the Difference

Perforce published a report on July 8, 2026 based on a survey of 820 technology professionals, fielded by research firm Panterra in November and December 2025 (Perforce, 2026; a vendor-commissioned study conducted by an independent research firm). Its central finding: platform engineering maturity, not raw AI adoption, is what separates organizations getting value from AI infrastructure tooling from those getting instability. Among organizations the report classified as having mature platform engineering practices, 73% said that maturity was a critical or significant factor in their AI success, against 44% at less mature organizations. Governance told a similar story: 79% of platform-mature organizations reported strong governance automation, compared with 14% of immature ones. Separately, 66% of respondents said they were already using AI in infrastructure workflows. Only 31% described that AI as fully autonomous, which suggests most organizations are further along in adopting AI tools than in trusting them to act without a human in the loop.

Read next to the AWS pattern, the finding is a useful corrective. The incidents in Oregon and Virginia were provider-side failures, not customer misconfiguration, and no platform engineering practice would have prevented a chiller failure or a bad line card in a Direct Connect facility. But they are also exactly the kind of event a mature platform team is built to absorb. That means cross-region failover, traffic shaping away from a broken path, and a runbook that does not depend on the AWS Management Console staying reachable during an AWS incident. The Perforce data suggests that gap between resilient and fragile responses is now measurable, and it is showing up most clearly wherever AI-driven infrastructure automation is involved.

What to Watch Next

AWS has not said whether it will publish detailed post-event summaries for the July or August US-WEST-2 incidents, something it has done for past major outages. Whether it does, and what it says about the recurring Seattle Metro dependency, is worth watching. AWS's re:Invent conference runs November 30 through December 4, 2026, in Las Vegas (Founderpath, 2026), and reliability announcements at that event, if the pattern continues into the autumn, would be a clear signal about how seriously AWS is treating the run of incidents rather than treating each one as isolated.


Sources

  • Tech Insider. "AWS Outage Hits US-West-2: 4th Incident in 4 Months." 2026. https://tech-insider.org/aws-outage-us-west-2-2026/
  • IncidentHub. "The July 24, 2026 AWS us-west-2 Outage: Network Routing and a Long Recovery Tail." 2026. https://blog.incidenthub.cloud/aws-us-west-2-outage-jul-24-2026
  • Tech Times. "AWS Knocks Out Apple Pay, Reddit, Hulu for 80 Minutes in Third Outage Since May." July 25, 2026. https://www.techtimes.com/articles/321567/20260725/aws-knocks-out-apple-pay-reddit-hulu-80-minutes-third-outage-since-may.htm
  • OutageCost. "AWS Outage Cost - History, Real Losses & Your Exposure." 2026. https://outagecost.com/aws-outages
  • CRN (citing Synergy Research Group). "Q1 2026 Global Cloud Market Share Figures." 2026. https://x.com/CRN/status/2052053282546163943
  • Kubernetes Blog. "Kubernetes v1.35: Timbernetes (The World Tree Release)." December 17, 2025. https://kubernetes.io/blog/2025/12/17/kubernetes-v1-35-release/
  • The New Stack. "Kubernetes 1.35 'Timbernetes' Introduces Vertical Scaling." 2025. https://thenewstack.io/kubernetes-1-35-timbernetes-introduces-vertical-scaling/
  • AWS. "Amazon EKS and Amazon EKS Distro now supports Kubernetes version 1.35." January 2026. https://aws.amazon.com/about-aws/whats-new/2026/01/amazon-eks-distro-kubernetes-version-1-35/
  • K3s. "Kubernetes v1.35 is out!" January 15, 2026. https://docs.k3s.io/blog/2026/01/15/K3s-1.35-release
  • Perforce. "Perforce's 2026 Platform Engineering Report Finds Platform Engineering Maturity Separates AI Advantage from Instability." July 8, 2026. https://www.perforce.com/press-releases/state-of-platform-engineering-2026
  • Founderpath. "AWS re:Invent 2026: Nov 30 in Las Vegas." 2026. https://founderpath.com/events/saas/aws-reinvent/2026-las-vegas