What happened
On November 25, 2020, the day before US Thanksgiving, AWS engineers did something routine: they added capacity to the Amazon Kinesis front-end fleet in us-east-1. The problem was hiding in an implementation detail. Each front-end server maintained operating-system threads to communicate with every other server in the fleet. As the fleet grew, the total thread count climbed, and this particular addition pushed it past the maximum threads the operating-system configuration allowed.
Once servers hit that ceiling, they could no longer build their shard-map, the structure that tells each server how to route stream data, and they began to fail. Kinesis Data Streams in us-east-1 started erroring on both produce and consume operations. Then the cascade began, because a surprising amount of AWS runs on Kinesis internally. CloudWatch metrics and alarms degraded, Cognito authentication faltered, Lambda functions wired to Kinesis event sources stopped firing, and parts of the AWS console went sideways.
Recovery was painfully slow, and instructively so. AWS could not simply restart the fleet, because bringing thousands of servers back at once would have them all rebuild shard-maps simultaneously and overwhelm the system again. So they recovered in small increments and raised the thread limit, finishing after roughly ten hours.
SLA credit analysis (the tier and dollar logic)
Amazon Kinesis Data Streams carries a 99.9% monthly uptime commitment, measured per Region. That is a tight budget: 99.9% allows only about 43 minutes of downtime across a full 30-day month. This outage ran roughly ten hours in us-east-1, which is more than thirteen times that entire monthly allowance in a single afternoon.
The Kinesis credit schedule is tiered against your monthly Kinesis charges in the affected Region:
- Below 99.9% but at or above 99.0%: a 10% service credit.
- Below 99.0% but at or above 95.0%: a 25% service credit.
- Below 95.0%: a 100% service credit.
A ten-hour outage lands every affected account at least in the 10% band, and heavier or longer-affected accounts that dropped below 99.0% for the month reach the 25% band. If you spent 15,000 dollars on Kinesis in us-east-1 that month, the 10% band is a 1,500 dollar credit and the 25% band is 3,750 dollars, on Kinesis alone.
The cascade is where a lot of value was quietly forfeited. CloudWatch, Cognito, and Lambda each carry their own SLAs. If your usage of those services breached their targets during the same window, because the Kinesis fault dragged them down, those are separate claims, priced against each service's own bill. Most customers filed for Kinesis, if they filed at all, and never chased the dependent services.
There is a second, easy-to-miss detail in this specific outage: your own monitoring may have been blind during it. CloudWatch was one of the degraded services, so the very dashboards and alarms you would normally use to prove downtime were themselves unreliable for part of the window. That does not weaken your claim, but it does mean you should reconstruct the timeline from multiple sources: the AWS post-event summary, your application logs, and any third-party monitoring you run outside AWS. When your provider's telemetry is part of the failure, independent evidence is what turns a plausible claim into a documented one, and it is the difference AWS support looks for when it decides whether to apply the credit.
How to claim
- Confirm the window. Impact ran from midday November 25 to just after midnight UTC on November 26, 2020, in us-east-1.
- Inventory affected services. Start with Kinesis, then add every dependent service you were billed for that also failed: CloudWatch, Cognito, Lambda event sources, and any downstream product.
- Compute monthly uptime per service and map each to its own credit tier.
- Price the total in the free SLA credit calculator, then file a case per service in the AWS Support Center.
The AWS SLA credits overview walks the case flow, and the cascade-claim guide shows how to chase dependent-service credits, not just the headline service.
Lessons
- The biggest money in a cascade is usually not the service that failed first. Kinesis was the trigger, but CloudWatch and Cognito breaches were separate, claimable events that most teams ignored.
- Slow recovery is a feature, not a bug, of large distributed systems, which means outages are often measured in hours, not minutes. Your SLA math should assume that.
- us-east-1 is the recurring center of gravity; the same pattern reappears five years later in the AWS DynamoDB October 2025 outage. Start from the calculator to see what a breach is worth, then read how it works.
For cross-provider outage histories see awsdown.com, azuredown.com, and gcpdown.com. Our sponsor Next Signal automatically detects breaches and drafts both SLA credit and overcharge claims across your accounts, and cloud-credits.com covers credit recovery in depth.