AssemblyAI incident
Outage on US Async Endpoint
Started
September 16, 2026 at 08:22 PM UTC
Duration
1h 14m
Resolved
September 16, 2026 at 09:37 PM UTC
Updates timeline
- Investigating
We are currently investigating an issue that is affecting all customers using our Async API on our US endpoint. Users will be receiving a server error. We will update with more information as we learn more. This began at 8pm UTC.
- Identified
The issue has been identified. This is affecting approximately 50% of all async transcriptions on the US endpoint.
- Identified
We are continuing to investigate these elevated errors. We will provide an update as soon as possible.
- Identified
We are seeing a reduction in the number of errors, but we are still working to fully resolve the issue.
- Monitoring
We have mitigated the issue and our Async API service is recovering.
- Resolved
The issue, which was the result of an AWS service issue, has been resolved and we have continue to see good performance since implementing a fix so we are now closing this issue.
- Postmortem
**Summary** On September 16, 2026, between 20:04 and 21:08 UTC, our US Async API returned elevated server errors. At peak, roughly 50% of async transcription requests on the US endpoint failed, and a portion of the requests that did succeed took longer than usual to complete. Our EU endpoint and real-time/streaming services were not affected. **Root cause** The failure originated in AWS SQS, the managed queuing service our transcription pipeline uses to distribute work. Specifically, the Fair Queues feature began rejecting valid messages with an `InvalidParameterValue` error on a parameter our services had been sending successfully for months. There was no deploy or infrastructure change on our side that triggered this; AWS confirmed it as a service-side issue and rolled back the change on their end later that evening. Because the affected queues sit in the path between our API and our transcription workers, requests that could not be queued failed outright, and a portion of the work that did get through was delayed behind reduced throughput. We mitigated by disabling Fair Queues across our services and reverting to standard SQS queues, which restored normal operation roughly an hour after the first error. **What we're doing about it** * **Faster configuration changes.** Much of our service configuration lives in environment variables, which require task restarts to propagate. We are moving this to a global feature flag system so mitigations like this one take seconds rather than minutes. * **Expedited emergency deploys.** We are adding a reviewed break-glass path so urgent fixes can bypass non-essential CI steps. * **Multi-region failover.** We are prioritizing work to fail the US pipeline over to a second region when a regional provider dependency degrades, rather than relying on feature-level mitigations alone. We're sorry for the disruption. No customer audio or transcript data was lost, and any requests that failed during this window can be safely resubmitted.
Get an email the next time AssemblyAI goes down
Outage alerts for AssemblyAI, straight to your inbox. No account needed, unsubscribe in every email.
Watching more than one service? A free account covers 5 services + a daily digest — and Pro is currently free.
Live AssemblyAI status
Current indicator + 24h latency
All incidents
Cross-service timeline
Subscribe via RSS
Atom feed for any reader