Incidents added from:
On September 2, from 16:48 to 19:28 PT, sandbox-management requests and some requests to sandboxes experienced errors. Sandbox-management performance recovered at 19:20 PT. The remaining errors affecting sandbox traffic were resolved at 19:28 PT.
We sincerely apologize for the disruption to your workloads.
Sandbox-management requests experienced elevated error rates and intermittent unavailability. Some requests to sandboxes failed when routing or automatic resume could not complete. Other sandbox traffic continued to succeed. Delayed cleanup also caused some sandboxes to run longer than intended.
The disruption began when load exceeded the capacity of a shared internal state service. As the number of running sandboxes grew, each sandbox-listing request required more work. Listing requests also became more frequent. Together, this increased demand on the shared service, delaying other sandbox operations and slowing cleanup.
Delayed cleanup left more sandboxes running, which made subsequent listing requests more expensive and prolonged the disruption.
Two weaknesses allowed the problem to escalate:
Our listing rate limits allowed more traffic than the shared service could handle alongside other sandbox operations.
Cleanup retried without completing, adding load to an already overloaded service.
We also took too long to identify the cause. Alerts detected the disruption, but our monitoring did not make the source of the overload clear. This delayed escalation and mitigation.
We increased capacity in the shared service, which provided temporary relief. Restarting the sandbox API later provided partial recovery, but errors and slow requests continued.
We then tightened listing rate limits. This reduced load and allowed cleanup to catch up. Sandbox-management performance recovered at 19:20 PT, and the remaining errors affecting sandbox traffic were resolved by 19:28 PT.
16:48 - Errors begin. Alerts fire, and we declare an incident.
16:58 - We add capacity to the shared service, providing temporary relief.
17:50 - Performance degrades as sandbox count and request load continue to grow.
18:30 - We identify frequent listing requests as a suspected source of the overload.
19:00 - We restart the sandbox API. Performance improves, but errors and slow requests continue.
19:18 - Stricter listing rate limits take effect. Load drops, and cleanup catches up.
19:20 - Sandbox-management performance recovers.
19:28 - The remaining errors affecting sandbox traffic are resolved.
Tightened sandbox-listing rate limits to reduce listing load on the shared service.
Updated listing and cleanup to process work in smaller batches.
Fixed excessive cleanup retries to reduce the load from repeated attempts.
Spread retry attempts over time to avoid simultaneous bursts.
Separating heavy listing work from other sandbox operations.
Adding protections against repeated connection attempts to a sandbox.
Improving monitoring and alerts to help us identify the source of overload and respond sooner.