Summary
On 2026-07-15, API requests experienced elevated error rates and increased tail latency (p99). Median latency (p50) was not affected and there was little effect on 95th percentile latency (p95), so most requests performed normally. We began investigating at 14:34 UTC and applied a fix at approximately 19:56 UTC, after which performance returned to normal.
What happened?
Through the affected period, a portion of API requests returned errors, responded with reason code SITO (Sardine internal timeout) or timed out rather than completing, with intermittent spikes during which a larger share of requests were affected.
Why did it happen?
A scaling configuration prevented part of our platform from adding capacity as traffic increased. As traffic rose through the day, available capacity was exhausted, causing timeouts and errors for some requests.
What are we doing about this?
We corrected the configuration at 19:56 UTC, which immediately restored normal performance. We are also making the affected part of the platform more resilient to sudden traffic increases and adding proactive monitoring so we can detect and respond to this class of issue faster in future. We apologize for the disruption.