Replicate incident

Delayed scaling due to node failure

Minor Resolved

Started

August 10, 2026 at 09:05 PM UTC

Duration

40 min

Resolved

August 10, 2026 at 09:45 PM UTC

Updates timeline

  1. Monitoring

    Scaling decisions were delayed by nearly 1 hour after the controller responsible for emitting queue metrics failed to schedule on a soft-failed node. We have since cordoned and drained the node, and the controller is emitting queue metrics again.

  2. Resolved

    All affected models have been scaling correctly for more than 30 minutes at this point, and we see no residual prediction queues. Thank you for your patience!

Get an email the next time Replicate goes down

Outage alerts for Replicate, straight to your inbox. No account needed, unsubscribe in every email.

Watching more than one service? A free account covers 5 services + a daily digest — and Pro is currently free.

Live Replicate status

Current indicator + 24h latency

All incidents

Cross-service timeline

Subscribe via RSS

Atom feed for any reader