GPUHedge: Hedging serverless GPU providers improves cold start p95 latency from 117s to 30s [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
Disclosure: I built it, it is open source, Apache-2.0 licensed, and currently alpha. Repository: https://github.com/mireklzicar/gpuhedge
I started working on it after benchmarking a 17 GB AI model across several serverless GPU providers.
On the primary provider, requests usually either completed in roughly 6–8 seconds or took around 90–122 seconds after a fresh GPU cold start. Simply switching to another provider did not remove the problem because every provider had its own tail.
GPUHedge treats this as a speculative-execution problem.
It starts a request on a primary provider, watches the job’s lifecycle state, and conditionally launches or switches to a backup. The first result that passes a validator wins, and the losing job is cancelled through the provider’s native API.
You can try the policy engines without creating provider accounts or spending money:
pip install gpuhedge
In the initial benchmark, a fixed RunPod → Cerebrium hedge launched after 10 seconds. On the 36-request evaluation portion, it changed:
- observed p95 latency from 116.6 s to 29.4 s;
- requests over 60 seconds from 11/36 to 0/36;
- modeled active-compute cost from $0.0114 to $0.0083 per request.
What is your experience with cold start latency? Which provider to add next? Can something like this help what you are building?
[link] [comments]
More from r/MachineLearning
-
For the people who got reviews back from neurips, cvpr, eccv, etc and also tested their paper through an agentic reviewer like the stanford one, how different were the reviews? [D]
Aug 14
-
Building text to ASCII diffusion model , need advice and guidance [P]
Aug 14
-
A collision-entropy floor for watermark/retrieval AI-text detection. Looking for a sanity check before I take this further [D]
Aug 14
-
Are supervised and unsupervised learning still relevant today? [D]
Aug 14
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.