We'll benchmark an Open weights LLM on any GPU you choose — drop your model + hardware and we'll run it. [D]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
We run HexGrid Cloud, a platform for deploying open-source models on GPUs, and we're heads-down optimizing our serving/deployment layer.
To pressure-test it we're benchmarking real models under real concurrency — and instead of guessing, we'd rather run what you actually want to see.
---
Models available for benchmarking:
- Nemotron-3 Super 120B-A12B (only NVFP4)
- Nemotron-3 Nano 30B A3B
- Qwen-3.6 27B
- Llama 3.3 70B Instruct
- Gemma-4 31B
- Devstral-Small-2-24B-Instruct-2512
- ?? (you suggest a model to us)
We're focused on chat/instruct models for now (that's what most of our users deploy), so pick one from the list above — or suggest another open-weight chat model that fits on a single H200 (141GB).
---
Hardware & quant choices:
- GPU (up to H200 for this round): RTX PRO 6000 · L40S · H100 · H200
- Quant: FP8 / AWQ / BF16
- Context length: (8K, 32K, 64K, 128K)
- What you want measured: max throughput? single-stream speed? long-context prefill?
---
We'll run the top picks and post full results — tokens/sec, TTFT, TPOT, throughput under concurrency, and cost-per-million-tokens — config and flags included so it's reproducible.
Let us know in comments.
[link] [comments]
More from r/MachineLearning
-
For the people who got reviews back from neurips, cvpr, eccv, etc and also tested their paper through an agentic reviewer like the stanford one, how different were the reviews? [D]
Aug 14
-
Building text to ASCII diffusion model , need advice and guidance [P]
Aug 14
-
A collision-entropy floor for watermark/retrieval AI-text detection. Looking for a sanity check before I take this further [D]
Aug 14
-
Are supervised and unsupervised learning still relevant today? [D]
Aug 14
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.