Got my Ascent GX10 two days ago, ran REAP-pruned NVFP4 DeepSeek-V4-Flash on a single Spark, and it stays consistent at long context
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Got my Ascent GX10 two days ago and spent the last couple of days pushing a REAP-pruned NVFP4 DeepSeek-V4-Flash setup on a single Spark by patching the Credit where it’s due: the REAPs were done by 0xSero. I’m just the person who wired it up, validated it, and pushed it through the machine. The main thing I wanted to check was long-context consistency, and the interesting part is how steady the throughput stays as context scales up. I also vibecoded a Grafana dashboard in Hermes so I can watch the Spark(served at 262k context with VLLM) without living in raw logs. Here are the numbers:
What stood out to me is that this thing stays surprisingly consistent at long context on a single Spark. The prefill and tg numbers don’t collapse the way you might expect as you stretch from Next up I’ll post the 180B REAP benchmarks too, and if the hardware cooperates I want to try longer contexts, maybe up to [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.