Deepseek V4 Flash just hit Colibri, does anyone have numbers?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I'm mosty interested in 128-192GB VRAM with 128-256GB RAM to spare, so SSD streaming is basically not even necessary. Seems only FP4 is supported, so older hardware will likely be slow - no Unsloth GGUF supported either.
I'd be curious what people are getting with V100s, R9700s, etc, just to have some comparison.
What's prefill like >200k context? Tg/s high enough to support agentic workloads?
It's probably wishful thinking, but when I saw the release, my immediate thought was Sonnet 5 level model being "affordable" to consumers.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.