r/LocalLLaMA · · 1 min read

DeepSeek-V4-Flash-Vision-Exp (285B MoE) on 10-12x RTX 3090 — spec decoding, vision

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Running the full deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

on consumer Ampere — 10-12x RTX 3090, SM86-compatible vLLM build.

285B MoE, FP4 experts + FP8 attention, 157 GB weights.

Highlights:

- **60+ tok/s** decode, DSpark spec (k=3)

on 10 GPUs (TP2xPP5), at a 240 W cap

- **120+ tok/s** on 12 GPUs (TP4xPP3)

- **Vision + spec + tool calls all working**

- **1M context** (no offload) / **4M** (RAM offload)

- ~3,500 tok/s long-context prefill

Fully documented + reproducible:

- Pre-built image:

`docker pull ghcr.io/ciprianveg/3090-vllm:dsv4-flash-vision-sm86`

- Repo: https://github.com/ciprianveg/3090-vllm

(build guide, start scripts, runtime patches)

The patches cover the DSpark propose-gate (spec + vision fix),

scheduler mm x spec row-crossing, grammar-bitmask validation,

vision ViT OOM fix, FlashInfer workspace-lane keying, and more.

submitted by /u/ciprianveg
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA