r/LocalLLaMA · · 1 min read

Is DeepSeek v4 (Flash) really extremely cheap to run? If yes, how?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hi. I don't have a GPU. So my biggest "local LLM" experience has been running ~26B models with single-digits tps values.

However, the "serving economy" of DSv4 models look like a riddle to me. The Flash model has 284B parameters, but providers (e.g. OpenRouter) charge so little for it it's ridiculous. It's for example cheaper than 27B Qwen, A tenth of its (total) size! How is it viable?

Are the providers just doing dumping here? Or is DSv4 architecture somehow different in making it extremely cheaper to serve? Those of you who have had the equipment to host DSv4, is there something making it different?

Thanks

submitted by /u/ihatebeinganonymous
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA