r/LocalLLaMA · · 1 min read

Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

In tests:
~80 tok/s decoding
2,500–3,500 tok/s long-input prefilling
Smooth use by 3–4 concurrent users

Private, on-device inference for coding, agents, and offline batch jobs

submitted by /u/niacolhealth
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA