r/LocalLLaMA · · 1 min read

[2608.16157] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

[2608.16157] FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Source of Claims: https://x.com/Andy_ShuoYang/status/2090856976880472439

Your gaming PC can now serve frontier models at interactive speed using official checkpoints without extreme quantization!

Qwen3.6 35B → 8GB RTX 4060 laptop @ 39 tok/s

DeepSeek-V4-Flash 284B → RTX 5090 desktop @ 22-25 tok/s

GLM-5.2 753B → RTX PRO 6000 workstation @ 15 tok/s

Run your claude code or codex now with frontier model for $0

FreeToken is fast. Comparing to Ollama, we have 3–4× faster decode, and 6–30× faster prefill

How? We introduce bandwidth-adaptive CPU–GPU execution + semantic-aware caching across agent turns.

submitted by /u/SteppenAxolotl
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA