r/LocalLLaMA · · 1 min read

Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen

I wonder someone will figure out a way to do this with 27B?

Throw Qwen3 on this page for demo
https://kishida.github.io/webdemos/llkvapprox/

submitted by /u/T_rex2700
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA