r/LocalLLaMA · · 1 min read

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Hi.

I saw some feedback that halogen was degrading at context depth. So I fixed that.

Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0:

  • decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter)
  • decode at 258,794: 42.9 to 45.0
  • prefill at 1,004,581: 790 to 937 tok/s, 21.2 to 17.9 minutes cold
  • prefill at 258,794: 1,086 to 1,114 tok/s

Conditions: Ryzen AI Max+ 395, 128 GB. The 262k and 1M rows are one cold request each at the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576), greedy, 64 tokens, the rates the response's `timings` report. The 32k row is the standard ten-prompt served mean and did not change. A follow-up turn over the prompt cache at 1M reaches its first token in about 0.55 s; the numbers above are the cold path.

To run it at 1M: add -e HALOGEN_ROPE_YARN=4 -e HALOGEN_CTX=1048576to the README's podman line; it needs the 128 GB box. Release notes and the full table:

https://github.com/peonist-ai/halogen-flash-server

If you have a 1M sweep of your own, I would like to see it rerun on 0.12.0.

Thanks for all your support, especially https://huggingface.co/nightvich

submitted by /u/peonist-ai
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA