llama.cpp releases · · 1 min read

b10121: ui: reduce per-token render cost when streaming (#26053)

Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.

  • performance harness - the empirical root

Assisted-by: Claude Opus 4.8

  • 210.36ms -> 2.67ms per streamed token

Assisted-by: Claude Opus 4.8

  • 11.58ms -> 0.62ms per streamed token

Assisted-by: Claude Opus 4.8

  • 22.02ms -> 3.33ms per streamed token

Assisted-by: Claude Opus 4.8

  • 3.07ms -> 1.36ms per streamed token at 40 messages

Assisted-by: Claude Opus 4.8


Co-authored-by: Zach Winter [email protected]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from llama.cpp releases