b10121: ui: reduce per-token render cost when streaming (#26053)
Mirrored from llama.cpp releases for archival readability. Support the source by reading on the original site.
- performance harness - the empirical root
Assisted-by: Claude Opus 4.8
- 210.36ms -> 2.67ms per streamed token
Assisted-by: Claude Opus 4.8
- 11.58ms -> 0.62ms per streamed token
Assisted-by: Claude Opus 4.8
- 22.02ms -> 3.33ms per streamed token
Assisted-by: Claude Opus 4.8
- 3.07ms -> 1.36ms per streamed token at 40 messages
Assisted-by: Claude Opus 4.8
Co-authored-by: Zach Winter [email protected]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.