r/LocalLLaMA · · 1 min read

bonsai's document reveal how much cherry picked their headlines are

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai_2/q3.5 60.8 / 72.4 = 84.0%

Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release [2], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

also 3.8 35b qwhen? plsss

submitted by /u/KURD_1_STAN
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA