r/LocalLLaMA · · 4 min read

Ternary Bonsai 2 27B (1.75bpw) vs. Gemma 26B-A4B MoE

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Introduction

My audiobook pipeline that has to decide who speaks each line of dialogue in a novel, so the TTS can cast voices per character. It's been running on Gemma 4 26B-A4B (QAT Q4). Bonsai 2 27B looked like it should win: a 27B-class model in 5.9 GB means a stronger base model fully resident on my 8 GB card, instead of a 4B-active MoE streaming experts over DDR5.

I gave it four chapters. Here's what happened.

Setup

Hardware: NVIDIA T1000 8GB — Turing TU117, ~6.6 GB free. Ryzen 5 7600, 93 GB DDR5.

Incumbent: unsloth/gemma-4-26B-A4B-it-qat-GGUF UD-Q4_K_XL, 14 GB. 26B total / ~4B active MoE, QAT weights. --cpu-moe parks every expert in system RAM with attention + KV on the GPU, plus a 241 MB MTP draft head for speculative decoding.

Challenger: prism-ml/Ternary-Bonsai-2-27B-gguf PTQ1_0, 5.95 GB. Ternary {−1,0,+1} weights with FP16 group scales, ~1.75 bpw, on a dense Qwen3.8-27B base — 64 layers, hybrid 3× Gated DeltaNet → 1× full attention. Needs the PrismML-Eng/llama.cpp fork (tag prism-b10687-5d80cff); stock llama.cpp rejects PTQ1_0 as an unknown type. No draft head exists.

Task: Batched JSON extraction. Each request carries N quotations with ~700 chars of preceding and ~400 chars of following prose each, the character roster, an intonation vocabulary, and 6 turns of prior attributions. Non-thinking, temp 0.3.

Three things it took to get running:

  1. Turing decodes The fork has an explicit PTQ1_0 DP4A path for cc >= TURING in ggml_cuda_should_use_mmvq. But MMQ prefill is gated on turing_mma_available(), so TU117 emulates mma on FP16 units: ~18 tok/s prefill, identical at ubatch 128/256/512. Their cuBLAS escape hatch (GGML_CUDA_PTQ1_0_MMQ_MAX_BATCH) OOM'd in my headroom.

  2. One CPU layer costs 2.6×. llama.cpp's auto-fit keeps a 1 GiB margin, which spilled a single layer → 2.5 tok/s. Forcing -ngl all → 6.5 tok/s. Two compounding reasons: the fork's PTQ1_0 CPU kernels are generic C with no AVX path, and one CPU layer disables the fused DeltaNet ops for every layer:

    resolve_fused_ops: layer 0 is assigned to device CPU but fused Gated Delta Net is assigned to device CUDA0

    resolve_fused_ops: fused Gated Delta Net (chunked) not supported, set to disabled

  3. Fitting 5.9 GB into 6.6 GB. Qwen3.8's large vocab makes the logits buffer scale hard with ubatch — at -ub 2048 it alone is ~2 GB. Final config: -ngl all -b 2048 -ub 256 -ctk q8_0 -ctv q8_0, giving 5,395 MiB weights + 265 MiB host buffer, ~5,980 MiB resident, 7,650 / 8,192 MiB total with Firefox open. f16 KV missed by 71 MiB. I also dropped from 32 to 16 quotations per request: Bonsai's longer reasoning overflowed 8192 ctx (n_tokens = 8191, truncated = 1) and the halve-and-retry cost ~13 min per overflow.

Sampling per Qwen3.8's non-thinking guidance: top_k 20, top_p 0.8, presence_penalty 1.5 (to avoid looping).

Methodology

Ground truth is four chapters (480 quotations) of previously-attributed, hand-corrected output. I score exact speaker match on (start_offset, end_offset) keys, plus: attributions to a character flagged silent in the roster, correct narrator identification, off-roster speakers routed to the unknowns file instead of hallucinated onto the roster, and intonation agreement.

The reference was produced by the incumbent and then corrected by hand, so it measures Bonsai against a human-verified gold standard, not a head-to-head against Gemma under identical conditions. Single sample per chapter, one book, one prompt.

Findings

chapter quotations speakers correct A 11 11/11 B 110 108/108 C 154 142/150 (94.7%) D 244 187/211 (88.6%) total 480 448/480 (93.3%) 

Narrator identity drift. In several scenes it decided a different roster character was narrating and handed the narrator's "I said" lines to them — roughly half of D's 24 errors, each with a confident wrong rationale ("The first-person narrator is X, as established by…").

Adjacent-turn swaps in rapid untagged dialogue (8 errors in chapter C).

Refusal to commit on short exclamations, where the reasoning identifies "the female speaker" but won't name her → 5 lines dropped from the chapter file entirely.

Intonation regression (100/480. Systematic: it returns the reporting verb (said, asked, accused) where the reference carries a delivery tone (deadpan, earnest, curt)

~27–30 s per quotation, steady across chapters. 110 quotes = 55 min; 244 quotes = 121 min. ~6 tok/s decode, ~18 tok/s prefill.

Conclusion

Not adopting it. The 9×-smaller claim holds true. Runs 27B-class weights on an 8 GB Turing card, which is a real achievement. But the only advantage was better comprehension than a 4B-active MoE, and on the two hard chapters it produced 32 speaker errors, 5 dropped lines, a systematic first-person narrator failure and an intonation regression at 5–10× the wall-clock over Gemma.

If your task is short-horizon and generation-bound, that footprint is remarkable. Mine is long-context, prefill-heavy reading comprehension, and 1.75 bpw seems to hobble the capability it needs: tracking who is speaking in a long scene.

Follow-up

Any other models I should try?

submitted by /u/autonoma_2042
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA