r/LocalLLaMA · · 1 min read

Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

submitted by /u/xenovatech
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA