Ling Tiny, King of Speed
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Ling Tiny has now replaced Gemma4-12B in my rig as an auxiliary model doing hindsight operations. This is on a 4060Ti, which is a reasonable GPU available out there, and the speed is phenomenal. Don’t enable MTP, set up the vLLM fork for BailingMoE3. Hope this is useful to [link] [comments] |
More from r/LocalLLaMA
-
you can now use MTP in GLM-Air
Aug 23
-
Qwen 3.8 27b helped me with something unique that Opus 4 couldn't - Firmware + Software preservation and emulation on an early 2000's ARM based POS system
Aug 23
-
We quantized Qwen 3.8 27B and compared the quants on an RTX 6000
Aug 23
-
Nvidia Customers Notified About AI-Related Price Hikes Above 15%
Aug 23
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.