Thinking Machines' best public Tinker result used Qwen3-235B, not Inkling. is the base model actually that important?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I Went through the Inkling model card and the Bridgewater/Tinker case study instead of the press coverage. Coverage mostly quoted the 97.1% AIME 2026 number; the rest of the table tells a more mixed story.
AIME 2026: Inkling 97.1%, GLM 5.2 99.2%, Fable 5 and GPT-5.6 Sol both 99.9%. Everyone on the list is above 94%, so this one doesn't say much on its own.
On HLE text-only, Inkling scores 29.7%. Ahead of Nemotron 3 Ultra, behind GLM 5.2, DeepSeek V4 Pro, and both Kimi models.
Same pattern on SWEBench Pro and Terminal Bench 2.1: beats Nemotron 3 Ultra and Kimi K2.5, loses to Kimi K2.6, GLM 5.2, and DeepSeek V4 Pro.
The one that doesn't get mentioned much: Inkling actually leads IFBench (instruction following) at 79.8%, second only to Nemotron 3 Ultra.
The Bridgewater case study everyone points to as proof fine-tuning beats frontier models used Qwen3-235B as the base, not Inkling. 84.7% accuracy across six financial document-filtering tasks, roughly 13.8x lower inference cost per task than the frontier models tested. Published two weeks before Inkling existed.
Two different claims keep getting collapsed into one: that Tinker can turn an open model into a strong specialist, and that Inkling specifically is a good base for that. The public evidence backs the first. Nothing public backs the second yet.
For anyone who's actually fine-tuned large MoE models: does Inkling's IFBench score and 41B active-parameter setup make it worth trying as a base, or would you still reach for Qwen or Kimi since their fine-tuning behavior is better documented?
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.