Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| Ran 10 realistic product tasks (classification, RAG QA, multi turn conversation, an agentic plan then execute task, etc) against the live APIs of OpenAI, Anthropic, Gemini and Kimi, using each provider's cost optimized tier. Total cost spread was 10.6x despite published rates differing by only 2x, mostly driven by reasoning and thinking tokens billed at the output rate but never shown in the response. Clearest example, a one word classification answer where one model burned 197 tokens of invisible reasoning to produce it. Ties into recent work on this exact blind spot, CostBench (ACL 2026) finds leading models routinely fail to choose cost optimal plans, and TerminalWorld reports failed agent attempts burn disproportionately more tokens than successful ones (r = -0.62). Full methodology, raw results, and the 10 task prompts are public on GitHub. [link] [comments] |
More from r/MachineLearning
-
TMLR Relevance and Prestige [D]
Aug 13
-
Reproducible canvas-aligned low-level patterns in somerandomllm-generated images and their possible relation to iterative editing artifacts [D]
Aug 13
-
worldproof: diagnosing where world-model predictions break and a measurement of when pixel metrics stop being able to rank models at all [P]
Aug 13
-
UrgenT Help Detecting Performance Regressions Using Machine Learning and Hardware Counters [P]
Aug 13
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.