r/MachineLearning · · 1 min read

Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]

Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.

Real task cost across GPT, Claude, Gemini and Kimi, 10.6x spread on models with only 2x price difference [R]

Ran 10 realistic product tasks (classification, RAG QA, multi turn conversation, an agentic plan then execute task, etc) against the live APIs of OpenAI, Anthropic, Gemini and Kimi, using each provider's cost optimized tier.

Total cost spread was 10.6x despite published rates differing by only 2x, mostly driven by reasoning and thinking tokens billed at the output rate but never shown in the response. Clearest example, a one word classification answer where one model burned 197 tokens of invisible reasoning to produce it.

Ties into recent work on this exact blind spot, CostBench (ACL 2026) finds leading models routinely fail to choose cost optimal plans, and TerminalWorld reports failed agent attempts burn disproportionately more tokens than successful ones (r = -0.62).

Full methodology, raw results, and the 10 task prompts are public on GitHub.

useweckr.com/benchmark

https://preview.redd.it/9zahx5t74xeh1.png?width=1642&format=png&auto=webp&s=4922d8c25eca5b903403d277dc4d52dec393921e

https://preview.redd.it/8180vuh24xeh1.png?width=2236&format=png&auto=webp&s=69b06b605a79c33455ccf66266e6023538597231

submitted by /u/pixelo2323
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/MachineLearning