DeepSeek-V4-Flash-0731: Oneshot evals, surprisingly not token efficient??
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I ran the newly released DeepSeek-V4-Flash-0731 in my oneshot eval harness across 34 prompts and here are the results. https://oneshotlm.com/model/deepseek-deepseek-v4-flash-0731/ The providers on openrouter were unstable and I had to retry generation multiple times. Surprisingly it costed $1.29 to go through all 34 prompts failing to produce 5 outputs whereas kimi k3 only costed $0.44 without any failures. DeepSeek V4 Flash 0731: 2.7/5 score, $1.29 cost, 753k tokens Kimi K3: 3.2/5 score, $0.44 cost, 233k tokens Am I doing something wrong?? How is your experience with this model compared to Kimi K3? [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.