DeepSeek v4 Flash has a nice bump in Capability
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
DeepSeek V4 Flash: Preview → 2026-07-31
| Benchmark | Preview | 0731 | Δ |
|---|---|---|---|
| Terminal Bench* | 56.9 | 82.7 | +25.8 |
| Toolathlon | 51.8 | 70.3 | +18.5 |
| NL2Repo | — | 54.2 | new |
| Cybergym | — | 76.7 | new |
| DeepSWE | — | 54.4 | new |
| Agent Last Exam | — | 25.2 | new |
| Automation Bench | — | 25.1 | new |
| DSBench-FullStack | — | 68.7 | new |
| DSBench-Hard | — | 59.6 | new |
* Terminal Bench changed from v2.0 → v2.1, so the improvement isn't a strict apples-to-apples comparison.
Compared to GPT-5.6 Terra
| Benchmark | GPT-5.6 Terra | DeepSeek V4 Flash | Advantage |
|---|---|---|---|
| Terminal Bench | 78.4 | 82.7 | Flash (+4.3) |
| Toolathlon | 53.1 | 70.3 | Flash (+17.2) |
| DeepSWE | 69.6 | 54.4 | Terra (+15.2) |
| Agents' Last Exam | 50.4 | 25.2 | Terra (+25.2) |
Trading blows, but no clear winner... very interesting!
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.