r/LocalLLaMA · · 1 min read

GLM 5.2 FP8 with FP8 KV - Terminal-Bench 2.1 = 79.8 (with one time-out that I didnt re-run)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I wanted to test the official results vs fp8 + fp8 kv. basic sglang setup on H200.

If anyone wants one of the official tests do ping me.

I didnt rerun the one so it might go up a bit :)

 TERMINAL-BENCH 2.1 — FINAL RESULTS (mymodel via mini-swe-agent) TOTAL: 89 tasks PASSED: 71 (79.8%) FAILED: 17 ERRORED: 1 Input tokens 218,656,815 Cache tokens 216,036,672 Output tokens 4,659,650 Cache hit rate 98.8% New input tokens 2,620,143 FAILED (17) — completed but wrong answer configure-git-webserver db-wal-recovery dna-assembly dna-insert extract-moves-from-video filter-js-from-html gcode-to-text install-windows-3.11 model-extraction-relu-logits mteb-leaderboard protein-assembly query-optimize raman-fitting regex-chess torch-pipeline-parallelism video-processing winning-avg-corewars ERRORED (1) — timed out torch-tensor-parallelism [VerifierTimeoutError — not rerun] 
submitted by /u/Daemonix00
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA