Local agentic coding Benchmark : Qwen3.8-Flash-Next NVFP4 vs 27B (and the others...)
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Using https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 and https://old.reddit.com/r/BlackwellPerformance/comments/1w04xb7/qwen38_flashnext_on_1x_rtx_pro_6000_171_ts_c1_428/ As usual, all the details in https://wonderrico.github.io/local_llm_benchmark/benchmark-main.html?filter=3.8 and even more in https://wonderrico.github.io/local_llm_benchmark/benchmark-detail.html?filter=3.8 (the bad score one is a "random" uncensored version from HF https://huggingface.co/dealignai/Qwen3.8-Flash-Next-UNCENSORED-NVFP4 ) I shall test other ones Bottom line : almost highest score of all local model I tested, the most efficient in both nb requests / point and fewer generated tokens / pt, all in medium reasoning. (xhigh is not useful, again, in this benchmark) and if it was not enough very fast All that for an undertrained model... [link] [comments] |
More from r/LocalLLaMA
-
Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
Sep 9
-
Surveillance plagiarism by OpenAI
Sep 9
-
Don't let FOMO win if you're interested in local llm from a hobby/learning aspect
Sep 9
-
Server rebuild to custom loop. 2x RTX Titans 24gb, 1x 22gb 2080ti | T: 70GB VRAM.
Sep 9
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.