r/LocalLLaMA · · 1 min read

Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode

Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_nonogram_logic_puzzle/. That thread shaped v1.2:

  • Reasoning effort is explicit per run
  • Every prompt and output is public.
  • All current top ranking private and open weight models have been added
  • Someone spotted Grok miscounting a 400-character answer. Turns out that trips up most models, so Hard mode answers row by row rather than a single text string.

Results: - GPT-6 Astra: 30/30, the first perfect run on 15x15 puzzles - Best open weights: DeepSeek V4 Pro 83% (tied 4th), DeepSeek V4.1 Flash 77% for $0.84 total - Hard mode (10 random 20×20s, one solution each): Opus 5.5 8/10. Every open-weight model: 0/10

Still OpenRouter-only, so no way to run locally yet. PRs welcome.

nonobench.com (raw data, API, and code on GitHub)

submitted by /u/mauricekleine
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA