r/LocalLLaMA · · 1 min read

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

I see big/large models(Opus-4.7, GPT-5.5, Kimi-K2.6, MiMo-V2.5-Pro, GLM-5.1, MiniMax-M2.7, DeepSeek-V4-Pro) on benchmarks. Curious to know how medium size models(Ex: Qwen3.6-27B, Gemma-4-31B) would perform on this.

Hopefully we get great medium size(30-70B) models with performance of 200B+ models(on everything .... at least on coding & writing) by end of this year.

submitted by /u/pmttyji
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA