r/LocalLLaMA · June 17, 2026 · 1 min read

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?

arXiv : https://arxiv.org/abs/2606.17861
Full Paper : https://arxiv.org/pdf/2606.17861
HuggingFace : https://huggingface.co/papers/2606.17861
GitHub : https://github.com/tongxuluo/gamecraft-bench
Project : https://tongxuluo.github.io/gamecraft-bench-website/

I see big/large models(Opus-4.7, GPT-5.5, Kimi-K2.6, MiMo-V2.5-Pro, GLM-5.1, MiniMax-M2.7, DeepSeek-V4-Pro) on benchmarks. Curious to know how medium size models(Ex: Qwen3.6-27B, Gemma-4-31B) would perform on this.

Hopefully we get great medium size(30-70B) models with performance of 200B+ models(on everything .... at least on coding & writing) by end of this year.

submitted by /u/pmttyji
[link] [comments]

Discussion (0)

No comments yet. Sign in and be the first to say something.

Discussion (0)

More from r/LocalLLaMA