ClashRoyaleAi: an open-source, deterministic Clash Royale simulator for RL, with recurrent PPO, lookahead search and expert iteration [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| Our PPO agent learned to park its Cannon behind its own King. Losing a building in a fight cost reward, and letting it decay cost nothing, so it found the loophole. It's one of many things we learned building a Clash Royale simulator from scratch so an agent could learn the game. The engine is deterministic C++ with Python bindings, plays a full match in about 10 ms on one laptop core, and can fork any game state in microseconds, so lookahead is cheap. Best result so far: a simple 1-ply lookahead took the policy from 0.625 to 0.944 win rate against a heuristic bot (160 paired matches). Distilling it back into the network kept only +0.045. The agent isn't strong yet, and RL isn't my home field, so feedback from people who know it better would mean a lot. Repo: https://github.com/itzik123/ClashRoyaleAi Built with my friend Ambash (most of the card roster). I used AI coding tools as a pair programmer. [link] [comments] |
More from r/MachineLearning
-
How can I turn an industry ML project into a publication? [R]
Sep 28
-
Are there any good research papers around Text clustering using LLMs [R]
Sep 28
-
Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books [P]
Sep 28
-
Two-stage shelf audit: YOLO finds the products, embeddings can't tell sibling SKUS apart. What should Stage 2 be? [P]
Sep 27
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.