Looking for feedback on my GPU-accelerated Snake AI project [P]
Mirrored from r/MachineLearning for archival readability. Support the source by reading on the original site.
| I've been building an AI that learns to play the classic Snake game through reinforcement learning. The goal is to reach high scores while keeping training time as low as possible. The current version averages 86 points (87 is the maximum) after less than 10 hours of training on a single free Google Colab T4 GPU. To keep training fast, it runs 4,096 Snake games directly on the GPU, combines GPU-native environment simulation with PPO + GAE, and uses a spatially-preserving CoordConv architecture that maintains the full game grid throughout training. I'm sure there's still room to improve. If you've worked on reinforcement learning or efficient training systems, what would you try next? Better exploration, reward design, network architecture, or something else? Repository: (https://github.com/siddhartha399/PPO-CoordConv-Snake) I'd really appreciate any feedback or criticism. [link] [comments] |
More from r/MachineLearning
-
I built an "honest" CS conference ranking: sorted by how good the trip is, not the CORE ranking [P]
Aug 12
-
Context-Induced Activation Drift: Long benign context passively decouples RLHF alignment without adversarial prompts (Mechanistic Interpretability + Ablation) [D]
Aug 12
-
Decoupled Descent: Enforcing Exact Train-Test Error Tracking Via AMP Onsager Corrections [R]
Aug 11
-
Research direction: Intelligent Model Weight transfer between LLMs [R]
Aug 11
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.