Snake Simulation
Smooth rendering for pathfinding heuristics and endgame play.
Experience overview
You are watching a live snake agent blend classic pathfinding with learned policy choices. The board renders each step while the agent prioritizes safe routes, food collection, and endgame survivability.
- Adjust playback speed, board size, and the BFS/Hamilton overlay to see how routing decisions change.
- Switch algorithms, toggle curriculum training, and tune reward sliders to change what the agent values.
- The end-game parameter (especially the tail space threshold) is often crucial—small tweaks can determine when the snake switches into safer late-game behavior.
Learning
Prioritized replay, n-step returns, and dueling heads make DQN stable and sample efficient.
Auto-run
Training progress
Average per 100 episodesReward telemetry
Rolling averages per component| Component | Last | Avg 100 | Avg 500 | Share | Trend |
|---|
Reward & tuning
Adjust incentives and hyperparameters.Reward model
Tempo & direction
Loops & revisits
Crashes & stalling
High-impact rewards
BFS / Hamilton Rewards
Matches with the planner award the full bonus; ignoring the helper costs roughly 25% of the slider value so the agent leans on guidance only when confident.