← Back to Lab
idle Dueling DQN ε 1.00 γ 0.98 LR 0.0005

Snake Simulation

Smooth rendering for pathfinding heuristics and endgame play.

Smooth realtime

Experience overview

You are watching a live snake agent blend classic pathfinding with learned policy choices. The board renders each step while the agent prioritizes safe routes, food collection, and endgame survivability.

  • Adjust playback speed, board size, and the BFS/Hamilton overlay to see how routing decisions change.
  • Switch algorithms, toggle curriculum training, and tune reward sliders to change what the agent values.
  • The end-game parameter (especially the tail space threshold) is often crucial—small tweaks can determine when the snake switches into safer late-game behavior.
20×20

Learning

Auto handles the curriculum, save logic, and hyperparameters for you.
Add MP3 files to the Audio folder to enable the in-game soundtrack.
3
Sets the starting length; with the curriculum enabled longer start shapes are used throughout the run.
Episodes0
Avg reward (100)0.0
Best length0
Best length (test)0
Fruit / ep0.0
Greedy Fruit (avg)—
Greedy Reward (avg)—
Gap (Fruit − Greedy)—

Prioritized replay, n-step returns, and dueling heads make DQN stable and sample efficient.

Auto-run

0 = unlimited auto-run cycles.
Adjustment tempo

Training progress

Average per 100 episodes
Run at least 100 episodes to see the trends.
Average reward Average fruit Greedy Reward Greedy Fruit
Episodes —

Reward telemetry

Rolling averages per component
Net reward (avg 100) 0.00 (trend +0.00)
Component Last Avg 100 Avg 500 Share Trend

Reward & tuning

Adjust incentives and hyperparameters.
Reward model

Tempo & direction

Loops & revisits

Crashes & stalling

High-impact rewards

BFS / Hamilton Rewards

Matches with the planner award the full bonus; ignoring the helper costs roughly 25% of the slider value so the agent leans on guidance only when confident.



Advanced settings

Shared

DQN family

Use the online network for action selection while the target net evaluates the value.