STUDY / 04
Explore Reinforcement Learning
Learn better actions through rewards.
Tabular / grid
Can rewards reveal a route?
Press Step and watch the TD target update one action value. Then train and look for a route to the goal.
Q ← Q + α [r + γ max Q(next) − Q]Q-learning
0 episodes0.00last reward
0.00TD target
0.00TD error
↑ 0.00→ 0.00↓ 0.00← 0.00
Action values · selected state 0
| Method | Episodes | Real steps | Avg return (10) |
|---|---|---|---|
| Q-learning | 0 | 0 | — |
What this experiment includes ↗
Cliff cells return the agent to the start with reward −100; other cliff-world moves cost −1. Maze moves cost −0.02 and the goal pays +1. Grid traps end an episode. Episodes are capped at 500 moves, with bootstrapping across time truncation. Tables start at zero; action ties are randomized. Planning reuses a learned deterministic model. Double Q alternates which table selects and which evaluates.
Source / ToolJar concept reference ↗Deep Q creatures · Open Rajiv’s original Eat Melon demo ↗