Rajiv Shah /ML / AI ARCADETeaching ↗
STUDY / 04

Explore Reinforcement Learning

Learn better actions through rewards.

Tabular / grid

Can rewards reveal a route?

Press Step and watch the TD target update one action value. Then train and look for a route to the goal.

Q ← Q + α [r + γ max Q(next) − Q]

Q-learning

0 episodes
0.00last reward
0.00TD target
0.00TD error
0.00 0.00 0.00 0.00

Action values · selected state 0

011Episode return · rolling average (10)Episode 12
MethodEpisodesReal stepsAvg return (10)
Q-learning00
What this experiment includes

Cliff cells return the agent to the start with reward −100; other cliff-world moves cost −1. Maze moves cost −0.02 and the goal pays +1. Grid traps end an episode. Episodes are capped at 500 moves, with bootstrapping across time truncation. Tables start at zero; action ties are randomized. Planning reuses a learned deterministic model. Double Q alternates which table selects and which evaluates.

Source / ToolJar concept reference