Cart Balancing Challenge

Train adaptive agents, tune forces and rewards, survive harder physics, replay decisions, and discover how reinforcement learning keeps the pole upright longer each round.

Live cart-pole arena

Use and in manual mode.
Ready
Balance time0.00 s
Reward0.0
Pole angle0.0°
ActionNone

Live training dashboard

Statistics update after every episode.
0
Episode
0.00 s
Longest balance
0.00
Average reward
0%
Success rate
1.000
Exploration rate
0
Training steps
0.0000
Learning loss
Latest failure
#AlgorithmDurationRewardStepsResultFailure
0.25×
0 / 0
Replay frame
VS
Choose two strategies to compare average duration, reward, and success rate.

Mode and agent

Current decision

Selected actionNone
Confidence0%
Decision typeWaiting
Expected return0.00
Important signalPole angle
ReasonStart an episode.
Decision trace ready.

Experiment configuration

Changes apply when the next episode starts.

Learning controls

0.0010.0800.5
0.50.9700.999
01.001
0.90.99201
00.030.3

Training schedule

Physics settings

39.820 m/s²
0.250.75 m1.5
0.030.15 kg1
0.31.0 kg3
210.0 N30
24.8 m8

Environment variation

0.5°±5.0°18°
00.0000.2
00 steps12
00.0150.25

Reward design

Challenge target

5 s30 s120 s
24°45°
Current target progress0%

Progressive challenges

Achievements

0 unlocked

Data, model, and reporting

Progress is stored locally in your browser. Exports contain no server-side data.

Learning guide

State, action, and reward

The state describes cart position, velocity, pole angle, and angular velocity. The agent chooses a force, then receives a reward.

Exploration and exploitation

Exploration tests unfamiliar actions. Exploitation chooses actions currently expected to produce the highest future reward.

Discount factor

A higher discount factor values future balance more strongly. A lower value emphasizes immediate rewards.

Browser algorithms

Tabular methods use discretized states. DQN-style modes use lightweight linear value approximators suitable for an interactive browser lesson.

Related Calculators

Maze Learning AgentGrid World ExplorerRobot Navigation ChallengeTreasure Hunt AgentTraffic Light ControllerMulti-Armed Bandit GameEnergy Management AgentWarehouse Robot GameAdaptive Game Opponent

Important Note: All the Calculators listed in this site are for educational purpose only and we do not guarentee the accuracy of results. Please do consult with other sources as well.