Continuous Control of Lunar Lander Using Deep Reinforcement Learning (TD3)
Teaching an Agent to Land a Rocket
We trained a reinforcement learning agent to solve the continuous control variant of Lunar Lander-v3 from Gymnasium. In this post, I’ll walk you through how our agent successfully solves the environment using Twin Delayed Deep Deterministic Policy Gradient (TD3).
The finished agent iteratively learning to solve the enviornment
Environment Dynamics
The task is treated as a Markov decision process with an 8-dimensional state space and a 2-dimensional continuous action space. The environment presents several unique challenges:
- Dense Reward Shaping: A dense shaping reward is paid on every step, encouraging the lander to act centered, slow, and upright.
- Sparse Terminal Reward: A sparse terminal reward of +100 for rest or -100 for crashing is only paid at the end.
- Action Dead Zones: Side boosters only fire beyond ±0.5, creating flat policy gradients near initialization.
- Irreducible Variance: Every reset applies a random initial force, meaning even a flawless policy occasionally draws an unrecoverable start.
Hyperparameter Sensitivity & Behavioral Impact
We systematically swept three algorithmic hyperparameters against the defaults from Fujimoto et al. (2018) to observe how fine-tuning alters physical behavior in the environment:
1. Soft Target-Update Rate ()
Controls target network dynamics and the speed at which the bootstrap target moves.
- Behavioral Effect: High values lead to unstable, jerky corrections, causing the lander to overshoot its descent vector.
2. Policy Delay ()
Governs the actor-critic coupling by determining how often the actor updates relative to the critic.
- Behavioral Effect: Setting (no delay) causes a complete collapse. The actor perpetually chases transient critic errors, leading to wild spinning and rapid structural failure.
3. Exploration Noise ()
Acts as the sole exploration mechanism since the TD3 policy remains deterministic during training.
- Behavioral Effect: Insufficient noise () leaves the agent trapped inside action dead zones, unaware that firing the side boosters past the threshold is a viable mechanism to stabilize lateral drift.
Tested on 100 fresh landings with the learning turned off, the finished agent averaged 241 — comfortably above passing. The occasional low score isn’t a mistake by the agent; the simulator gives every drop a random shove, and once in a while it’s simply unrecoverable. That’s exactly why the bar is an average over 100 tries rather than a single perfect run.