Every chapter carries the same three layers: complete mathematics, an interactive visual for each hard idea, and working code.
Part I
Foundations of Sequential Decision-Making
Sutton & Barto's spine, retold with robots and interactive math.
- 01Why Reinforcement Learning for Robotics?The see–think–act loop, why hand-coding breaks, and a tour of real-world successes graded L0–L5.Kober §1Tang §1–3Akinola
- 02The Mathematical ToolkitProbability, linear algebra, ODE discretization, contraction mappings and the Robbins–Monro theorem that makes every RL algorithm tick.S&B prerequisitesKober §1.3
- 03Multi-Armed Bandits: Exploration & ExploitationThe atom of RL: action values, regret, UCB derived from Hoeffding, gradient bandits as the seed of policy gradients.S&B ch. 2
- 04Markov Decision Processes: The FormalismThe MDP tuple, Bellman equations derived in full, γ-contraction proofs, and POMDPs — because robots never see state.S&B ch. 3Tang §3.2
- 05Dynamic Programming: Planning with a Known ModelPolicy evaluation, the policy improvement theorem, value iteration, and generalized policy iteration — the book's master pattern.S&B ch. 4
- 06Learning from Experience: Monte Carlo & Temporal-DifferenceLearning without a model: MC, importance sampling, TD(0), SARSA vs Q-learning, and the maximization bias that Double Q fixes.S&B ch. 5–6
- 07Unifying Learning & Planning: n-step, Traces, Dyna & MCTSThe λ-dial between TD and MC, eligibility traces, Dyna-Q learning from imagination, and MCTS at decision time.S&B ch. 7–8, 12
Part II
Scaling Up: Function Approximation & Deep RL
From tables to tensors — the leap robots require.
- 08Function Approximation & the Deadly TriadWhy robots outgrow tables: semi-gradient TD, tile coding, and Baird’s counterexample diverging live.S&B ch. 9–11Kober §4.2
- 09Deep Value-Based Methods: DQN & DescendantsReplay buffers and target networks as variance surgery; Double, Dueling, PER, Rainbow, and distributional RL.Mnih 2015Tang Table 5
- 10Policy Gradients: REINFORCE → PPOThe policy gradient theorem derived step by step, baselines and GAE, trust regions, and PPO — the workhorse of real robot RL.S&B ch. 13Kober §2.2.2Schulman 2017
- 11Off-Policy Continuous Control: DDPG, TD3 & SACSample efficiency as a robotics imperative: the DPG theorem, overestimation bias, and maximum-entropy RL.Silver 2014Fujimoto 2018Haarnoja 2018
- 12Model-Based RL & World ModelsMental rehearsal: ensemble dynamics, model-bias compounding bounds, CEM-MPC, and learning inside a latent dream.Kober §6Chua 2018Hafner 2023
Part III
The Robotics Side
Kober's bridge, rebuilt with modern materials.
- 13The Robot as an Environment: Kinematics, Dynamics & ControlWhat is inside the env black box when it is a robot: FK/IK, Jacobians, the manipulator equation, PID and LQR baselines.Kober §1.3Spong 2006Siciliano 2009
- 14The Four Curses of Robot RLDimensionality, real-world samples, under-modeling and goal specification — each quantified, plus the reward-hacking zoo.Kober §3Tang §5Ng 1999
- 15Simulation & the Sim-to-Real BridgeIntegrators, contact models, domain randomization as distributional robustness, system identification and teacher–student transfer.Kober §6.1Tang §3.3
- 16Demonstrations, Imitation & Offline RLBehavior cloning and its T² compounding error, DAgger, MaxEnt IRL, CQL/IQL pessimism, and offline-to-online fine-tuning.Kober §5Akinola LfDRoss 2011Kostrikov 2021
- 17Motor-Skill Policy RepresentationsWhat the policy outputs matters as much as how it learns: action-space levels, DMPs, CPGs, residual RL, and ball-in-a-cup then vs now.Kober §4, §7Ijspeert 2013Tang §3.2
Part IV
Competencies: RL on Real Robots
Tang's taxonomy as deep dives: what worked, why, and rebuilt in Rust.
- 18Learning LocomotionDRL’s flagship competency: reward anatomy, terrain curricula, teacher–student privileged learning, and Ferris learning to walk.Tang §4.1Lee 2020Kumar 2021
- 19Learning Navigation & Mobile ManipulationBelief MDPs and recurrent policies, end-to-end vs modular pipelines, hierarchical skills, and whole-body control.Tang §4.2Tang §4.4
- 20Learning ManipulationThe hard competency: grasp wrench space, impedance action spaces, in-hand dexterity, and honest success-rate methodology.Tang §4.3OpenAI 2020
Part V
Frontiers & Capstone
Where the field is going — and one project that uses all of it.
- 21Frontiers: HRI, Multi-Robot & Foundation ModelsShared autonomy, Dec-POMDPs and CTDE, vision-language-action models, and a Kober-2013-vs-Tang-2024 open-problem diff.Tang §4.5–4.6, §5Kober §8
- 22Capstone: An End-to-End Learned Robot in RustOne project using the whole book: specify, formalize, randomize, train teacher–student PPO, evaluate honestly, and ship to the browser.Kober §7 disciplineTang §3.4 rubric
