Signal
Issue 10 · September 17, 2026 · 6 min read
Imitation Learning: How Robots Learn Jobs From 50 Demos
Imitation learning explained: behavior cloning, the DAgger problem, and how pi0 and ALOHA train robots from demos. Build one this week on GRADuateML.
"The ChatGPT moment for general robotics is just around the corner." - Jensen Huang (CEO, NVIDIA), CES 2025
A warehouse robot does not read a manual. Somebody drives it through the task a few dozen times, and a network copies the moves.
This week we're talking about imitation learning, the training recipe behind nearly every robot manipulation demo you saw this year, and yes, this can get you a very comfortable job.
Here's the idea: imitation learning is supervised learning on expert demonstrations. You record what a human operator did, observations paired with actions, then train a policy that maps each observation to the expert's action. No reward function, no trial-and-error exploration. The simplest form, behavior cloning, turns robotics into a regression or classification problem over recorded trajectories:
That is the same cross-entropy objective we covered in issue 8, pointed at motor commands instead of class labels. If you know reinforcement learning, notice what is missing: no reward, no environment interaction during training, no exploration. That is the whole appeal. Teleoperation data is expensive, but it is far cheaper than letting a physical arm explore.
The industry has gone all-in on this. Physical Intelligence's π0 policy was trained on more than 10,000 hours of demonstrations across 7 robot embodiments and 68 tasks, from folding laundry to bussing tables, on top of the open X-Embodiment dataset. Diffusion Policy, the Columbia and Toyota Research Institute method, reported a 46.9% average improvement over prior imitation methods across 15 manipulation tasks. The ALOHA bimanual rig learned fine tasks from roughly 50 human demonstrations each. Tesla says Optimus learns tasks from human video, and 1X runs remote "experts" whose sessions double as training data. Data vendors now quote 200 to 500 demonstrations per day per collection site. The bottleneck is no longer the algorithm. It is logistics: stations, operators, calibration, and quality control on the demos.
How does imitation learning work?
You collect a dataset of (observation, action) pairs from an expert, usually a human with a teleoperation rig or a VR controller. You train a network to map each observation, typically camera images plus joint states, to the expert's next action or a short chunk of future actions. At deployment the policy runs closed-loop: observe, predict, act, repeat.
The catch is the top StackExchange question on the topic: the DAgger problem. A cloned policy makes small mistakes, drifts into states the expert never visited, and has no idea what to do there. Errors compound along the trajectory. This is called covariate shift, and DAgger (Ross et al., 2011) is the fix: run the learner, have the expert label the states it actually reached, add those to the dataset, retrain. Modern robot labs use action chunking, diffusion and flow-matching policies, and huge demo counts to blunt the same failure.
Is imitation learning reinforcement learning?
No. Imitation learning needs demonstrations and no reward. Reinforcement learning needs a reward and no demonstrations. The two are often combined, with imitation providing a warm start and RL fine-tuning on top, but the training signals are different.
| Imitation learning | Reinforcement learning | |
|---|---|---|
| Training signal | Expert actions | Reward from the environment |
| Data source | Teleoperation, human video | Agent's own rollouts |
| Exploration | None | Required, often unsafe on hardware |
| Sample efficiency | High (tens to hundreds of demos) | Low (millions of steps) |
| Failure mode | Compounding error off-distribution | Reward hacking, instability |
| Where you see it | π0, ALOHA, Diffusion Policy | Locomotion, games, RLHF |
Now I'm going to explain what this means for you:
The Skill Employers EXPECT
Learn behavior cloning end to end. Robot learning teams assume you can:
- Build a demonstration dataset with aligned observations and actions at a fixed control rate
- Train a behavior cloning policy in PyTorch and pick the right output head (continuous regression vs discretized bins)
- Predict action chunks, not single steps, and explain why chunking reduces compounding error
- Evaluate by task success rate over many rollouts, never by training loss alone
- Name covariate shift on sight when a policy works for 3 seconds then drifts
If your only robotics metric is validation MSE, you have not tested the policy.
The Skill That Separates You
Learn DAgger and multimodal policies.
Real demonstrations are inconsistent: two operators grasp the same mug differently, and a mean-regression policy averages them into a grasp that hits neither. You should know how to:
- Run DAgger rounds with a scripted or human expert and show the success curve bend upward
- Train a diffusion or flow-matching policy head that can represent several valid actions
- Weigh demo quality vs quantity, the exact trade-off the π0 team documented
- Decide when to add RL fine-tuning on top of an imitation warm start
- Read a vision-language-action model card and identify the imitation objective under the transformer
The Project To Learn These Skills This Week
Clone an expert, break it, then fix it with DAgger. Environment: Gymnasium LunarLander (continuous) or the PushT benchmark from the Diffusion Policy release. LunarLander gives you a free expert; PushT gives you real human demos.
Requirements:
- Get an expert. Train PPO with Stable-Baselines3 until it solves the task, or use the PushT human dataset.
- Collect demos at 10, 25, 50, and 100 episodes. Train a behavior cloning MLP on each and plot success rate vs demo count.
- Add action chunking (predict the next 8 actions, execute 4). Compare success rate against single-step prediction at the same demo count.
- Break it on purpose: evaluate your 50-demo policy with small observation noise (5% of state range). Record how fast success collapses. That is compounding error.
- Fix it: run 3 rounds of DAgger using the PPO expert to relabel the states your policy actually visits. Plot success per round.
- Write a half-page note mapping your pipeline to π0 or ALOHA: where their teleoperation rig replaces your PPO expert, and why they need 10,000 hours when you needed 100 episodes.
This project teaches a practical lesson for production ML:
A policy is only as good as the states it has seen, and the job is getting it to see the right ones.
If you want to sharpen your machine learning skills even more, I also selected a challenge problem for you this week:
If you learned something from this newsletter, make sure to forward it to a friend.