All cursus
Roboticsadvanced
Vision-Language-Action: How Robots Learn to Act
The frontier where robotics meets large multimodal models. This cursus builds the modern robot-learning stack from the ground up. First, why robots learn skills from demonstrations instead of hand-written controllers, and the compounding-error problem that makes it hard. Then the two ideas that fixed it: action chunking and diffusion policies, plus the pooled multi-robot datasets behind generalist control. Finally, vision-language-action models, RT-2, OpenVLA, and pi0, that put a web-pretrained brain behind the robot, with an honest look at what they can and cannot yet do.
0 of 3 lessons complete
Sign in to track progress and earn a certificate.

