Robotics & Automation
AI That Acts in the Physical World
Robotics is what happens when AI stops generating text and starts moving things. Every robot — industrial arm, self-driving car, surgical assistant, warehouse picker, humanoid prototype — runs the same closed loop: perceive the world, plan a next action, execute it, observe what changed, repeat. The difference from purely digital AI is that mistakes have physical consequences and the environment talks back. That changes the engineering, the safety story, and the economics in important ways.
- Perception: sensors (cameras, LiDAR, IMUs, force/touch) feed a model that estimates the state of the world
- Planning: the robot decides what to do — a path to walk, an object to grasp, a sequence of sub-tasks to execute
- Action: actuators (motors, grippers, wheels, joints) move the robot in the physical world
- Closed loop: every action changes the world, the sensors re-observe, and the cycle repeats — this is what makes robotics fundamentally different from a chatbot
Where Robotics Lives Today
Robotics is not one industry — it is four very different ones, each at a different stage of maturity. Industrial robotics has been profitable for forty years. Autonomous vehicles are scaling driverless services in specific cities. Humanoids are the 2025–2027 frontier, with multiple billion-dollar companies racing to ship general-purpose robots into homes and factories. Service robotics — surgical, agricultural, inspection — is the quiet category where individual deployments are worth millions per unit.
- Industrial & warehouse: mature and profitable — assembly lines, pick-and-place, Amazon fulfilment, collaborative robots (cobots) on the factory floor
- Autonomous vehicles: Waymo expanding driverless service, Tesla FSD shipping in cars, trucking and middle-mile delivery emerging fast
- Humanoid robots: Tesla Optimus, Figure, 1X, Boston Dynamics Atlas — the bet that a general-purpose human form factor unlocks home and labour markets
- Service & surgical: da Vinci surgical robots, cleaning fleets, agricultural pickers, inspection drones — high value, regulated, often invisible to consumers
Foundation Models for Robotics: VLAs
For most of robotics history, every behaviour was hand-coded by an engineer. Pick up a red block? Write a controller. Walk up a stair? Write another controller. The 2024–2026 shift is that foundation models — specifically Vision-Language-Action (VLA) models — are starting to do for robotics what LLMs did for text. A single trained model can be told "pick up the red mug and put it on the counter" and figure out the rest, generalising across robots, environments, and tasks it has never seen before.
- RT-2 (Google DeepMind): the first model to fuse vision, language, and robotic action in a single end-to-end network
- PaLM-E: shows that scaling laws apply to robotics — bigger models generalise to new robots and new tasks zero-shot
- π0 (Physical Intelligence): cross-embodiment foundation model trained across many robot types and tasks
- OpenVLA: open-weight VLA, lowering the barrier for academic and industrial labs to build on top of foundation models
- The shift: from "program every behaviour" to "give the robot a goal in natural language and let it figure out the steps"
Prefer slides, quizzes, and saved progress? Read this lesson in the library — free, no sign-up.