Everyone is talking about robots again.
Humanoids. Agents with arms. “LLMs controlling the physical world.”
Nice demos.
But we’re still missing the thing that actually matters.
Robots don’t fail because they can’t talk. They fail because they don’t understand reality.
A 17-year-old can learn to drive in about 20 hours. Yet our best autonomous systems still need:
- perfect maps
- constrained zones
- heavy sensors
- controlled conditions
That gap is not about data scale.
It’s about how intelligence works.
Language was the easy part
Language models succeeded fast for a simple reason.
Language is already compressed reality.
Words are discrete. Symbols are stable. Meaning is mostly predictable from context.
The physical world is none of that.
It’s:
- continuous
- noisy
- ambiguous
- partially observable
- full of edge cases
You don’t “autocomplete” reality.
You anticipate consequences.
And that difference changes everything.
The real inversion people miss
Language AI predicts what comes next.
Real intelligence predicts what will happen next if I act.
That sounds subtle. It’s not.
This is the difference between:
- generating text
- and planning in the world
Between:
- reacting
- and deciding
If you remember one sentence, remember this one:
Intelligence is prediction in the service of action.
Why “agents everywhere” is the wrong mental model
Most current “agentic” systems work when:
- the environment is scripted
- the task repeats
- failure is cheap
That’s fine for workflows.
It collapses in the real world.
Reality doesn’t follow scripts. It changes. It surprises you. It punishes mistakes.
If a system cannot predict the consequences of its actions, it is not autonomous.
It is just executing.
World models: not buzzwords, the missing layer
A world model is simple to describe:
Given the current state of the world, and an action I might take, can I predict what happens next?
Not at pixel level. Not with fake precision.
At the right level of abstraction.
That is the key.
You don’t plan a trip by simulating every muscle movement. You plan hierarchically:
- goal
- sub-goals
- actions
Humans do this constantly.
AI systems mostly don’t.
Planning is the forgotten superpower
Planning is what separates:
- pattern recognition
- from intelligence
It requires:
- a model of the world
- a model of yourself
- a sense of time
- the ability to imagine futures
This is why animals can solve new problems instantly.
A child can clear a table the first time. A beginner can learn to ski. A teenager can learn to drive.
Not because they were trained on the task.
Because they have a world model.
Watch first. Act later.
The important shift is this:
You don’t start by acting. You start by observing.
Learn how the world behaves. What’s stable? What’s predictable. What violates expectations?
Only then do you connect actions to outcomes.
This is how biology works. It’s also how scalable physical AI will work.
Observation dominates. Action fine-tunes.
Reinforcement learning is not the hero
Reinforcement learning has a role.
But it’s the polish, not the foundation.
Learning by trial and error in the real world is:
- slow
- dangerous
- expensive
No system should learn not to drive off a cliff by driving off cliffs.
The heavy lifting comes from self-supervised learning:
- building representations
- learning dynamics
- predicting change
Reinforcement adjusts. It does not create intelligence.
Energy is the quiet constraint
Here’s a reality check people avoid.
The human brain runs on ~20 watts.
Modern AI systems need data centers.
Why?
Because today’s hardware:
- moves data constantly
- separates memory and compute
- wastes energy shuffling information
Brains don’t do that.
They compute where the data lives. Massively parallel. Slow, but efficient.
Physical AI will not scale without rethinking this.
This is not sci-fi. It’s physics and engineering.
This is not about humanoids
Humanoids are the headline.
World models are the value.
The real impact is in:
- factories
- logistics
- energy systems
- industrial operations
- maintenance
- infrastructure
Anywhere decisions meet physics.
The companies that win won’t have the best chatbot.
They’ll own the best operational understanding of reality.
The next AI wave won’t talk more
It will decide better
Chatbots made AI visible.
World models will make AI consequential.
Because language is an interface.
Reality is the operating system.
And intelligence is the ability to run on it.
Robots are not coming to replace us.
Systems that understand how the world works are.
Reality is the ultimate reasoning system.
First published in the OG Approved newsletter on 24/02/2026. Read it on Substack or subscribe to get the next one.


