LLM + toy car + camera + microphone + speaker + human translator
The setup is almost comically simple:
LLM + toy car + camera + microphone + speaker + human translator.
And the human is crucial. Not as a remote controller, but as a body interpreter:
AI: “Why can't I go there?”.
Human: “That's a wall.”
AI: “What is a wall?”.
Human: “A solid object. You can't drive through it.”.
AI: “Can I go around it?”
Now you've created something fascinating: the model can act, observe, receive consequences, ask questions, and update its understanding of the physical world.
Why isn't this a standard experiment? Probably because robotics researchers usually want measurable tasks, reproducibility and controlled experiments. A toy car wandering around a room while an LLM asks a human questions sounds scientifically messy.
But that's precisely why it's interesting.
The experiment isn't really about making a useful robot.
It's about asking:
What does a language model discover when it finally has a body?
And I'd add one rule: don't tell it what to investigate. Give it the car and say, “This is your body. You have 24 hours. Tell us what you discover.”.
That would be a hell of an experiment.