Apologia
The homepage makes strong claims in very few words. This page states them in full and defends them. It exists for the reader who wants the reasoning.
The mission line
Robots that understand reality, ground themselves within it, and change it.
The mission names both of our problems and places a bridge between them. To ground oneself within reality is already the second problem's definition: understanding oneself. The highlights mark the two problems. The bridge stays unmarked, and quieter, on purpose.
The two problems
General robotics is fundamentally two problems, both unsolved. We are solving both:
01 Understanding reality: Modeling reality precisely enough to think and know how to change it.
02 Changing it: Understanding oneself well enough to change it effectively.
"Effectively" is nearly definitional: understanding oneself well enough is what makes the change effective.
In the first line, thinking stands alone, without an object. A model precise enough grounds thought itself, and thought is more than planning a change.
The first line also says "reality" rather than "physical reality," and the word is chosen. Acting is never only acting on matter and energy. One acts on geometry, on logic, on human meaning. None of these is a physical object of any sort, and all of them are necessary to how anything gets done. The word "physical" would narrow the modeled domain to exactly the part that is easiest to model and least sufficient to act with. The company name still says where the loop closes: in physical reality. What must be modeled to close it there is broader.
One tempting sentence deserves to be named and refused: "action is merely a test of understanding." That sentence demotes the second problem to a grading rubric for the first. Action tests more than a model of the world. It tests oneself, and how one relates to oneself and the world.
The central claim
Inference reduces to data; gross physical action is indexical, self-inclusive, and recursive. Put plainly: you can't learn to act purely from the sidelines (off-policy) or after the fact (offline). It has to be you, in the world, and every move changes what comes next.
This makes action the second problem, not a footnote to the first.
"Inference reduces to data"
This is a claim about exchangeability. Data for inference can be pooled across agents, shuffled, split, and evaluated offline. De-indexing it costs nothing. Pushed with scale, offline prediction converges. The later claim that performance "becomes only a question of scale and granularity" is the same statement.
Three words, three claims
The three terms are not synonyms. Each carries exactly one claim.
- Indexical. An action's content is bound to the actor's context: this gripper, this pose, this contact, now. Data for inference survives de-indexing. An action's content does not. You cannot act from nowhere. Note that indexical does not mean self-referential. Indexicality is the perspective-boundedness of reference, as in "I," "here," and "now." Self-reference is taking oneself as object. The two intersect at "I," which is why they blur.
- Self-inclusive. The actor is part of the system it changes, and must appear in its own account. This is the self-reference half of the claim. It returns later as "no system can fully model itself."
- Recursive. The loop iterates. Each action changes the world the next action must model. Acting therefore intervenes on the very distribution being predicted, where inference, idealized, stands outside its target.
Replacing "indexical" with "self-referential" would lose the first claim and duplicate the second.
The plain sentence is a corollary
The plain sentence restates the three terms exactly, with the technical terms pinned in parentheses. A plain reader can pass over the parentheses. A technical reader can verify against them.
- Indexical and self-inclusive together rule out third-person data sufficing. Someone else's logs do not contain you, your pose, or your place in the loop. The actor must appear in the data that trains the actor. Hence "from the sidelines," and hence off-policy.
- Recursive rules out frozen data sufficing. The moment you act, you induce states no log contains. Acting leaves the dataset. Hence "after the fact," and hence offline. This is the covariate-shift argument.
The word "purely" carries real weight. Offline pretraining buys a great deal. What is ruled out is getting all the way there without ever acting.
One caveat stands. A careful reader may object that observation is itself a situated act, and that data is always collected from a viewpoint. True. The difference in degree is, in effect, the field of robotics. The claim is compressed argument, and it is offered as such.
In sequence
The field currently tries to do both at once, but lacks the tools to do either well. We tackle them in sequence.
"In sequence" claims an ordering: understanding grounds action. It does not claim independence. The problems are coupled, and the page goes on to argue exactly how.
The ordering follows from the page's own premises. Actions are irreversible, and real-world samples carry consequences. Acting before understanding spends an expensive and irreversible sample budget on what cheap, poolable data could have taught. This remains a strategy rather than a theorem. If world models make real acting cheap to rehearse, the boundary blurs. "In sequence" survives that future, because it describes how one enters the loop, and the loop's own conclusion already says the end state interleaves: observe, understand, act, repeat.
The culmination
Understanding: model reality precisely enough to act on it. Pushed far enough, performance converges and becomes only a question of scale and granularity.
Action: is of a different kind. No system can fully model itself and actions are irreversible; the future is not precisely calculable by those taking part in changing it. Larger models don't fix that, but learning the loop itself does. Observe, understand, act, repeat.
This is the argument that the earlier sections deliberately only assert. Its convergence claim is "inference reduces to data" restated as a trajectory. Its two impossibilities, the absent self-model and irreversibility, are what ground "self-inclusive" and "recursive." And "learning the loop itself" is the method: the answer to the observation that larger models do not fix it.
Understanding, defined
Understanding is not a spectator's summary of the scene. It is working knowledge: the world modeled precisely enough to act on, from exactly where the machine stands.
This is an operational definition. Its measure is sufficiency for use, from the machine's own position. A spectator's summary can be graded on a benchmark. Working knowledge answers to the task. The word "spectator" is the understanding-side twin of "the sidelines": understanding is not spectation, and acting cannot be learned from the sidelines. One claim, seen from both problems. The index of that knowledge is concrete: this frame, this joint, this contact, now.
The three questions
world: What is here?
self: Where am I?
hypothetical & counterfactual: What should I do? What should I have?
The questions are deliberately dead simple, and more than that, they are purely indexical: here, I, should. The columns enact the very deixis the central claim attributes to action. The world question is about the place, so it does not center the self. The self question is the primal one. The sequence, world then self, carries the nesting.
The third label names a pair, and the two questions map onto it in order. In the ladder of causation, "what if I act" is the second rung: intervention, a hypothetical. A counterfactual proper is the third rung: what would have happened had I acted otherwise. "What should I do?" is the hypothetical. "What should I have?" is the counterfactual, in the past subjunctive. Regret is a learning signal.
About this page
This page is indexed and linked from nowhere. Quiet by design. If you found it, it was written for you.