Guides

Deep-dives on the ideas behind MoreSight™ and GINN — geometric priming, robotics foundation models, and where classical vision ends and spatial intelligence begins.

What is spatial intelligence?

Spatial intelligence is the ability to perceive, represent, and reason about the three-dimensional structure and physical behavior of the world — not as pixels, but as geometry, objects, and motion. For robots, this is the difference between guessing from images and understanding from structure.

Classical computer vision identifies what a camera sees. Spatial intelligence goes further: it reconstructs where things are, how they are oriented, how they move, and how they can be manipulated. That representation is what makes fast, reliable skill mastery possible on edge devices.

Robotics foundation models and the sim-to-real gap

Robotics foundation models promise general-purpose robot control learned from large, diverse datasets. But many of today's RFMs are built on the same paradigm as large language models: they tokenize pixels and predict actions from language and image prompts. That works well for text, but the physical world is continuous, noisy, and governed by geometry and physics.

This is why the sim-to-real gap remains so stubborn. A model trained on internet-scale video or simulation images learns patterns, not physical laws. When lighting shifts, textures change, or a part is rotated a few degrees, pixel similarity breaks down and the robot fails. Geometric priming — the approach behind MoreSight™ and GINN — sidesteps that fragility by grounding perception in depth, geometry, and physics.

The guides below explore each layer of that shift: from VLA model comparisons like OpenVLA vs. RT-2, to the role of 3D datasets like Objaverse-XL, to the difference between spatial intelligence and classical computer vision. Each one is written for robotics researchers and engineers who want to move beyond pixel-level imitation and build robots that master real-world skills.