Spatio-Temporal Embodied Intelligence

Understanding space, time, and interaction for embodied intelligence.

← Back to Physical AI

Building embodied agents that understand where interaction matters, how the 3D world is organized, and how it changes through action.

Humans do not interact with the world as a sequence of isolated images. We continuously reason about how they relate to one another, which parts of the environment matter for interaction, and how the world may change as we act.

Many robotic systems, however, still treat perception as a transient input to action. This makes it difficult to maintain a consistent understanding of the environment, reason beyond the current observation, or represent the many possible ways an interaction may unfold.

Our research explores how spatial and temporal structure can become an organizing principle for embodied intelligence. We study how robots can identify interaction-relevant structure, maintain efficient representations of the surrounding 3D world, and use these representations to generate adaptive behavior.


Spatial Intelligence for Interaction

Intelligent behavior depends not only on what a robot sees, but on where interaction can occur and how actions are constrained by the surrounding space.

Our research investigates models that organize action generation around this interaction-relevant spatial structure. Rather than treating an observation as an undifferentiated input, the goal is to focus computation on the parts of the environment that determine how the robot can act.

Across problems such as dexterous grasping and precise manipulation, we explore generative models that can represent multiple physically meaningful ways of interacting with the same scene, while remaining robust to visual variation.

Spatially conditioned action generation for dexterous robotic interaction

Spatial conditioning provides a way to express where action-relevant information lies in the scene and to use that structure when generating robot behavior. In dexterous grasping, for example, this allows a model to represent diverse multi-finger configurations rather than collapsing to a small set of solutions.

The same principle extends to visuomotor control. Even from a single RGB observation, a robot can learn to extract the spatial information that matters for precise interaction while ignoring distractors and other irrelevant visual changes.

Precise and distractor-robust manipulation from a single RGB observation

Persistent 3D World Understanding

Embodied intelligence requires an understanding of the world that persists beyond a single observation. We study compact representations of 3D structure and dynamics that allow robots to maintain and update their understanding of the environment over time.

Our goal is to move beyond static reconstruction toward action-oriented world representations that support perception, prediction, planning, and interaction.

Efficient persistent 3D representation of a robotic scene Learning dynamic 3D structure for robotic interaction

Toward Spatio-Temporal Embodied Intelligence

These directions reflect a common view of embodied intelligence: a robot should not experience the world as disconnected observations and actions.

Our long-term goal is to build spatio-temporal world models that capture the structure and dynamics of the physical world, and enable robots to use these models for perception, prediction, and action.

Such world models can allow embodied agents to understand dynamic 3D environments, anticipate the consequences of their actions, and generate adaptive behavior across a wide range of physical tasks.


Representative Publications

  1. Spatially Conditioned Diffusion Policy: Learning Precise and Robust Manipulation with a Single RGB Camera
    Seoyoon Kim*, Kanghyun Kim*, Dongwoo Ko, Yeong Jin Heo, Min Jun Kim
    CoRL 2026, accepted
    * Co-first authors

  2. SPLIT: SE(3)-diffusion via Local Geometry-based Score Prediction for 3D Scene-to-Pose-Set Matching Problems
    Kanghyun Kim, Min Jun Kim
    arXiv preprint, 2024