Motivation
Embodied AI needs physical-world data from household activities, yet such data is hard to capture.
Household activity is not a single atomic motion but a long-horizon process. Making tea, wiping a counter, or carrying a cup across the room weaves together perception, locomotion, dexterous manipulation, contact, and sound, and it unfolds over minutes rather than seconds. What embodied AI needs from a dataset is exactly this: how all of these signals evolve together over the course of a real activity, not just how one action looks in a cropped clip.
Unfortunately, existing datasets capture only pieces of this picture. Egocentric video shows natural behavior and the person's observing perspective, but rarely carries measured body, hand, and object states. Motion-capture datasets provide precise trajectories, but mostly in sparse laboratory spaces stripped of the clutter, furniture, and occlusion that make homes difficult, and seldom with the first-person view, audio, or touch. And most interaction clips last seconds, while household goals take minutes and depend on what has already changed in the scene. The gap is not just scale but also completeness: few records contain ego view, exo views, body motion, hand articulation, object state, audio, and touch for the same event, in a real home, over a full activity.
ACE addresses this gap through long-horizon, goal-directed activity capture. Instead of prescribing a sequence of atomic actions, ACE specifies a household goal and allows participants to complete it naturally. A single activity can span multiple rooms, involve many objects, and contain a sequence of interdependent interactions. The resulting recordings preserve the full evolution of the activity—from the initial scene state to intermediate changes and final task completion—providing continuous trajectories for studying long-term planning, state tracking, and memory in embodied agents.
Equally important, ACE provides continuous synchronization across all modalities. Ego and exo video, body and hand motion, object states, audio, and tactile signals are synchronized to a shared temporal reference, so observations from different sensors correspond to the same physical event at any point in a take. This is particularly critical for long-horizon activities, where clock drift can accumulate over time and progressively break cross-modal correspondence. By maintaining alignment throughout the complete recording, ACE enables precise associations between visual observations, physical motion, contact, and sound over extended activities.
Holistic sensory record
Ego-view, exo-view, motion, object state, audio, and touch describe the same event.
Long-horizon activity
Goal-directed household tasks preserve planning, memory, and scene-state evolution.
Multi-modal synchronization
All streams stay aligned so visual, physical, contact, and audio signals remain paired.
ACE Capture System
Two complementary capture scales, one shared pipeline.
Household interaction spans very different spatial scales. Fingertip manipulation requires close-range observations of subtle grasp changes and contact, while household activities require room-scale coverage of locomotion, human-scene interaction, and long motion trajectories. A single capture setup cannot resolve both equally well. ACE is therefore built at two complementary spatial scales: one for fine-grained manipulation and another for room-scale embodied activity.
The table-scale configuration concentrates close-range cameras, optical motion capture, and tactile gloves around table-top workspaces. At this scale, small spatial changes matter: fingers make and break contact, grasps transition, tools move through short trajectories, and objects rotate within the hand. The close-range setup is designed to preserve these details, providing high-resolution observations of hand-object interaction and tactile events.
The room-scale configuration, on the other hand, extends capture across a furnished apartment with a kitchen, dining area, living room, and bedroom. Wide-baseline cameras and motion-capture coverage follow participants as they walk between spaces, interact with furniture, and manipulate objects over extended trajectories. Rather than constraining activities to avoid occlusion, the setup observes them from multiple viewpoints, preserving natural room-scale behavior in a complex home environment.
Despite their different spatial scales and sites, the two configurations share the same sensing and representation framework. Participants use the same egocentric, body-motion, hand-tracking, and tactile systems, while the environments provide synchronized exocentric video, human and object tracking, and multi-channel audio. A shared calibration, synchronization, and data representation pipeline then standardizes recordings across the two sites, allowing fine-grained manipulation and room-scale household activity to be studied within a unified dataset.
Table-scale configuration
Room-scale configuration
Wearable sensing
Four ego cameras, IMU, headset pose, full-body motion, and hand articulation.
Scene sensing
Stable third-person views plus optical tracking for bodies, objects, and the room.
Physical sensing
Tactile pressure and multi-channel audio capture contact beyond what video can show.
Data capturing and processing
Each recording is aligned in time and space, not a folder of unrelated streams.
Multi-modal capture is only useful when different signals agree on when and where an event happens. ACE therefore aligns every sensor stream along two dimensions: synchronization establishes a common timeline, while calibration establishes a common spatial frame. Together, they make video, human and object motion, audio, and tactile signals directly comparable across modalities.
Temporal synchronization resolves the mismatch between sensors that operate at different rates and on different clocks. Without precise alignment, contact may appear in the video before it is registered by the tactile glove, or an object trajectory may lag behind the observed motion. ACE uses the motion-capture clock as a common temporal reference and associates each camera frame with the corresponding motion, audio, and tactile measurements.
Spatial calibration resolves the mismatch between sensors that observe the environment from different viewpoints and coordinate systems. Static exocentric cameras, the moving egocentric headset, tracked objects, and human motion are registered into a common world frame, so measurements from different modalities refer to the same physical location. Together, temporal synchronization and spatial calibration turn independent sensor streams into a coherent multi-modal activity record.
Temporal alignment
A common clock aligns camera frames, motion states, tactile readings, and audio.
Spatial registration
Cameras, the ego rig, tracked objects, and human motion share one world frame.
Capture protocol
Each take follows setup, goal briefing, recording, and post-capture quality checks.
Modality stack
Five modalities travel together through every aligned take.
Once synchronized and calibrated, each ACE sequence becomes a paired record of observations, measured states, and physical interaction signals. Raw sensor streams are released together with camera intrinsics and poses, headset motion, object trajectories, scene geometry, and the shared timeline and coordinate frame that connect them.
This is what separates an embodied AI dataset from a collection of videos. Different modalities describe the same physical state at the same time and place, allowing one signal to serve as input and another as supervision without uncertain cross-modal correspondence.
Egocentric view
4 fisheye cameras, IMU, and tracked headset pose.
First-person observations for robot-like perception.Exocentric view
Multi-view RGB from table-scale and room-scale cameras.
Body, scene, and object context beyond the actor's view.Motion state
Full-body, hand, and object trajectories at 60 Hz.
Metric supervision for state recovery and imitation.Tactile signal
Full-palm tactile gloves recording pressure maps across the palm and fingers.
Contact evidence for grasp, slip, and pressure.Multi-channel audio
Audio recorded by the egocentric headset and exocentric cameras.
Contact sounds, appliance events, and ambient cues.Video data examples
Synchronized examples with ego-view and exo-view cameras.
Each ACE sequence provides paired egocentric and exocentric observations of the same activity. The ego view captures the scene from the participant's perspective, while the multi-view exocentric cameras preserve the surrounding environment, full-body motion, and object context.
The two views are complementary. Egocentric video resolves hands and nearby interactions but moves with the head and often loses whole-body context. Exocentric views provide stable scene-level observations, yet fine hand motion may be small or occluded. By recording both simultaneously, ACE preserves these complementary perspectives for the same physical events.
The same moment is observed from the actor's perspective and from eight synchronized exocentric cameras.
A clip from a 20-minute room-scale long-horizon sequence with synchronized multi-modal data.
A synchronized HSI sequence follows the participant's ego view and eight fixed exocentric cameras across the full interaction.
Data annotation
Annotation examples from the released annotation stack.
Each ACE sequence is paired with annotations of both human state and object state. Human and hand motion describe how the participant moves and interacts, while object trajectories and 6-DoF poses capture how the surrounding scene changes over time.
Most annotations are derived directly from the calibrated capture system. Body and hand poses, object poses, bounding boxes, motion trails, contact states, and camera projections are generated from measured physical states rather than inferred independently from video. Text descriptions are then aligned to the same timeline, adding a semantic layer to the multi-modal record.
Together, the annotation stack describes where the sensors are, where the human and objects are, when physical interaction occurs, and what activity is unfolding. Because the underlying geometry and motion come from calibrated measurements, these annotations remain consistent even under occlusion, clutter, and visual ambiguity.
Camera and timeline
Intrinsics, poses, headset trajectory, and synchronized timestamps.
Human and hand state
Body skeletons, hand articulation, and per-view projected overlays.
Object state
Object meshes, 6-DoF poses, boxes, and motion trails.
Contact labels
Hand-object contact grounded by motion state and tactile readings.
Language descriptions
Natural-language descriptions aligned with the physical timeline.
Human pose annotations align egocentric video, multi-view exocentric video, and 3D human motion across the same activity timeline.
Object annotations expose object identity, spatial state, 6-DoF pose, and motion trails in the same synchronized timeline.
Hand annotations focus on hand state, articulation, and projected overlays aligned with the physical activity record.
Tactile annotations add contact evidence that complements visual observations and motion state during manipulation.
Language annotations provide semantic descriptions aligned to the same measured activity timeline.
Dataset
ACE-Data-0: A Large-Scale Multi-Modal Dataset for Long-Horizon Household Activities
Through ACE, we release ACE-Data-0, a large-scale dataset of long-horizon, multi-modal, synchronized household activities. It comprises 150 hours of capture across 200 task categories, recorded as continuous takes that each run for minutes and are segmented into 75,000 interaction episodes. Each take preserves egocentric and exocentric observations together with human motion, hand motion, object states, audio, and tactile signals over a continuous activity timeline.
ACE-Data-0 is built around goal-directed activity capture. Participants receive household goals rather than fixed sequences of atomic actions and complete them naturally. The same goal may involve different routes, object choices, grasp strategies, subtask orders, pauses, and recoveries. Long-horizon behavior therefore emerges from task completion itself, rather than from manually concatenating short interaction clips.
The dataset contains three complementary types of activities. Atomic HOI takes focus on one to three everyday manipulation tasks each, such as pouring water, preparing food, or tidying a surface. Chain of HOI connects multiple atomic HOI into continuous household routines that require intermediate state tracking and sequential decision making. HSI emphasizes whole-body interaction with the environment, including locomotion, posture transitions, and contact with furniture and other scene structures.
Across all three categories, ACE-Data-0 preserves the complete context of an interaction. The approach, contact, manipulation, release, and resulting scene changes remain within the same timeline, so an action can be interpreted through both its history and its consequence. A grasp is not recorded as an isolated label, but as one event within an evolving human-object-scene state.
Natural variation is retained as part of the data rather than removed through scripting. Participants choose how to move, where to stand, which objects to use, and how to recover from small deviations, while object configurations and trajectories change with the activity. Because this variation remains grounded by synchronized, measured multi-modal signals, ACE-Data-0 captures household interaction as performed behavior rather than a collection of isolated action instances.
Why this division matters
Short clips (atomic HOI) can show what a grasp looks like; long takes (chain of HOI) reveal what the grasp is for. A household routine has memory: objects move, surfaces become clean or cluttered, a person leaves one area and returns later, and the scene state changes over time. ACE keeps that continuity so perception, action, and physical state can be studied together.
This also makes the data harder in the right way. Errors can accumulate over time, objects can leave and re-enter view, and the meaning of an action can depend on what happened much earlier in the take. Those are exactly the conditions an embodied agent will face in a real home.
Benchmark
A diagnostic benchmark from signals, to states, to interactions.
ACE-Data-0 organizes evaluation as a bottom-up hierarchy of embodied perception. Rather than measuring task success, the benchmark isolates where an interaction understanding pipeline fails—from recovering low-level physical signals, to reconstructing scene components, to understanding coordinated human-object interaction.
The first level focuses on low-level signals, asking whether physical cues such as tactile contact can be recovered from visual observations. The second level evaluates scene components, including metric 3D recovery of human body motion in furnished home scenes. The third level moves to interaction understanding, where models must recover and reason about coordinated hand, body, and object motion from egocentric and exocentric observations.
This hierarchy is diagnostic by design. Errors at one level expose the missing information required by the next: missed contact obscures whether a grasp succeeds, inaccurate human or object states break physical correspondence, and lost world-space motion limits interaction understanding and downstream embodied planning. The benchmark therefore reveals what a model understands, and where that understanding breaks down.
All three levels are evaluated on the same synchronized and calibrated recordings. Low-level signals, 3D scene states, and interaction trajectories refer to the same physical events in a shared timeline and coordinate frame. This allows predictions across different tasks and viewpoints to be compared against consistent measured ground truth, turning ACE-Data-0 into a unified benchmark for embodied perception and interaction understanding.
Low-level signal inference
Predict tactile readings and contact events from video, especially when the hand hides the fingertips at the exact moment contact matters. ACE pairs the video frame with the tactile reading from the same instant, so the target is measured contact rather than a visual proxy.
This track asks whether an abundant sensor, vision, can recover a scarce one, touch. It measures not only when contact occurs, but also where pressure is distributed across the hand, which is where current methods remain weakest.
Scene component recovery
Estimate human body motion under furniture occlusion, uncommon household postures, and minutes of continuous movement. The same measured ground truth compares per-frame, temporal, and scene-aware methods across single-view, multi-view, and egocentric inputs.
Embodied interaction understanding
Recover hand motion and trajectories from egocentric and exocentric views, then compare complementary failure modes under the same ground truth. Ego-view captures close hand detail; exo-view preserves global trajectory and scene context.
This level connects perception to action. The hand approach, grasp, adjustment, and release are the parts of human behavior most directly useful for robot imitation and manipulation policy learning.
Conclusion
ACE-Data-0 is an invitation to ambient capture for embodied AI.
Embodied intelligence requires understanding activities as evolving physical processes, not isolated visual clips. An agent must track how the scene changes, how humans and objects move and interact, when contact occurs, and how individual actions contribute to a longer-term goal. ACE captures these signals as a unified, synchronized, and calibrated record of household activity.
Together, the two ACE configurations span fine-grained manipulation and room-scale embodied behavior. Their complementary capture scales connect detailed hand-object interaction with whole-body motion, scene contact, and long activity trajectories, providing measured observations of household interaction across different spatial and temporal scales.
With ACE-Data-0, we aim to support embodied models that reason over long horizons, recover physical states, learn from natural behavioral variation, and connect perception to interaction in everyday environments. The accompanying benchmark evaluates these capabilities from low-level signal recovery to scene reconstruction and interaction understanding, treating them as connected stages of embodied perception rather than isolated tasks.
ACE provides the capture system, the synchronized dataset, and the diagnostic benchmark needed to study household interaction as a continuous embodied experience.