Research blog / July 30, 2026

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

A synchronized home-scene capture system for turning everyday human activity into structured perception, action, and contact data.

Technical report · soon Hugging Face · dataset
Overview of ACE-Data-0 showing ambient capture scenes, data modalities, annotations, and pipeline.
We introduce Ambient Capture Engine (ACE), a capture system that transforms real home environments into recording studios. ACE observes household activity at two complementary scales, table-scale and room-scale. Via ACE, we capture and release ACE-Data-0 with synchronized multiple modalities and rich data annotations.

Demo video

See everyday activity become embodied-AI data.

ACE keeps objects, body motion, camera views, tactile, and audio aligned to one timeline, turning a room-scale scene into supervision a model can learn from.

ACE-Data-0 teaser: synchronized multi-modal capture for embodied AI.

Motivation

Embodied AI needs physical-world data from household activities, yet such data is hard to capture.

Household activity is not a single atomic motion but a long-horizon process. Making tea, wiping a counter, or carrying a cup across the room weaves together perception, locomotion, dexterous manipulation, contact, and sound, and it unfolds over minutes rather than seconds. What embodied AI needs from a dataset is exactly this: how all of these signals evolve together over the course of a real activity, not just how one action looks in a cropped clip.

Unfortunately, existing datasets capture only pieces of this picture. Egocentric video shows natural behavior and the person's observing perspective, but rarely carries measured body, hand, and object states. Motion-capture datasets provide precise trajectories, but mostly in sparse laboratory spaces stripped of the clutter, furniture, and occlusion that make homes difficult, and seldom with the first-person view, audio, or touch. And most interaction clips last seconds, while household goals take minutes and depend on what has already changed in the scene. The gap is not just scale but also completeness: few records contain ego view, exo views, body motion, hand articulation, object state, audio, and touch for the same event, in a real home, over a full activity.

ACE addresses this gap through long-horizon, goal-directed activity capture. Instead of prescribing a sequence of atomic actions, ACE specifies a household goal and allows participants to complete it naturally. A single activity can span multiple rooms, involve many objects, and contain a sequence of interdependent interactions. The resulting recordings preserve the full evolution of the activity—from the initial scene state to intermediate changes and final task completion—providing continuous trajectories for studying long-term planning, state tracking, and memory in embodied agents.

Equally important, ACE provides continuous synchronization across all modalities. Ego and exo video, body and hand motion, object states, audio, and tactile signals are synchronized to a shared temporal reference, so observations from different sensors correspond to the same physical event at any point in a take. This is particularly critical for long-horizon activities, where clock drift can accumulate over time and progressively break cross-modal correspondence. By maintaining alignment throughout the complete recording, ACE enables precise associations between visual observations, physical motion, contact, and sound over extended activities.

01

Holistic sensory record

Ego-view, exo-view, motion, object state, audio, and touch describe the same event.

02

Long-horizon activity

Goal-directed household tasks preserve planning, memory, and scene-state evolution.

03

Multi-modal synchronization

All streams stay aligned so visual, physical, contact, and audio signals remain paired.

ACE Capture System

Two complementary capture scales, one shared pipeline.

Household interaction spans very different spatial scales. Fingertip manipulation requires close-range observations of subtle grasp changes and contact, while household activities require room-scale coverage of locomotion, human-scene interaction, and long motion trajectories. A single capture setup cannot resolve both equally well. ACE is therefore built at two complementary spatial scales: one for fine-grained manipulation and another for room-scale embodied activity.

The table-scale configuration concentrates close-range cameras, optical motion capture, and tactile gloves around table-top workspaces. At this scale, small spatial changes matter: fingers make and break contact, grasps transition, tools move through short trajectories, and objects rotate within the hand. The close-range setup is designed to preserve these details, providing high-resolution observations of hand-object interaction and tactile events.

The room-scale configuration, on the other hand, extends capture across a furnished apartment with a kitchen, dining area, living room, and bedroom. Wide-baseline cameras and motion-capture coverage follow participants as they walk between spaces, interact with furniture, and manipulate objects over extended trajectories. Rather than constraining activities to avoid occlusion, the setup observes them from multiple viewpoints, preserving natural room-scale behavior in a complex home environment.

Despite their different spatial scales and sites, the two configurations share the same sensing and representation framework. Participants use the same egocentric, body-motion, hand-tracking, and tactile systems, while the environments provide synchronized exocentric video, human and object tracking, and multi-channel audio. A shared calibration, synchronization, and data representation pipeline then standardizes recordings across the two sites, allowing fine-grained manipulation and room-scale household activity to be studied within a unified dataset.

Table-scale configuration

Table-scale capture setup with dense close-range sensing around table-top workspaces.
It concentrates dense sensing around table-top workspaces for close-range hand-object manipulation.

Room-scale configuration

Room-scale capture setup spanning a furnished apartment with camera and motion-capture coverage.
It spans a furnished apartment for room-scale motion, both human-object and human-scene interaction, and wide-baseline capture.

Wearable sensing

Four ego cameras, IMU, headset pose, full-body motion, and hand articulation.

Scene sensing

Stable third-person views plus optical tracking for bodies, objects, and the room.

Physical sensing

Tactile pressure and multi-channel audio capture contact beyond what video can show.

Data capturing and processing

Each recording is aligned in time and space, not a folder of unrelated streams.

Multi-modal capture is only useful when different signals agree on when and where an event happens. ACE therefore aligns every sensor stream along two dimensions: synchronization establishes a common timeline, while calibration establishes a common spatial frame. Together, they make video, human and object motion, audio, and tactile signals directly comparable across modalities.

Temporal synchronization resolves the mismatch between sensors that operate at different rates and on different clocks. Without precise alignment, contact may appear in the video before it is registered by the tactile glove, or an object trajectory may lag behind the observed motion. ACE uses the motion-capture clock as a common temporal reference and associates each camera frame with the corresponding motion, audio, and tactile measurements.

Spatial calibration resolves the mismatch between sensors that observe the environment from different viewpoints and coordinate systems. Static exocentric cameras, the moving egocentric headset, tracked objects, and human motion are registered into a common world frame, so measurements from different modalities refer to the same physical location. Together, temporal synchronization and spatial calibration turn independent sensor streams into a coherent multi-modal activity record.

ACE workflow showing data collection, export and synchronization, and annotation.
The pipeline proceeds from multi-modal capture to temporal alignment, then to annotations derived from measured states.

Temporal alignment

A common clock aligns camera frames, motion states, tactile readings, and audio.

Spatial registration

Cameras, the ego rig, tracked objects, and human motion share one world frame.

Capture protocol

Each take follows setup, goal briefing, recording, and post-capture quality checks.

Modality stack

Five modalities travel together through every aligned take.

Once synchronized and calibrated, each ACE sequence becomes a paired record of observations, measured states, and physical interaction signals. Raw sensor streams are released together with camera intrinsics and poses, headset motion, object trajectories, scene geometry, and the shared timeline and coordinate frame that connect them.

This is what separates an embodied AI dataset from a collection of videos. Different modalities describe the same physical state at the same time and place, allowing one signal to serve as input and another as supervision without uncertain cross-modal correspondence.

EGO

Egocentric view

4 fisheye cameras, IMU, and tracked headset pose.

First-person observations for robot-like perception.
EXO

Exocentric view

Multi-view RGB from table-scale and room-scale cameras.

Body, scene, and object context beyond the actor's view.
3D

Motion state

Full-body, hand, and object trajectories at 60 Hz.

Metric supervision for state recovery and imitation.
TACT

Tactile signal

Full-palm tactile gloves recording pressure maps across the palm and fingers.

Contact evidence for grasp, slip, and pressure.
AUD

Multi-channel audio

Audio recorded by the egocentric headset and exocentric cameras.

Contact sounds, appliance events, and ambient cues.

Video data examples

Synchronized examples with ego-view and exo-view cameras.

Each ACE sequence provides paired egocentric and exocentric observations of the same activity. The ego view captures the scene from the participant's perspective, while the multi-view exocentric cameras preserve the surrounding environment, full-body motion, and object context.

The two views are complementary. Egocentric video resolves hands and nearby interactions but moves with the head and often loses whole-body context. Exocentric views provide stable scene-level observations, yet fine hand motion may be small or occluded. By recording both simultaneously, ACE preserves these complementary perspectives for the same physical events.

Data annotation

Annotation examples from the released annotation stack.

Each ACE sequence is paired with annotations of both human state and object state. Human and hand motion describe how the participant moves and interacts, while object trajectories and 6-DoF poses capture how the surrounding scene changes over time.

Most annotations are derived directly from the calibrated capture system. Body and hand poses, object poses, bounding boxes, motion trails, contact states, and camera projections are generated from measured physical states rather than inferred independently from video. Text descriptions are then aligned to the same timeline, adding a semantic layer to the multi-modal record.

Together, the annotation stack describes where the sensors are, where the human and objects are, when physical interaction occurs, and what activity is unfolding. Because the underlying geometry and motion come from calibrated measurements, these annotations remain consistent even under occlusion, clutter, and visual ambiguity.

Camera and timeline

Intrinsics, poses, headset trajectory, and synchronized timestamps.

Human and hand state

Body skeletons, hand articulation, and per-view projected overlays.

Object state

Object meshes, 6-DoF poses, boxes, and motion trails.

Contact labels

Hand-object contact grounded by motion state and tactile readings.

Language descriptions

Natural-language descriptions aligned with the physical timeline.

Dataset

ACE-Data-0: A Large-Scale Multi-Modal Dataset for Long-Horizon Household Activities

Through ACE, we release ACE-Data-0, a large-scale dataset of long-horizon, multi-modal, synchronized household activities. It comprises 150 hours of capture across 200 task categories, recorded as continuous takes that each run for minutes and are segmented into 75,000 interaction episodes. Each take preserves egocentric and exocentric observations together with human motion, hand motion, object states, audio, and tactile signals over a continuous activity timeline.

ACE-Data-0 is built around goal-directed activity capture. Participants receive household goals rather than fixed sequences of atomic actions and complete them naturally. The same goal may involve different routes, object choices, grasp strategies, subtask orders, pauses, and recoveries. Long-horizon behavior therefore emerges from task completion itself, rather than from manually concatenating short interaction clips.

The dataset contains three complementary types of activities. Atomic HOI takes focus on one to three everyday manipulation tasks each, such as pouring water, preparing food, or tidying a surface. Chain of HOI connects multiple atomic HOI into continuous household routines that require intermediate state tracking and sequential decision making. HSI emphasizes whole-body interaction with the environment, including locomotion, posture transitions, and contact with furniture and other scene structures.

Across all three categories, ACE-Data-0 preserves the complete context of an interaction. The approach, contact, manipulation, release, and resulting scene changes remain within the same timeline, so an action can be interpreted through both its history and its consequence. A grasp is not recorded as an isolated label, but as one event within an evolving human-object-scene state.

Natural variation is retained as part of the data rather than removed through scripting. Participants choose how to move, where to stand, which objects to use, and how to recover from small deviations, while object configurations and trajectories change with the activity. Because this variation remains grounded by synchronized, measured multi-modal signals, ACE-Data-0 captures household interaction as performed behavior rather than a collection of isolated action instances.

ACE-Data-0 Ambient-captured interaction dataset
Atomic HOI Single-task, long-horizon interactions
Drink Pour Fill Wash Clean Wipe Peel Chop Cook Stir Fold Tidy
Chain of HOI Composed household activities
Open fridge Take out Rinse Cut Cook Season Serve Clean up
HSI Human-scene interactions
Sit Navigate Exercise Open/close Pick/place Posture transition Scene arrangement
Task examples from ACE-Data-0 showing atomic HOI, human-scene interaction, chain of HOI, and total statistics panels.
Task examples: ACE-Data-0 spans atomic HOI, human-scene interaction, and chain of HOI household activities.
ACE-Data-0 statistics panel for task categories, duration, and distribution.
Overview statistics of ACE-Data-0: task categories, task composition, and the distribution of frames across atomic HOI, chain of HOI, and HSI takes.

Why this division matters

Short clips (atomic HOI) can show what a grasp looks like; long takes (chain of HOI) reveal what the grasp is for. A household routine has memory: objects move, surfaces become clean or cluttered, a person leaves one area and returns later, and the scene state changes over time. ACE keeps that continuity so perception, action, and physical state can be studied together.

This also makes the data harder in the right way. Errors can accumulate over time, objects can leave and re-enter view, and the meaning of an action can depend on what happened much earlier in the take. Those are exactly the conditions an embodied agent will face in a real home.

Benchmark

A diagnostic benchmark from signals, to states, to interactions.

ACE-Data-0 organizes evaluation as a bottom-up hierarchy of embodied perception. Rather than measuring task success, the benchmark isolates where an interaction understanding pipeline fails—from recovering low-level physical signals, to reconstructing scene components, to understanding coordinated human-object interaction.

The first level focuses on low-level signals, asking whether physical cues such as tactile contact can be recovered from visual observations. The second level evaluates scene components, including metric 3D recovery of human body motion in furnished home scenes. The third level moves to interaction understanding, where models must recover and reason about coordinated hand, body, and object motion from egocentric and exocentric observations.

This hierarchy is diagnostic by design. Errors at one level expose the missing information required by the next: missed contact obscures whether a grasp succeeds, inaccurate human or object states break physical correspondence, and lost world-space motion limits interaction understanding and downstream embodied planning. The benchmark therefore reveals what a model understands, and where that understanding breaks down.

All three levels are evaluated on the same synchronized and calibrated recordings. Low-level signals, 3D scene states, and interaction trajectories refer to the same physical events in a shared timeline and coordinate frame. This allows predictions across different tasks and viewpoints to be compared against consistent measured ground truth, turning ACE-Data-0 into a unified benchmark for embodied perception and interaction understanding.

Three benchmark levels: low-level signals, scene components, and interaction.
The three levels mirror the capabilities an embodied agent must chain together: perceive, understand, and interact.
01

Low-level signal inference

Predict tactile readings and contact events from video, especially when the hand hides the fingertips at the exact moment contact matters. ACE pairs the video frame with the tactile reading from the same instant, so the target is measured contact rather than a visual proxy.

This track asks whether an abundant sensor, vision, can recover a scarce one, touch. It measures not only when contact occurs, but also where pressure is distributed across the hand, which is where current methods remain weakest.

02

Scene component recovery

Estimate human body motion under furniture occlusion, uncommon household postures, and minutes of continuous movement. The same measured ground truth compares per-frame, temporal, and scene-aware methods across single-view, multi-view, and egocentric inputs.

03

Embodied interaction understanding

Recover hand motion and trajectories from egocentric and exocentric views, then compare complementary failure modes under the same ground truth. Ego-view captures close hand detail; exo-view preserves global trajectory and scene context.

This level connects perception to action. The hand approach, grasp, adjustment, and release are the parts of human behavior most directly useful for robot imitation and manipulation policy learning.

Conclusion

ACE-Data-0 is an invitation to ambient capture for embodied AI.

Embodied intelligence requires understanding activities as evolving physical processes, not isolated visual clips. An agent must track how the scene changes, how humans and objects move and interact, when contact occurs, and how individual actions contribute to a longer-term goal. ACE captures these signals as a unified, synchronized, and calibrated record of household activity.

Together, the two ACE configurations span fine-grained manipulation and room-scale embodied behavior. Their complementary capture scales connect detailed hand-object interaction with whole-body motion, scene contact, and long activity trajectories, providing measured observations of household interaction across different spatial and temporal scales.

With ACE-Data-0, we aim to support embodied models that reason over long horizons, recover physical states, learn from natural behavioral variation, and connect perception to interaction in everyday environments. The accompanying benchmark evaluates these capabilities from low-level signal recovery to scene reconstruction and interaction understanding, treating them as connected stages of embodied perception rather than isolated tasks.

Ambient capture turns everyday activity into aligned supervision.

ACE provides the capture system, the synchronized dataset, and the diagnostic benchmark needed to study household interaction as a continuous embodied experience.