{
  "meta": {
    "en": {
      "title": "ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine",
      "description": "A research blog page for ACE-Data-0, presenting ambient capture as an embodied data engine."
    },
    "zh": {
      "title": "ACE-Data-0：以人为中心的环境式采集构建具身数据引擎",
      "description": "ACE-Data-0 研究博客：一个面向具身智能的同步、多模态、长时程家庭活动采集数据集。"
    }
  },
  "text": {
    "Skip to content": "跳到正文",
    "ACE Capture System": "ACE 采集系统",
    "Workflow": "工作流",
    "Dataset": "数据集",
    "Benchmark": "评测",
    "Conclusion": "总结",
    "Research blog / July 30, 2026": "研究博客 / 2026 年 7 月 30 日",
    "ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine": "ACE-Data-0：以人为中心的环境式采集构建具身数据引擎",
    "A synchronized home-scene capture system for turning everyday human activity into structured perception, action, and contact data.": "一个同步的家庭场景数据采集系统，将日常人类活动转化为结构化的感知、动作与触觉数据。",
    "Technical report · soon": "技术报告 · 即将发布",
    "Hugging Face · soon": "Hugging Face · 即将开放",
    "We introduce Ambient Capture Engine (": "我们提出 Ambient Capture Engine（",
    "), a capture system that transforms real home environments into recording studios. ACE observes household activity at two complementary scales, table-scale and room-scale. Via ACE, we capture and release ACE-Data-0 with synchronized multiple modalities and rich data annotations.": "），一个将真实家庭环境转化为可采集同步、且不同种类的空间智能数据的系统。ACE 在桌面尺度与房间尺度这两个互补尺度上观察家庭活动。通过 ACE，我们采集并发布 ACE-Data-0，包含同步的多模态数据与丰富的数据标注。",
    "and": "和",
    "Demo video": "演示视频",
    "See everyday activity become embodied-AI data.": "将日常活动转化为具身智能数据。",
    "ACE keeps objects, body motion, camera views, tactile, and audio aligned to one timeline, turning a room-scale scene into supervision a model can learn from.": "ACE 将物体与人体运动、相机视角、触觉、音频等匹配到同一时间线上，把房间尺度的场景转化为模型可以学习的监督信号。",
    "ACE-Data-0 teaser: synchronized multi-modal capture for embodied AI.": "ACE-Data-0 预览：面向具身智能的同步多模态采集。",
    "Data capturing and processing": "数据采集与处理",
    "Motivation": "动机",
    "Embodied AI needs physical-world data from household activities, yet such data is hard to capture.": "具身智能需要来自家庭活动的物理世界数据，然而这类数据很难采集。",
    "Household activity is not a single atomic motion but a long-horizon process. Making tea, wiping a counter, or carrying a cup across the room weaves together perception, locomotion, dexterous manipulation, contact, and sound, and it unfolds over minutes rather than seconds. What embodied AI needs from a dataset is exactly this: how all of these signals evolve together over the course of a real activity, not just how one action looks in a cropped clip.": "家庭活动不是单个原子动作，而是一个跨越很长时间与空间距离的过程。比如说泡茶、擦拭台面，或把杯子带到房间另一侧，会把感知、移动、灵巧操作、接触和声音交织在一起，并在数分钟而非数秒内展开。具身智能真正需要的数据正是这些信号在真实活动过程中如何共同演化，而不只是一个裁剪片段里的单个动作过程。",
    "Unfortunately, existing datasets capture only pieces of this picture. Egocentric video shows natural behavior and the person's observing perspective, but rarely carries measured body, hand, and object states. Motion-capture datasets provide precise trajectories, but mostly in sparse laboratory spaces stripped of the clutter, furniture, and occlusion that make homes difficult, and seldom with the first-person view, audio, or touch. And most interaction clips last seconds, while household goals take minutes and depend on what has already changed in the scene. The gap is not just scale but also completeness: few records contain ego view, exo views, body motion, hand articulation, object state, audio, and touch for the same event, in a real home, over a full activity.": "遗憾的是，现有数据集往往只覆盖这幅图景的某一部分。第一视角视频能够展示人的自然行为与其观察视角，却很少同时包含同一时段下的人体、手部和物体状态。动捕数据集可以提供精确轨迹，但大多来自稀疏实验室环境，缺少真实家庭中的杂物、家具和遮挡，也很少同时包含第一视角、音频或触觉。并且，多数交互片段只有几秒，而家庭内的日常活动通常需要数分钟，并依赖场景中已经发生的变化。这样一来，现有数据集所存在的差距就不只是规模，更是完整性：很少有数据集能在真实家庭、完整活动中，为同一事件同时提供第一人称视角、第三人称视角、人体运动、手部关节、物体状态、音频和触觉。",
    "ACE addresses this gap through": "ACE 通过",
    "long-horizon, goal-directed activity capture": "长时程、目标驱动的活动采集",
    ". Instead of prescribing a sequence of atomic actions, ACE specifies a household goal and allows participants to complete it naturally. A single activity can span multiple rooms, involve many objects, and contain a sequence of interdependent interactions. The resulting recordings preserve the full evolution of the activity—from the initial scene state to intermediate changes and final task completion—providing continuous trajectories for studying long-term planning, state tracking, and memory in embodied agents.": "来弥合这一差距。ACE 并不给参与者规定他们需要进行的动作，而是给定家庭目标，让参与者根据他们自身的习惯自然完成这一任务。并且，参与者的每一次活动都可以跨越多个房间、涉及多个物体，并包含一系列相互依赖的交互。最终记录保留了活动从初始场景状态、中间变化到最终完成的完整演化，为研究具身智能体的长期规划、状态跟踪和记忆提供连续轨迹。",
    "Equally important, ACE provides": "同样重要的是，ACE 提供",
    "continuous synchronization across all modalities": "跨模态的连续同步",
    ". Ego and exo video, body and hand motion, object states, audio, and tactile signals are synchronized to a shared temporal reference, so observations from different sensors correspond to the same physical event at any point in a take. This is particularly critical for long-horizon activities, where clock drift can accumulate over time and progressively break cross-modal correspondence. By maintaining alignment throughout the complete recording, ACE enables precise associations between visual observations, physical motion, contact, and sound over extended activities.": "。第一/第三人称的视频、人体和手部运动、物体状态、音频与触觉信号都同步到同一个时间线上，因此不同模态的数据在任意时刻都对应同一物理事件。这对于长时程活动尤其关键，因为时钟漂移会随时间累积并逐渐破坏跨模态对应关系。ACE则在完整记录中保持对齐，使视觉观察、物理运动、接触和声音能在长活动中被精确关联。",
    "Holistic sensory record": "完整感知记录",
    "Ego-view, exo-view, motion, object state, audio, and touch describe the same event.": "第一人称、第三人称、运动、物体状态、音频和触觉共同描述同一事件。",
    "Long-horizon activity": "长时程活动",
    "Goal-directed household tasks preserve planning, memory, and scene-state evolution.": "目标驱动的家庭活动可以有效地保留规划、记忆和场景状态演化。",
    "Multi-modal synchronization": "多模态同步",
    "All streams stay aligned so visual, physical, contact, and audio signals remain paired.": "所有数据模态保持对齐，使视觉、物理、接触和音频信号始终成对出现。",
    "Two complementary capture scales, one shared pipeline.": "两个互补的采集尺度，一条共享的数据流水线。",
    "Household interaction spans very different spatial scales. Fingertip manipulation requires close-range observations of subtle grasp changes and contact, while household activities require room-scale coverage of locomotion, human-scene interaction, and long motion trajectories. A single capture setup cannot resolve both equally well. ACE is therefore built at two complementary spatial scales: one for fine-grained manipulation and another for room-scale embodied activity.": "家庭交互横跨非常不同的空间范围。指尖操作需要近距离观察细微抓取变化和接触，而家庭活动需要房间尺度覆盖移动、人与场景交互以及长轨迹运动。单一采集设置很难同时兼顾两者，因此 ACE 在两个互补的空间尺度上构建：一个面向细粒度操作，另一个面向房间尺度的具身活动。",
    "The table-scale configuration concentrates close-range cameras, optical motion capture, and tactile gloves around table-top workspaces. At this scale, small spatial changes matter: fingers make and break contact, grasps transition, tools move through short trajectories, and objects rotate within the hand. The close-range setup is designed to preserve these details, providing high-resolution observations of hand-object interaction and tactile events.": "桌面尺度配置将近距离相机、光学动捕和触觉手套集中布置在桌面工作区周围。在这个尺度上，细小空间变化非常重要：手指建立和解除接触，抓取姿态发生转换，工具沿短轨迹移动，物体在手中旋转等。近距离设置旨在保留这些细节，提供手-物交互和触觉事件的高分辨率观察。",
    "The room-scale configuration, on the other hand, extends capture across a furnished apartment with a kitchen, dining area, living room, and bedroom. Wide-baseline cameras and motion-capture coverage follow participants as they walk between spaces, interact with furniture, and manipulate objects over extended trajectories. Rather than constraining activities to avoid occlusion, the setup observes them from multiple viewpoints, preserving natural room-scale behavior in a complex home environment.": "另一方面，房间尺度配置将采集扩展到包含厨房、餐区、客厅和卧室的带家具公寓。宽基线相机与动捕覆盖能够跟随参与者在空间之间移动、与家具交互，并沿长轨迹操作物体。该设置并不通过约束活动来避免遮挡，而是从多个视角观察它们，在复杂家庭环境中保留自然的房间内的人体活动。",
    "Despite their different spatial scales and sites, the two configurations share the same sensing and representation framework. Participants use the same egocentric, body-motion, hand-tracking, and tactile systems, while the environments provide synchronized exocentric video, human and object tracking, and multi-channel audio. A shared calibration, synchronization, and data representation pipeline then standardizes recordings across the two sites, allowing fine-grained manipulation and room-scale household activity to be studied within a unified dataset.": "尽管空间尺度和场地不同，两个配置共享同一套感知与表示框架。参与者使用相同的第一视角、身体运动、手部追踪和触觉系统，环境则提供同步的第三视角视频、人和物体追踪以及多通道音频。共享的标定、同步和数据表示流水线将两个场地的记录标准化，使细粒度操作和房间尺度家庭活动能够在统一数据集中被研究。",
    "Table-scale configuration": "桌面尺度配置",
    "It concentrates dense sensing around table-top workspaces for close-range hand-object manipulation.": "它在桌面工作区周围集中布置密集感知，用于近距离手-物操作。",
    "Room-scale configuration": "房间尺度配置",
    "It spans a furnished apartment for room-scale motion, both human-object and human-scene interaction, and wide-baseline capture.": "它覆盖带家具公寓，用于房间尺度运动、人-物与人-场景交互，以及宽基线采集。",
    "Wearable sensing": "可穿戴感知",
    "Four ego cameras, IMU, headset pose, full-body motion, and hand articulation.": "四个第一人称视角相机、IMU、头显位姿、全身运动和手部关节。",
    "Scene sensing": "场景感知",
    "Stable third-person views plus optical tracking for bodies, objects, and the room.": "稳定第三人称视角加光学追踪，用于人体、物体和房间。",
    "Physical sensing": "物理感知",
    "Tactile pressure and multi-channel audio capture contact beyond what video can show.": "触觉压力和多通道音频捕捉视频无法直接呈现的接触信息。",
    "Each recording is aligned in time and space, not a folder of unrelated streams.": "每段记录都在时间与空间上对齐，而不是一组互不相关的数据流。",
    "Multi-modal capture is only useful when different signals agree on": "只有当不同信号在事件发生的",
    "when": "时间",
    "where": "位置",
    "an event happens. ACE therefore aligns every sensor stream along two dimensions: synchronization establishes a common timeline, while calibration establishes a common spatial frame. Together, they make video, human and object motion, audio, and tactile signals directly comparable across modalities.": "上达成一致时，多模态采集才真正有用。因此 ACE 从两个维度对齐所有传感器流：同步建立共同时间线，标定建立共同空间坐标系。二者共同使视频、人和物体运动、音频与触觉信号能够跨模态直接比较。",
    "Temporal synchronization": "时间同步",
    "resolves the mismatch between sensors that operate at different rates and on different clocks. Without precise alignment, contact may appear in the video before it is registered by the tactile glove, or an object trajectory may lag behind the observed motion. ACE uses the motion-capture clock as a common temporal reference and associates each camera frame with the corresponding motion, audio, and tactile measurements.": "解决不同采样率、不同系统时钟之间的不匹配。若没有精确对齐，接触可能先出现在视频中，随后才被触觉手套记录；物体轨迹也可能滞后于观察到的运动。ACE 使用动捕时钟作为共同时间参考，将每一帧相机图像与对应的运动、音频和触觉测量关联起来。",
    "Spatial calibration": "空间标定",
    "resolves the mismatch between sensors that observe the environment from different viewpoints and coordinate systems. Static exocentric cameras, the moving egocentric headset, tracked objects, and human motion are registered into a common world frame, so measurements from different modalities refer to the same physical location. Together, temporal synchronization and spatial calibration turn independent sensor streams into a coherent multi-modal activity record.": "解决不同视角、不同坐标系观察环境所带来的不匹配。静态第三人称视角相机、移动的第一人称视角头显、被追踪物体和人体运动都被注册到同一世界坐标系中，因此不同模态的测量指向同一物理位置。时间同步与空间标定共同将独立传感器流转化为一致的多模态活动记录。",
    "The pipeline proceeds from multi-modal capture to temporal alignment, then to annotations derived from measured states.": "数据流水线从多模态采集开始，经过时间对齐，再生成来自测量的数据标注。",
    "Temporal alignment": "时间对齐",
    "A common clock aligns camera frames, motion states, tactile readings, and audio.": "共同时间钟对齐相机帧、运动状态、触觉读数和音频。",
    "Spatial registration": "空间标定",
    "Cameras, the ego rig, tracked objects, and human motion share one world frame.": "动捕相机、第一人称相机设备、被追踪物体和人体运动共享同一世界坐标系。",
    "Capture protocol": "采集准则",
    "Each take follows setup, goal briefing, recording, and post-capture quality checks.": "每次采集都遵循设备准备、目标说明、录制和采后质量检查流程。",
    "Modality stack": "模态栈",
    "Five modalities travel together through every aligned take.": "五种不同的模态在每个对齐的采集数据中同步发生。",
    "Once synchronized and calibrated, each ACE sequence becomes a paired record of observations, measured states, and physical interaction signals. Raw sensor streams are released together with camera intrinsics and poses, headset motion, object trajectories, scene geometry, and the shared timeline and coordinate frame that connect them.": "一旦完成同步和标定，每个 ACE 序列就成为观察、测量状态和物理交互信号的成对记录。原始传感器流与相机内参和位姿、头显运动、物体轨迹、场景几何，以及连接它们的共享时间线和坐标系一起发布。",
    "This is what separates an embodied AI dataset from a collection of videos. Different modalities describe the same physical state at the same time and place, allowing one signal to serve as input and another as supervision without uncertain cross-modal correspondence.": "这正是具身智能数据集区别于视频集合的地方。不同模态在同一时间、同一位置描述同一物理状态，使一个信号可以作为输入，另一个信号可以作为监督，而不必依赖不确定的跨模态对应关系。",
    "Egocentric view": "第一视角",
    "4 fisheye cameras, IMU, and tracked headset pose.": "4 个鱼眼相机、IMU 和被追踪的头显位姿。",
    "First-person observations for robot-like perception.": "面向类机器人感知的第一人称观察。",
    "Exocentric view": "第三视角",
    "Multi-view RGB from table-scale and room-scale cameras.": "来自桌面尺度和房间尺度相机的多视角 RGB。",
    "Body, scene, and object context beyond the actor's view.": "超出操作者视野的人体、场景和物体上下文。",
    "Motion state": "运动状态",
    "Full-body, hand, and object trajectories at 60 Hz.": "60 Hz 的全身、手部和物体轨迹。",
    "Metric supervision for state recovery and imitation.": "用于状态恢复和模仿学习的度量监督。",
    "Tactile signal": "触觉信号",
    "Full-palm tactile gloves recording pressure maps across the palm and fingers.": "记录手掌与手指压力分布的全掌触觉手套。",
    "Contact evidence for grasp, slip, and pressure.": "抓取、滑动和压力的接触证据。",
    "Multi-channel audio": "多通道音频",
    "Audio recorded by the egocentric headset and exocentric cameras.": "由第一人称头显和第三人称视角相机录制的音频。",
    "Contact sounds, appliance events, and ambient cues.": "接触声音、设备事件和环境线索。",
    "Video data examples": "视频数据示例",
    "Synchronized examples with ego-view and exo-view cameras.": "包含第一人称视角相机 和 多个第三人称视角相机的同步示例。",
    "Each ACE sequence provides paired": "每个 ACE 序列都提供成对的",
    "egocentric and exocentric observations": "第一人称视角和第三人称视角观察",
    "of the same activity. The ego view captures the scene from the participant's perspective, while the multi-view exocentric cameras preserve the surrounding environment, full-body motion, and object context.": "来记录同一活动。第一人称视角从参与者视角捕捉场景，多个第三人称视角相机则保留周围环境、全身运动和物体上下文。",
    "The two views are complementary. Egocentric video resolves hands and nearby interactions but moves with the head and often loses whole-body context. Exocentric views provide stable scene-level observations, yet fine hand motion may be small or occluded. By recording both simultaneously, ACE preserves these complementary perspectives for the same physical events.": "这两种视角互为补充。第一人称视角视频能清楚呈现手部和近距离交互，但会随头部移动，并常常丢失全身上下文。第三人称视角提供稳定的场景级观察，但细微手部动作可能很小或被遮挡。ACE 同时记录两者，为同一物理事件保留互补视角。",
    "Example 01": "示例 01",
    "Example 02": "示例 02",
    "Example 03": "示例 03",
    "Example 04": "示例 04",
    "Example 05": "示例 05",
    "Table-scale long-horizon sequence": "桌面尺度长时程序列",
    "Room-scale long-horizon sequence": "房间尺度长时程序列",
    "HSI sequence": "人-场景交互序列",
    "Table-scale household interaction": "桌面尺度家庭交互",
    "One of ego views": "其中一个第一视角",
    "Exo 1": "第三视角 1",
    "Exo 2": "第三视角 2",
    "Exo 3": "第三视角 3",
    "Exo 4": "第三视角 4",
    "Exo 5": "第三视角 5",
    "Exo 6": "第三视角 6",
    "Exo 7": "第三视角 7",
    "Exo 8": "第三视角 8",
    "The same moment is observed from the actor's perspective and from eight synchronized exocentric cameras.": "同一时刻可以从操作者视角以及 8 个同步的第三人称视角相机中被观察。",
    "A clip from a 20-minute room-scale long-horizon sequence with synchronized multi-modal data.": "一段 20 分钟房间尺度长时程序列中的片段，包含同步的多模态数据。",
    "A synchronized HSI sequence follows the participant's ego view and eight fixed exocentric cameras across the full interaction.": "一个同步的人-场景交互序列通过参与者的第一人称视角和 8 个固定的第三人称视角相机，完整记录整个交互过程。",
    "Data annotation": "数据标注",
    "Annotation examples from the released annotation stack.": "标注示例。",
    "Each ACE sequence is paired with annotations of both": "每个 ACE 序列都配有",
    "human state and object state": "人体状态和物体状态",
    ". Human and hand motion describe how the participant moves and interacts, while object trajectories and 6-DoF poses capture how the surrounding scene changes over time.": "标注。人体和手部运动描述参与者如何移动和交互，物体轨迹与 6-DoF 位姿则捕捉周围场景如何随时间变化。",
    "Most annotations are derived directly from the calibrated capture system. Body and hand poses, object poses, bounding boxes, motion trails, contact states, and camera projections are generated from measured physical states rather than inferred independently from video. Text descriptions are then aligned to the same timeline, adding a semantic layer to the multi-modal record.": "大多数标注直接来自标定后的采集系统。身体和手部姿态、物体位姿、边界框、运动轨迹、接触状态和相机投影都由测量到的物理状态生成，而不是独立从视频中推断。文本描述随后被对齐到同一时间线，为多模态记录加入语义层。",
    "Together, the annotation stack describes": "整体来看，标注栈描述了",
    "where the sensors are, where the human and objects are, when physical interaction occurs, and what activity is unfolding": "传感器在哪里，人和物体在哪里，物理交互何时发生，以及正在进行什么活动",
    ". Because the underlying geometry and motion come from calibrated measurements, these annotations remain consistent even under occlusion, clutter, and visual ambiguity.": "。由于底层几何和运动来自标定测量，即使存在遮挡、杂乱和视觉歧义，这些标注也能保持一致。",
    "Camera and timeline": "相机与时间线",
    "Intrinsics, poses, headset trajectory, and synchronized timestamps.": "内参、位姿、头显轨迹和同步时间戳。",
    "Human and hand state": "人体与手部状态",
    "Body skeletons, hand articulation, and per-view projected overlays.": "身体骨架、手部关节和每个视角的投影叠加。",
    "Object state": "物体状态",
    "Object meshes, 6-DoF poses, boxes, and motion trails.": "物体 mesh、6-DoF 位姿、框和运动轨迹。",
    "Contact labels": "接触标签",
    "Hand-object contact grounded by motion state and tactile readings.": "由运动状态和触觉读数支撑的手-物接触。",
    "Language descriptions": "语言描述",
    "Natural-language descriptions aligned with the physical timeline.": "与物理时间线对齐的自然语言描述。",
    "Human pose": "人体姿态",
    "Hand motion": "手部运动",
    "Tactile": "触觉",
    "Language caption": "语言描述",
    "Annotation example 01": "标注示例 01",
    "Annotation example 04": "标注示例 04",
    "Annotation example 05": "标注示例 05",
    "Human pose annotation": "人体姿态标注",
    "Object state annotation": "物体状态标注",
    "Human pose annotations align egocentric video, multi-view exocentric video, and 3D human motion across the same activity timeline.": "人体姿态标注将第一人称视角视频、多视角第三人称视角视频和 3D 人体运动对齐到同一活动时间线。",
    "Object annotations expose object identity, spatial state, 6-DoF pose, and motion trails in the same synchronized timeline.": "物体标注在同一同步时间线中展示物体身份、空间状态、6-DoF 位姿和运动轨迹。",
    "Hand motion annotation": "手部运动标注",
    "Hand annotations focus on hand state, articulation, and projected overlays aligned with the physical activity record.": "手部标注关注与物理活动记录对齐的手部状态、关节运动和投影叠加。",
    "Tactile annotation": "触觉标注",
    "Tactile annotations add contact evidence that complements visual observations and motion state during manipulation.": "触觉标注补充操作过程中的接触证据，与视觉观察和运动状态互为补充。",
    "Language caption annotation": "语言描述标注",
    "Language annotations provide semantic descriptions aligned to the same measured activity timeline.": "语言标注提供与同一测量活动时间线对齐的语义描述。",
    "Annotation example 02": "标注示例 02",
    "Annotation example 03": "标注示例 03",
    "ACE-Data-0: A Large-Scale Multi-Modal Dataset for Long-Horizon Household Activities": "ACE-Data-0：面向长时程家庭活动的大规模多模态数据集",
    "Through ACE, we release": "通过 ACE，我们发布",
    ", a large-scale dataset of long-horizon, multi-modal, synchronized household activities. It comprises 150 hours of capture across 200 task categories, recorded as continuous takes that each run for minutes and are segmented into 75,000 interaction episodes. Each take preserves egocentric and exocentric observations together with human motion, hand motion, object states, audio, and tactile signals over a continuous activity timeline.": "，一个大规模、长时程、多模态、同步的家庭活动数据集。它包含 150 小时的采集，覆盖 200 个任务类别；每个 take 都是持续数分钟的连续记录，并被切分为总计 75,000 个交互片段（episode）。每个 take 都在连续活动时间线上保留第一人称视角和第三人称视角观察，以及人体运动、手部运动、物体状态、音频和触觉信号。",
    "ACE-Data-0 is built around": "ACE-Data-0 围绕",
    "goal-directed activity capture": "目标驱动活动采集",
    ". Participants receive household goals rather than fixed sequences of atomic actions and complete them naturally. The same goal may involve different routes, object choices, grasp strategies, subtask orders, pauses, and recoveries. Long-horizon behavior therefore emerges from task completion itself, rather than from manually concatenating short interaction clips.": "构建。参与者收到的是家庭目标，而不是固定的原子动作序列，并以自然方式完成它们。同一目标可能涉及不同路线、物体选择、抓取策略、子任务顺序、停顿和恢复。因此，长时程行为来自任务完成本身，而不是人工拼接短交互片段。",
    "The dataset contains three complementary types of activities.": "数据集包含三类互补活动。",
    "Atomic HOI": "Atomic HOI",
    "takes focus on one to three everyday manipulation tasks each, such as pouring water, preparing food, or tidying a surface.": "每次采集关注 1 到 3 个日常操作任务，例如倒水、准备食物或整理台面。",
    "Chain of HOI": "Chain of HOI",
    "connects multiple atomic HOI into continuous household routines that require intermediate state tracking and sequential decision making.": "将多个 Atomic HOI 连接成连续家庭流程，需要中间状态跟踪和序列决策。",
    "HSI": "HSI",
    "emphasizes whole-body interaction with the environment, including locomotion, posture transitions, and contact with furniture and other scene structures.": "强调全身与环境的交互，包括移动、姿态转换，以及与家具和其他场景结构的接触。",
    "Across all three categories, ACE-Data-0 preserves the complete context of an interaction. The approach, contact, manipulation, release, and resulting scene changes remain within the same timeline, so an action can be interpreted through both its history and its consequence. A grasp is not recorded as an isolated label, but as one event within an evolving human-object-scene state.": "在这三类活动中，ACE-Data-0 都保留交互的完整上下文。接近、接触、操作、释放以及由此产生的场景变化都位于同一时间线上，因此一个动作可以通过其历史和结果来理解。一次抓取不是孤立标签，而是不断演化的人-物-场景状态中的一个事件。",
    "Natural variation is retained as part of the data rather than removed through scripting. Participants choose how to move, where to stand, which objects to use, and how to recover from small deviations, while object configurations and trajectories change with the activity. Because this variation remains grounded by synchronized, measured multi-modal signals, ACE-Data-0 captures household interaction as performed behavior rather than a collection of isolated action instances.": "自然变化被保留为数据的一部分，而不是通过脚本消除。参与者自行选择如何移动、站在哪里、使用哪些物体以及如何从小偏差中恢复，物体配置和轨迹也会随活动变化。由于这些变化仍由同步、测量得到的多模态信号支撑，ACE-Data-0 将家庭交互捕捉为真实执行行为，而不是孤立动作实例的集合。",
    "Ambient-captured interaction dataset": "环境式采集交互数据集",
    "Single-task, long-horizon interactions": "单任务、长时程交互",
    "Drink": "饮用",
    "Pour": "倒入",
    "Fill": "装满",
    "Wash": "清洗",
    "Clean": "清洁",
    "Wipe": "擦拭",
    "Peel": "去皮",
    "Chop": "切碎",
    "Cook": "烹饪",
    "Stir": "搅拌",
    "Fold": "折叠",
    "Tidy": "整理",
    "Composed household activities": "组合式家庭活动",
    "Open fridge": "打开冰箱",
    "Take out": "取出",
    "Rinse": "冲洗",
    "Cut": "切割",
    "Season": "调味",
    "Serve": "上菜",
    "Clean up": "清理",
    "Human-scene interactions": "人-场景交互",
    "Sit": "坐下",
    "Navigate": "移动",
    "Exercise": "锻炼",
    "Open/close": "打开/关闭",
    "Pick/place": "拿取/放置",
    "Posture transition": "姿态转换",
    "Scene arrangement": "场景整理",
    "Task examples: ACE-Data-0 spans atomic HOI, human-scene interaction, and chain of HOI household activities.": "任务示例：ACE-Data-0 覆盖 Atomic HOI、人-场景交互和 Chain of HOI 家庭活动。",
    "Overview statistics of ACE-Data-0: task categories, task composition, and the distribution of frames across atomic HOI, chain of HOI, and HSI takes.": "ACE-Data-0 总体统计：任务类别、任务组合，以及 Atomic HOI、Chain of HOI 和 HSI 采集在帧数上的分布。",
    "Why this division matters": "为什么这样的划分重要",
    "Short clips (atomic HOI) can show what a grasp looks like; long takes (chain of HOI) reveal what the grasp is for. A household routine has memory: objects move, surfaces become clean or cluttered, a person leaves one area and returns later, and the scene state changes over time. ACE keeps that continuity so perception, action, and physical state can be studied together.": "短片段（Atomic HOI）可以展示抓取看起来是什么样；长片段（Chain of HOI）则揭示抓取是为了什么。家庭流程具有记忆：物体会移动，表面会变干净或杂乱，人会离开某一区域再返回，场景状态会随时间变化。ACE 保留了这种连续性，使感知、动作和物理状态能够被共同研究。",
    "This also makes the data harder in the right way. Errors can accumulate over time, objects can leave and re-enter view, and the meaning of an action can depend on what happened much earlier in the take. Those are exactly the conditions an embodied agent will face in a real home.": "这也让数据以正确的方式采集和处理变得更难。错误会随时间累积，物体可能离开视野再重新出现，一个动作的含义也可能依赖于采集早先发生的事情。这正是具身智能体在真实家庭中会遇到的问题。",
    "A diagnostic benchmark from signals, to states, to interactions.": "一个从信号、状态到交互的评测基准。",
    "ACE-Data-0 organizes evaluation as a": "ACE-Data-0 将评测划分为一个",
    "bottom-up hierarchy of embodied perception": "自底向上的具身感知层级",
    ". Rather than measuring task success, the benchmark isolates where an interaction understanding pipeline fails—from recovering low-level physical signals, to reconstructing scene components, to understanding coordinated human-object interaction.": "。它不是只衡量任务成功率，而是定位交互理解流水线在哪里失败：从底层物理信号恢复，到场景组件重建，再到协调的人-物交互理解。",
    "The first level focuses on": "第一层关注",
    "low-level signals": "底层信号",
    ", asking whether physical cues such as tactile contact can be recovered from visual observations. The second level evaluates": "，考察是否能从视觉观察中恢复触觉接触等物理线索。第二层评估",
    "scene components": "场景组件",
    ", including metric 3D recovery of human body motion in furnished home scenes. The third level moves to": "，包括在带家具的家庭场景中对人体运动进行度量 3D 恢复。第三层进入",
    "interaction understanding": "交互理解",
    ", where models must recover and reason about coordinated hand, body, and object motion from egocentric and exocentric observations.": "，模型需要从第一视角和第三视角观察中恢复并推理协调的手、身体和物体运动。",
    "This hierarchy is diagnostic by design. Errors at one level expose the missing information required by the next: missed contact obscures whether a grasp succeeds, inaccurate human or object states break physical correspondence, and lost world-space motion limits interaction understanding and downstream embodied planning. The benchmark therefore reveals": "这个层级本身就是诊断式设计。某一层的错误会暴露下一层所需信息的缺失：遗漏接触会模糊抓取是否成功，不准确的人体或物体状态会破坏物理对应关系，丢失世界坐标中的运动会限制交互理解和下游具身规划。因此，该基准揭示",
    "what a model understands, and where that understanding breaks down": "模型理解了什么，以及这种理解在哪里失效",
    ".": "。",
    "All three levels are evaluated on the same synchronized and calibrated recordings. Low-level signals, 3D scene states, and interaction trajectories refer to the same physical events in a shared timeline and coordinate frame. This allows predictions across different tasks and viewpoints to be compared against consistent measured ground truth, turning ACE-Data-0 into a unified benchmark for embodied perception and interaction understanding.": "三个层级都在同一套同步且标定的记录上评测。底层信号、3D 场景状态和交互轨迹都指向共享时间线和坐标系中的同一物理事件。这使不同任务和视角的预测可以与一致的测量真值比较，并将 ACE-Data-0 转化为具身感知与交互理解的统一基准。",
    "The three levels mirror the capabilities an embodied agent must chain together: perceive, understand, and interact.": "这三个层级对应具身智能体必须串联起来的能力：感知、理解和交互。",
    "Low-level signal inference": "底层信号推断",
    "Predict tactile readings and contact events from video, especially when the hand hides the fingertips at the exact moment contact matters. ACE pairs the video frame with the tactile reading from the same instant, so the target is measured contact rather than a visual proxy.": "从视频预测触觉读数和接触事件，尤其是在接触最关键的瞬间手部遮挡指尖时。ACE 将视频帧与同一时刻的触觉读数配对，因此目标是测量到的接触，而不是视觉代理信号。",
    "This track asks whether an abundant sensor, vision, can recover a scarce one, touch. It measures not only when contact occurs, but also where pressure is distributed across the hand, which is where current methods remain weakest.": "该任务考察丰富传感器（视觉）是否能恢复稀缺传感器（触觉）。它不仅衡量接触何时发生，还衡量压力在手上如何分布，而后者正是现有方法最薄弱的环节。",
    "Scene component recovery": "场景组件恢复",
    "Estimate human body motion under furniture occlusion, uncommon household postures, and minutes of continuous movement. The same measured ground truth compares per-frame, temporal, and scene-aware methods across single-view, multi-view, and egocentric inputs.": "在家具遮挡、非常见的家庭姿态和数分钟连续运动下估计人体运动。同一套测量真值可以在单视角、多视角和第一视角输入下比较逐帧、时序和场景感知方法。",
    "Embodied interaction understanding": "具身交互理解",
    "Recover hand motion and trajectories from egocentric and exocentric views, then compare complementary failure modes under the same ground truth. Ego-view captures close hand detail; exo-view preserves global trajectory and scene context.": "从第一人称视角和第三人称视角恢复手部运动与轨迹，并在同一真值下比较互补的失败模式。Ego-view 捕捉近距离手部细节；exo-view 保留全局轨迹和场景上下文。",
    "This level connects perception to action. The hand approach, grasp, adjustment, and release are the parts of human behavior most directly useful for robot imitation and manipulation policy learning.": "这一层将感知连接到行动。手部接近、抓取、调整和释放，是人类行为中最直接有助于机器人模仿和操作策略学习的部分。",
    "ACE-Data-0 is an invitation to ambient capture for embodied AI.": "ACE-Data-0 是对面向具身智能的环境式采集的一次邀请。",
    "Embodied intelligence requires understanding activities as evolving physical processes, not isolated visual clips. An agent must track how the scene changes, how humans and objects move and interact, when contact occurs, and how individual actions contribute to a longer-term goal. ACE captures these signals as a unified, synchronized, and calibrated record of household activity.": "具身智能需要把活动理解为不断演化的物理过程，而不是孤立的视频片段。智能体必须跟踪场景如何变化，人和物体如何移动与交互，接触何时发生，以及单个动作如何服务于更长期目标。ACE 将这些信号采集为统一、同步且标定的家庭活动记录。",
    "Together, the two ACE configurations span fine-grained manipulation and room-scale embodied behavior. Their complementary capture scales connect detailed hand-object interaction with whole-body motion, scene contact, and long activity trajectories, providing measured observations of household interaction across different spatial and temporal scales.": "综合来看，ACE 的两个配置覆盖细粒度操作和房间尺度具身行为。它们互补的采集尺度将细致手-物交互与全身运动、场景接触和长活动轨迹连接起来，为不同空间和时间尺度上的家庭交互提供测量观察。",
    "With": "借助",
    ", we aim to support embodied models that reason over long horizons, recover physical states, learn from natural behavioral variation, and connect perception to interaction in everyday environments. The accompanying benchmark evaluates these capabilities from low-level signal recovery to scene reconstruction and interaction understanding, treating them as connected stages of embodied perception rather than isolated tasks.": "，我们希望支持能够进行长时程推理、恢复物理状态、从自然行为变化中学习，并在日常环境中连接感知与交互的具身模型。配套基准从底层信号恢复到场景重建和交互理解评估这些能力，将它们视为具身感知的连续阶段，而不是孤立任务。",
    "Ambient capture turns everyday activity into aligned supervision.": "环境式采集将日常活动转化为对齐的监督信号。",
    "ACE provides the capture system, the synchronized dataset, and the diagnostic benchmark needed to study household interaction as a continuous embodied experience.": "ACE 提供研究家庭交互所需的采集系统、同步数据集和诊断式基准，使得具身模型能够学习到家庭活动的经验。",
    "ACE-Data-0 research blog": "ACE-Data-0 研究博客",
    "Back to top ↑": "回到顶部 ↑",
    "All participants volunteered and signed informed consent forms covering the recording and public release of the data.": "所有参与者均为自愿参加，并签署了涵盖数据采集与公开发布的知情同意书。"
  }
}
