Xperience-10M — Large-Scale Egocentric Multimodal Human Experience Dataset
Xperience-10M is a large-scale egocentric multimodal dataset containing 10 million interactions and 10,000 hours of synchronized first-person recordings for embodied AI, robotics, and world modeling research. Released in April 2026 under Apache 2.0 for commercial use, it provides six synchronized video streams, audio, stereo depth, 6-DoF camera pose, hand motion capture, full-body motion capture, IMU, and hierarchical language annotations per recording session. The dataset represents the largest egocentric multimodal collection available for robot learning pretraining, surpassing Ego4D (3,025 hours) by more than 3x in duration and adding sensor modalities beyond video. The Apache 2.0 license makes it commercially usable — a significant advantage over Ego4D's non-commercial restriction.
| Year | 2026 |
|---|---|
| Episodes | 10,000,000 |
| Total hours | 10,000 |
| Embodiments | human (wearable camera), egocentric rig |
| Modalities | rgb, audio, depth, imu, proprioception, language |
| Task categories | manipulation, cleaning, cooking, human-robot-interaction, long-horizon, inspection |
| Data format | mp4, json, hdf5 |
| License | Apache 2.0 |
| Access | gated — commercial use permitted |
| Maintainer | Xperience Research Consortium |
| Origin country | US |
What is it?
Xperience-10M is a large-scale egocentric multimodal dataset containing 10 million interactions and 10,000 hours of synchronized first-person recordings for embodied AI, robotics, and world modeling. Released in April 2026 under Apache 2.0 for commercial use, each recording session provides six synchronized video streams, audio, stereo depth, 6-DoF camera pose, hand motion capture, full-body motion capture, IMU, and hierarchical language annotations. At 10,000 hours it is 3x larger than Ego4D and adds sensor modalities beyond video that Ego4D lacks.
Who is it for?
Researchers pretraining visual representations, world models, and embodied AI systems at scale. Particularly valuable for teams that need commercially licensable egocentric data at a scale beyond what any robot-collected dataset can provide. The full-body mocap and hand mocap modalities enable retargeting to humanoid robot kinematics without requiring direct robot data collection.
Key specifications
- Interactions: 10,000,000
- Duration: 10,000 hours of synchronized recordings
- Video streams: 6 synchronized per session
- Modalities: RGB, audio, stereo depth, 6-DoF camera pose, hand mocap, full-body mocap, IMU, language
- Format: MP4, JSON, HDF5
- License: Apache 2.0 — commercial use permitted
- Access: Open — Hugging Face
How it compares
Xperience-10M supersedes Ego4D (3,025 hours, non-commercial) in both scale and modality coverage. EgoVerse (1,213 hours) provides robot-specific task labels but covers far less total data. Xperience-10M's combination of commercial license, 10K hours, and full-body mocap makes it the strongest foundation pretraining dataset currently available.
Limitations and access notes
Like all human activity datasets, Xperience-10M requires domain adaptation for robot policies — human motion must be retargeted to robot kinematics. Apache 2.0 permits unrestricted commercial use.
Linked professions
- Hotel Housekeeper
- Warehouse Picker Packer
- Fast Food Worker
- Commercial Floor Cleaner
- Assembly Line Worker Repetitive
Frequently asked questions
How does Xperience-10M compare to Ego4D?
Xperience-10M contains 10,000 hours vs Ego4D's 3,025 hours — more than 3x larger. Xperience-10M adds hand mocap, full-body mocap, stereo depth, and 6-DoF camera pose that Ego4D lacks. Crucially, Xperience-10M is Apache 2.0 licensed for commercial use, while Ego4D is non-commercial research only.
Can Xperience-10M be used commercially?
Yes. Xperience-10M is Apache 2.0 licensed, permitting unrestricted commercial use, modification, and redistribution.
What does Xperience-10M's full-body mocap enable?
Full-body motion capture provides skeletal joint angles and positions for the entire human body during each recording. This enables retargeting human demonstrations to humanoid robot kinematics without requiring robot hardware for data collection — the key advantage over pure video datasets.
How do I access Xperience-10M?
Xperience-10M is available on Hugging Face. No registration required. The Apache 2.0 license permits download and commercial use without restrictions.
What are the 6 synchronized video streams?
Each Xperience-10M recording session captures 6 synchronized video streams covering different perspectives — typically including forward-facing, downward-facing, and peripheral views from the head-mounted rig, enabling 3D reconstruction and multi-view learning that single-camera egocentric datasets cannot support.