Xperience-10M — Large-Scale Egocentric Multimodal Human Experience Dataset

Xperience-10M is a large-scale egocentric multimodal dataset containing 10 million interactions and 10,000 hours of synchronized first-person recordings for embodied AI, robotics, and world modeling research. Released in April 2026 under Apache 2.0 for commercial use, it provides six synchronized video streams, audio, stereo depth, 6-DoF camera pose, hand motion capture, full-body motion capture, IMU, and hierarchical language annotations per recording session. The dataset represents the largest egocentric multimodal collection available for robot learning pretraining, surpassing Ego4D (3,025 hours) by more than 3x in duration and adding sensor modalities beyond video. The Apache 2.0 license makes it commercially usable — a significant advantage over Ego4D's non-commercial restriction.

Dataset specifications
Year2026
Episodes10,000,000
Total hours10,000
Embodimentshuman (wearable camera), egocentric rig
Modalitiesrgb, audio, depth, imu, proprioception, language
Task categoriesmanipulation, cleaning, cooking, human-robot-interaction, long-horizon, inspection
Data formatmp4, json, hdf5
LicenseApache 2.0
Accessgated — commercial use permitted
MaintainerXperience Research Consortium
Origin countryUS

What is it?

Xperience-10M is a large-scale egocentric multimodal dataset containing 10 million interactions and 10,000 hours of synchronized first-person recordings for embodied AI, robotics, and world modeling. Released in April 2026 under Apache 2.0 for commercial use, each recording session provides six synchronized video streams, audio, stereo depth, 6-DoF camera pose, hand motion capture, full-body motion capture, IMU, and hierarchical language annotations. At 10,000 hours it is 3x larger than Ego4D and adds sensor modalities beyond video that Ego4D lacks.

Who is it for?

Researchers pretraining visual representations, world models, and embodied AI systems at scale. Particularly valuable for teams that need commercially licensable egocentric data at a scale beyond what any robot-collected dataset can provide. The full-body mocap and hand mocap modalities enable retargeting to humanoid robot kinematics without requiring direct robot data collection.

Key specifications

How it compares

Xperience-10M supersedes Ego4D (3,025 hours, non-commercial) in both scale and modality coverage. EgoVerse (1,213 hours) provides robot-specific task labels but covers far less total data. Xperience-10M's combination of commercial license, 10K hours, and full-body mocap makes it the strongest foundation pretraining dataset currently available.

Limitations and access notes

Like all human activity datasets, Xperience-10M requires domain adaptation for robot policies — human motion must be retargeted to robot kinematics. Apache 2.0 permits unrestricted commercial use.

Linked professions

Frequently asked questions

How does Xperience-10M compare to Ego4D?

Xperience-10M contains 10,000 hours vs Ego4D's 3,025 hours — more than 3x larger. Xperience-10M adds hand mocap, full-body mocap, stereo depth, and 6-DoF camera pose that Ego4D lacks. Crucially, Xperience-10M is Apache 2.0 licensed for commercial use, while Ego4D is non-commercial research only.

Can Xperience-10M be used commercially?

Yes. Xperience-10M is Apache 2.0 licensed, permitting unrestricted commercial use, modification, and redistribution.

What does Xperience-10M's full-body mocap enable?

Full-body motion capture provides skeletal joint angles and positions for the entire human body during each recording. This enables retargeting human demonstrations to humanoid robot kinematics without requiring robot hardware for data collection — the key advantage over pure video datasets.

How do I access Xperience-10M?

Xperience-10M is available on Hugging Face. No registration required. The Apache 2.0 license permits download and commercial use without restrictions.

What are the 6 synchronized video streams?

Each Xperience-10M recording session captures 6 synchronized video streams covering different perspectives — typically including forward-facing, downward-facing, and peripheral views from the head-mounted rig, enabling 3D reconstruction and multi-view learning that single-camera egocentric datasets cannot support.