REAL-I Industrial Humanoid Dataset

The REAL-I Industrial Humanoid Dataset, maintained by Leju Robot in collaboration with Ant Group and AliCloud, is a large-scale real-robot dataset specifically designed for industrial humanoid manipulation. Released in conjunction with the ICRA 2026 REAL-I Embodied Intelligence Challenge, it features over 180,000 high-quality demonstrations across a variety of complex industrial tasks, including metal parts flipping, parcel weighing, and chemical bottle alignment. The dataset utilizes the Kuavo 4 Pro and Kuavo 5 humanoid platforms, capturing synchronized multi-view RGB-D video, proprioception, and action trajectories at frequencies of 15 to 30 Hz. It forms a key component of the RoboCOIN and LingBot-VLA ecosystems, providing a standardized, open-source resource for researchers developing Vision-Language-Action (VLA) models. The data is publicly accessible via GitHub and Hugging Face under the Apache-2.0 license, aimed at advancing cross-embodiment generalization and temporal reasoning in real-world industrial environments for a global research community.

Dataset specifications
Year2026
Episodes180,000
Total hours60,000
Trajectories180,000
Frame rate15 fps
EmbodimentsKuavo 4 Pro, Kuavo 5, Kuavo
Task categoriesmanipulation, pick-and-place, warehouse, inspection, other
LicenseApache-2.0
Accessopen — commercial use permitted
MaintainerLeju Robot
Origin countryCN

Methodology

The REAL-I Industrial Humanoid Dataset was constructed using a large-scale data acquisition pipeline involving real-world teleoperation of Leju Kuavo humanoid robots. The collection process focused on 1:1 replications of actual industrial scenes to minimize the sim-to-real gap. The dataset employs the CoRobot processing framework, which utilizes Robot Trajectory Markup Language (RTML) for automated quality assessment and hierarchical annotation. Annotations are provided at the trajectory level (global objectives), segment level (subtask decomposition), and frame level (kinematic states and dense action labels), enabling multi-resolution learning.

Collection

Data collection was conducted primarily at Leju's dedicated humanoid skill center in Beijing. The repository includes 50,000 hours of robot interaction data augmented by 10,000 hours of egocentric human manipulation video for predictive dynamics distillation. The heterogeneous data is aligned into a unified 55-dimensional action space covering whole-body control, including head, waist, dual arms, and dexterous hands. Multi-view visual observations are captured using head-mounted and wrist-mounted RGB-D cameras to provide robust spatial reasoning capabilities.

Use-Cases

This dataset is optimized for training and benchmarking Vision-Language-Action (VLA) foundation models like LingBot-VLA. It is particularly suited for researchers focusing on bimanual coordination, long-horizon industrial workflows, and cross-embodiment generalization. Key industrial use cases include automotive assembly assistance, logistics sorting, and quality control in pharmaceutical or manufacturing environments. The structured nature of the data supports various learning paradigms, including imitation learning, offline reinforcement learning, and flow-matching-based action generation.

Frequently asked questions

What robotic platforms were used for the REAL-I dataset collection?

The dataset primarily utilizes the Kuavo 4 Pro and Kuavo 5 humanoid robots developed by Leju Robot, featuring full-size bipedal or wheeled configurations with dual 7-DoF arms.

What are the primary industrial tasks covered in the dataset?

The dataset covers specific industrial challenges such as metal part righting (flipping), express parcel weighing and sorting, and chemical product loading and alignment.

Is the REAL-I Industrial Humanoid Dataset available for commercial use?

Yes, the dataset and its associated models are released under the Apache-2.0 license, which permits both academic and commercial use.

How was the data annotated for model training?

The data features a hierarchical capability pyramid with three levels: trajectory-level for global planning, segment-level for subtask reasoning, and frame-level for precise kinematic control.

Where can I find the baseline models and codebase for this dataset?

The official baseline model is LingBot-VLA, and the dataset is publicly hosted via the Robbyant project on GitHub and Hugging Face as part of the RoboCOIN ecosystem.