RoboEdit-14M

RoboEdit-14M is a large-scale, synthetic-to-real robot manipulation dataset developed by researchers at UCLA, MIT, and the University of Utah. It addresses the embodiment gap in robot learning by utilizing the RoboEdit-ADC pipeline to transform human manipulation videos into action-consistent, physically plausible robot videos. The dataset comprises 174,547 aligned human-robot video pairs, totaling over 14 million frames and approximately 131 hours of data. By leveraging existing human-object interaction footage from sources like DexYCB and TACO, the dataset provides scalable supervision across seven different robot embodiments. It is designed for researchers aiming to improve robot policy training by overcoming the high costs associated with collecting embodiment-specific robot data. The dataset is publicly accessible via the project's official website, providing a robust resource for advancing generalizable manipulation skills in diverse environments through cross-embodiment video retargeting and 3D kinematic state recovery.

Dataset specifications
Year2026
Episodes174,547
Total hours131
Trajectories174,547
Frames14,138,307
Frame rate30 fps
Embodimentsvarious
Task categoriesmanipulation
LicensearXiv.org perpetual non-exclusive license
Accessopen
MaintainerYG Yaowei Guo et al.
Origin countryUS

RoboEdit-14M is a comprehensive dataset generated by the RoboEdit-ADC (Automatic Data Collection) pipeline. The methodology involves reconstructing 3D hand-object interactions from RGB human videos and retargeting these motions onto various robot embodiments. The pipeline employs a video editing engine, RoboEdit-Trans, to ensure the resulting robot videos are physically plausible and action-consistent. The dataset is primarily used for training robot manipulation policies by providing large-scale, diverse supervision that is not limited by the constraints of traditional robot teleoperation. It includes data derived from established sources such as DexYCB, GigaHands, H2O, HOT3D, and TACO, ensuring a wide variety of scenes and interaction types. The dataset is intended for academic research in robotics, computer vision, and machine learning, specifically focusing on cross-embodiment transfer and scalable imitation learning.

Frequently asked questions

What is the primary purpose of RoboEdit-14M?

It aims to bridge the embodiment gap by transforming abundant human manipulation videos into scalable, aligned robot training data.

How many frames are included in the dataset?

The dataset contains 14,138,307 frames.

Which institutions are involved in this research?

The research is a collaboration between UCLA, MIT, and the University of Utah.

Is the dataset real or synthetic?

It consists of aligned video pairs that leverage real human video sources to generate robot-specific supervision.

What is the total duration of the dataset?

The dataset provides approximately 130.91 hours of paired video data.