RoboReward — Vision-Language Reward Dataset for Robotics

Stanford University / UC Berkeley · 2026 · CC BY 4.0 · US
Episodes45,000
DurationNot disclosed
EmbodimentsOpen X-Embodiment robots, RoboArena robots
Data formatsmp4, parquet, json
Access modelopen
Commercial usePermitted
Year released2026

About This Dataset

RoboReward is a dataset and benchmark for training and evaluating general-purpose vision-language reward models for robotics, released in January 2026 by researchers at Stanford University and UC Berkeley (Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, Chelsea Finn). Each example pairs a task instruction with a real-robot rollout video and a discrete end-of-episode progress score from 1 to 5. The dataset contains 45,000 scored robot episodes built from large-scale real-robot corpora — Open X-Embodiment and RoboArena. Because these source corpora are heavily skewed toward successful demonstrations, RoboReward introduces counterfactual relabeling and temporal clipping to synthetically generate calibrated failure and near-miss examples from the same videos. It is used to train the RoboReward 4B and 8B reward models (fine-tuned from Qwen3-VL) and provides a standardized reward-accuracy benchmark. Released under CC BY 4.0 for commercial use.

What is it?

RoboReward is a dataset and benchmark for vision-language reward models in robotics, released in January 2026 by Stanford University and UC Berkeley researchers. Each example pairs a natural-language task instruction with a real-robot rollout video and a discrete end-of-episode progress reward score from 1 (no success) to 5 (perfect completion). The dataset contains 45,000 scored episodes across diverse tasks and embodiments, built from Open X-Embodiment and RoboArena.

Who is it for?

Researchers working on reinforcement-learning-based policy improvement who need automatic reward signals rather than labor-intensive human labeling or brittle hand-crafted objectives. RoboReward is distinct from demonstration datasets — it is used to train and evaluate reward models that judge whether a robot completed a task, closing much of the performance gap to human-given rewards.

Key specifications

  • Scored episodes: 45,000
  • Source corpora: Open X-Embodiment (OXE), RoboArena
  • Label: Discrete end-of-episode progress score 1 to 5
  • Augmentation: Counterfactual relabeling and temporal clipping to generate calibrated negatives and near-misses
  • Validation: Automated GPT-5-mini consistency check; test split additionally human-verified
  • Companion models: RoboReward 4B and 8B (fine-tuned from Qwen3-VL)
  • Format: MP4 rollout videos with task instructions and reward labels
  • License: CC BY 4.0 — commercial use permitted
  • Access: Open — Hugging Face

How it compares

RoboReward fills a category no other dataset in the catalogue covers: reward modeling and evaluation rather than demonstration data. Where Open X-Embodiment and DROID provide demonstrations for imitation learning, RoboReward provides scored success and failure examples for training reward functions. It is built on top of Open X-Embodiment, adding the failure and near-miss examples that success-heavy demonstration corpora lack.

Limitations and access notes

RoboReward is a reward-modeling dataset, not a source of demonstration trajectories for policy imitation. Its failure examples are partly synthetic, generated via counterfactual relabeling of successful episodes rather than collected from genuine robot failures. CC BY 4.0 permits commercial use with attribution.

Tasks Covered

This dataset covers manipulation, pick-and-place, inspection tasks across Open X-Embodiment robots, RoboArena robots robot platforms.

Access and Licensing

LicenseCC BY 4.0
Commercial usePermitted
Access modelFreely downloadable without registration

Citation

Stanford University & UC Berkeley (2026). RoboReward — Vision-Language Reward Dataset for Robotics. https://huggingface.co/datasets/teetone/RoboReward

Robot Embodiments

This dataset includes demonstrations collected using the following robot platforms.

  • Open X-Embodiment robots
  • RoboArena robots

Frequently Asked Questions

About Stanford University & UC Berkeley

RoboReward — Vision-Language Reward Dataset for Robotics is a Real Robot dataset maintained by Stanford University and UC Berkeley, released in 2026 under the CC BY 4.0 license. The dataset permits commercial use. It covers manipulation, pick-and-place, inspection tasks using Open X-Embodiment robots, RoboArena robots robot platforms. The full dataset is freely accessible at https://huggingface.co/datasets/teetone/RoboReward.