RoboReward — Vision-Language Reward Dataset for Robotics
About This Dataset
RoboReward is a dataset and benchmark for training and evaluating general-purpose vision-language reward models for robotics, released in January 2026 by researchers at Stanford University and UC Berkeley (Tony Lee, Andrew Wagenmaker, Karl Pertsch, Percy Liang, Sergey Levine, Chelsea Finn). Each example pairs a task instruction with a real-robot rollout video and a discrete end-of-episode progress score from 1 to 5. The dataset contains 45,000 scored robot episodes built from large-scale real-robot corpora — Open X-Embodiment and RoboArena. Because these source corpora are heavily skewed toward successful demonstrations, RoboReward introduces counterfactual relabeling and temporal clipping to synthetically generate calibrated failure and near-miss examples from the same videos. It is used to train the RoboReward 4B and 8B reward models (fine-tuned from Qwen3-VL) and provides a standardized reward-accuracy benchmark. Released under CC BY 4.0 for commercial use.
What is it?
RoboReward is a dataset and benchmark for vision-language reward models in robotics, released in January 2026 by Stanford University and UC Berkeley researchers. Each example pairs a natural-language task instruction with a real-robot rollout video and a discrete end-of-episode progress reward score from 1 (no success) to 5 (perfect completion). The dataset contains 45,000 scored episodes across diverse tasks and embodiments, built from Open X-Embodiment and RoboArena.
Who is it for?
Researchers working on reinforcement-learning-based policy improvement who need automatic reward signals rather than labor-intensive human labeling or brittle hand-crafted objectives. RoboReward is distinct from demonstration datasets — it is used to train and evaluate reward models that judge whether a robot completed a task, closing much of the performance gap to human-given rewards.
Key specifications
- •Scored episodes: 45,000
- •Source corpora: Open X-Embodiment (OXE), RoboArena
- •Label: Discrete end-of-episode progress score 1 to 5
- •Augmentation: Counterfactual relabeling and temporal clipping to generate calibrated negatives and near-misses
- •Validation: Automated GPT-5-mini consistency check; test split additionally human-verified
- •Companion models: RoboReward 4B and 8B (fine-tuned from Qwen3-VL)
- •Format: MP4 rollout videos with task instructions and reward labels
- •License: CC BY 4.0 — commercial use permitted
- •Access: Open — Hugging Face
How it compares
RoboReward fills a category no other dataset in the catalogue covers: reward modeling and evaluation rather than demonstration data. Where Open X-Embodiment and DROID provide demonstrations for imitation learning, RoboReward provides scored success and failure examples for training reward functions. It is built on top of Open X-Embodiment, adding the failure and near-miss examples that success-heavy demonstration corpora lack.
Limitations and access notes
RoboReward is a reward-modeling dataset, not a source of demonstration trajectories for policy imitation. Its failure examples are partly synthetic, generated via counterfactual relabeling of successful episodes rather than collected from genuine robot failures. CC BY 4.0 permits commercial use with attribution.
Tasks Covered
This dataset covers manipulation, pick-and-place, inspection tasks across Open X-Embodiment robots, RoboArena robots robot platforms.
Access and Licensing
Citation
Stanford University & UC Berkeley (2026). RoboReward — Vision-Language Reward Dataset for Robotics. https://huggingface.co/datasets/teetone/RoboReward
Robot Embodiments
This dataset includes demonstrations collected using the following robot platforms.
- Open X-Embodiment robots
- RoboArena robots
Related Datasets
SPARC (Spatial Annotations from Robot Demonstrations at Scale)
Intuitive Robots Research Team · 2026 · null
Real RobotGrandTour — Legged Robot Dataset in the Wild
Robotic Systems Lab · 2026 · MIT — confirm in repository LICENSE file before activation
Real RobotPalmDex — Embodiment-Agnostic Tactile Manipulation Dataset
PalmDex Research Team · 2026 · CC BY-SA 4.0
Real RobotGR1 Robot Teleoperation Dataset
NVIDIA · 2026 · MIT
Frequently Asked Questions
About Stanford University & UC Berkeley
RoboReward — Vision-Language Reward Dataset for Robotics is a Real Robot dataset maintained by Stanford University and UC Berkeley, released in 2026 under the CC BY 4.0 license. The dataset permits commercial use. It covers manipulation, pick-and-place, inspection tasks using Open X-Embodiment robots, RoboArena robots robot platforms. The full dataset is freely accessible at https://huggingface.co/datasets/teetone/RoboReward.
