Claru — Training Data for Physical AI

Claru is a purpose-built training data platform for physical AI, embodied AI, and world models. Backed by NVIDIA, Khosla Ventures, General Catalyst, and Y Combinator, Claru provides custom real-world video datasets, egocentric capture, manipulation data, warehouse and factory floor recordings, and synthetic augmentation across 100+ cities and 6 continents via a network of 10,000+ collectors. The platform delivers 4M+ human annotations across 100+ licensed datasets spanning egocentric, driving, manufacturing, cooking, warehouse, gaming, and human activity environments. Claru operates as a commercial data supplier to frontier AI labs training vision-language-action models, world models, and embodied AI systems — not an open dataset publisher. Data is delivered to client specification from brief to first delivery in days.

Dataset specifications
Year2024
Embodimentsrobot arms, egocentric cameras, autonomous vehicles, warehouse robots
Modalitiesrgb, depth, lidar, imu, tactile, thermal, stereo, point_cloud, audio, language, event_camera
Task categoriesmanipulation, pick-and-place, warehouse, cooking, human-robot-interaction, inspection, outdoor-navigation, long-horizon
Data formatmp4, json, parquet
LicenseCommercial — custom licensing per engagement
Accessproprietary — commercial use permitted
MaintainerClaru AI
Origin countryUS

What is it?

Claru is a commercial training data platform purpose-built for physical AI, embodied AI, and robotics foundation model development. Founded in 2024 and backed by NVIDIA, Khosla Ventures, General Catalyst, and Y Combinator, Claru provides end-to-end data pipelines for frontier labs — from real-world video capture and synthetic augmentation through enrichment, human annotation, and delivery in client-specified formats. The company operates a global collector network of 10,000+ contributors across 100+ cities on 6 continents, enabling data collection across any environment a model needs: residential kitchens, factory floors, city streets, warehouses, workshops, offices, and outdoor settings.

Who is it for?

Claru serves frontier AI labs and robotics companies training vision-language-action (VLA) models, world models, and embodied AI systems. Its primary customers are organisations that need large-scale, licensed, real-world data at specification — covering environments, tasks, or modalities that existing open datasets do not address. Case studies include 500K+ first-person egocentric videos for a world-modelling lab, 10,000+ hours of 3D game environment data for sim-to-real RL, and 2M+ cinematic video annotations for a frontier video generation company.

Key specifications

How it compares

Claru occupies a distinct position from open dataset publishers. Open datasets like EgoVerse (57,761 episodes, CC BY-SA 4.0), Open X-Embodiment (1M+ episodes, Apache 2.0), and DROID (76,000 episodes, Apache 2.0) provide freely available training data for academic and commercial use but are fixed collections covering specific tasks and embodiments. Claru provides custom data collection to specification — any environment, any task, any scale — but at commercial pricing with no public access. The two approaches are complementary rather than competing: open datasets provide broad baselines; Claru fills specific gaps that open datasets cannot address.

Limitations and access notes

Claru is not an open dataset provider. All data is licensed commercially and delivered under custom agreements. There is no public dataset download or Hugging Face repository. Pricing is not published — engagement begins with a scoping call. The platform is designed for organisations with dedicated ML infrastructure and data pipeline requirements, not individual researchers or academic labs with limited budgets.

Frequently asked questions

What is Claru?

Claru is a commercial training data platform for physical AI and robotics. Backed by NVIDIA, Khosla Ventures, General Catalyst, and Y Combinator, it provides custom real-world video datasets, egocentric capture, manipulation data, and synthetic augmentation across 100+ cities via 10,000+ collectors. Data is delivered to client specification rather than published as open datasets.

Is Claru data free to access?

No. Claru is a commercial data supplier — all datasets are licensed under custom agreements. There is no public download or open access. Engagement starts with a scoping call at claru.ai. Claru is distinct from open dataset publishers like EgoVerse or Open X-Embodiment.

What types of data does Claru provide?

Claru provides egocentric video, driving and urban navigation data, manufacturing and factory floor recordings, cooking and kitchen manipulation data, warehouse and logistics data, gaming environment data for sim-to-real RL, and human activity data. All data types are available in real-world capture and synthetic augmentation variants.

Who backs Claru?

Claru is backed by NVIDIA, Khosla Ventures, General Catalyst, and Y Combinator. The NVIDIA backing is notable given NVIDIA's central role in the physical AI infrastructure stack through Isaac GR00T, Isaac Lab, and the Jetson compute platform.

How does Claru compare to open robot training datasets?

Open datasets like Open X-Embodiment, DROID, and EgoVerse provide freely available training data for specific tasks and embodiments. Claru provides custom data collection to any specification at commercial scale — any environment, any task, any volume — but at commercial pricing with no public access. They serve different customer types: open datasets suit academic researchers and startups; Claru suits frontier labs needing proprietary data at scale.