OmniAction — Multimodal Proactive Robot Manipulation Dataset

OmniAction is a large-scale multimodal dataset for proactive robot manipulation containing 141,162 episodes across 112 skills and 748 objects. Released in March 2026 under CC BY-NC 4.0, the dataset is distinguished by its cross-modal contextual instruction approach — rather than explicit text commands, robots must infer intent from spoken dialogue, environmental sounds, and visual cues. The dataset includes 5,096 distinct speaker timbres, 2,482 non-verbal sound events, and 640 environmental backgrounds across six categories of contextual instructions. OmniAction addresses the gap between instruction-following datasets (where commands are explicit) and real-world deployment (where human intent must be inferred from context). With 117K downloads it is one of the most accessed physical AI datasets released in 2026.

Dataset specifications
Year2026
Embodimentsrobot arms, humanoid arms
Modalitiesrgb, audio, language
Task categoriesmanipulation, pick-and-place, human-robot-interaction, long-horizon
Data formatmp4, parquet, json
LicenseCC BY-NC 4.0
Accessopen
MaintainerLMMs Lab
Origin countryUS

What is it?

OmniAction is a large-scale multimodal dataset for proactive robot manipulation containing 141,162 episodes across 112 skills and 748 objects. Released in March 2026 under CC BY-NC 4.0, it is distinguished by cross-modal contextual instruction — robots must infer intent from spoken dialogue, environmental sounds, and visual cues rather than explicit text commands. The dataset includes 5,096 distinct speaker timbres, 2,482 non-verbal sound events, and 640 environmental backgrounds across six categories of contextual instructions. With 117K downloads it is one of the most accessed physical AI datasets of 2026.

Who is it for?

Researchers working on language-conditioned and context-aware robot manipulation — specifically the challenge of real-world deployment where humans give implicit rather than explicit instructions. OmniAction targets the gap between current instruction-following datasets (explicit commands) and real-world human-robot interaction (inferred intent from context).

Key specifications

How it compares

CALVIN is the standard benchmark for language-conditioned manipulation but uses explicit text instructions. OmniAction adds audio and contextual modalities that CALVIN lacks. Open X-Embodiment covers more embodiments but has no audio or contextual instruction component. OmniAction's multimodal instruction approach is unique among large-scale manipulation datasets.

Limitations and access notes

CC BY-NC 4.0 prohibits commercial use. The audio-contextual instruction approach requires models capable of processing multiple simultaneous input modalities — more complex pipeline than text-only datasets.

Linked professions

Frequently asked questions

What makes OmniAction different from other manipulation datasets?

OmniAction uses cross-modal contextual instructions — robots must infer what to do from spoken dialogue, environmental sounds, and visual cues rather than explicit text commands. This reflects real-world human-robot interaction where people don't always issue precise verbal commands. No other large-scale manipulation dataset at this scale includes audio and contextual instruction modalities.

Can OmniAction be used commercially?

No. OmniAction is licensed under CC BY-NC 4.0, which restricts use to non-commercial research. Commercial use requires a separate license from LMMs Lab.

How many episodes does OmniAction contain?

OmniAction contains 141,162 episodes across 112 manipulation skills and 748 distinct object classes, with contextual instructions derived from 5,096 speaker timbres and 2,482 non-verbal sound events.

How do I access OmniAction?

OmniAction is available on Hugging Face at huggingface.co/datasets/lmms-lab/OmniAction. No registration required.

What are the six categories of contextual instructions in OmniAction?

OmniAction's six instruction categories cover spoken dialogue (direct speech from humans), environmental sounds (appliances, objects), visual cues (gestures, pointing), implicit commands (statements implying an action), background context (ambient sounds indicating environment), and combined multimodal signals requiring fusion across all channels.