ICLR 2026 RoboOmni - Heungwoo/research GitHub Wiki
RoboOmni — proactive manipulation from spoken dialogue, sounds, and vision
Venue: ICLR 2026 · Authors: Siyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He, Huangxuan Wu, Junhao Shi, Kexin Huang, Zhaoye Fei, Jingjing Gong, Zuxuan Wu, Yu-Gang Jiang, See-Kiong Ng, Tat-Seng Chua, Xipeng Qiu · arXiv:2510.23763 · Category: agentic / omni-modal VLA · Trend tag: multimodal intention & proactivity.
Approach diagram
flowchart LR
Speech[Spoken dialogue] --> Perc[Perceiver]
Sound[Environmental sounds] --> Perc
Vision[Visual scene] --> Perc
Perc --> Think[Thinker: intention recognition]
Think --> Talk[Talker: speech confirmation]
Talk --> Exec[Executor: action policy]
Exec --> Robot[Robot manipulation]
Problem
Most manipulation systems assume an explicit text command. In real settings, user intent is often implicit — conveyed through conversation, ambient sounds, and what is visible — and the robot must proactively infer intent rather than wait for a literal instruction. Training data for this "cross-modal contextual instruction" setting did not exist.
Method
RoboOmni is a Perceiver–Thinker–Talker–Executor framework built on an end-to-end omni-modal LLM that unifies intention recognition, interaction confirmation, and action execution. It fuses auditory and visual signals spatiotemporally for robust intention recognition and supports direct speech interaction (it can speak to confirm intent before acting).
To enable training, the authors build OmniAction: 140k episodes, 5k+ speakers, 2.4k event sounds, 640 backgrounds, and six contextual-instruction types. A LIBERO-based variant (OmniAction-LIBERO) is also released.
Results
- In simulation and real-world experiments, RoboOmni surpasses text- and ASR-based baselines on success rate, inference speed, intention recognition, and proactive assistance.
- ASR-pipeline baselines (speech → text → policy) are outperformed by the end-to-end omni-modal approach.
Significance
Reframes manipulation as proactive, context-driven assistance rather than literal command-following, and contributes a large omni-modal dataset (OmniAction) that combines speech, environmental audio, and vision — a step toward robots that infer what you want from natural multimodal context.
Links
Related pages
← Back to ICLR-2026