ICLR 2026 RoboOmni - Heungwoo/research GitHub Wiki

RoboOmni — proactive manipulation from spoken dialogue, sounds, and vision

Venue: ICLR 2026 · Authors: Siyin Wang, Jinlan Fu, Feihong Liu, Xinzhe He, Huangxuan Wu, Junhao Shi, Kexin Huang, Zhaoye Fei, Jingjing Gong, Zuxuan Wu, Yu-Gang Jiang, See-Kiong Ng, Tat-Seng Chua, Xipeng Qiu · arXiv:2510.23763 · Category: agentic / omni-modal VLA · Trend tag: multimodal intention & proactivity.

Approach diagram

flowchart LR
  Speech[Spoken dialogue] --> Perc[Perceiver]
  Sound[Environmental sounds] --> Perc
  Vision[Visual scene] --> Perc
  Perc --> Think[Thinker: intention recognition]
  Think --> Talk[Talker: speech confirmation]
  Talk --> Exec[Executor: action policy]
  Exec --> Robot[Robot manipulation]

Problem

Most manipulation systems assume an explicit text command. In real settings, user intent is often implicit — conveyed through conversation, ambient sounds, and what is visible — and the robot must proactively infer intent rather than wait for a literal instruction. Training data for this "cross-modal contextual instruction" setting did not exist.

Method

RoboOmni is a Perceiver–Thinker–Talker–Executor framework built on an end-to-end omni-modal LLM that unifies intention recognition, interaction confirmation, and action execution. It fuses auditory and visual signals spatiotemporally for robust intention recognition and supports direct speech interaction (it can speak to confirm intent before acting).

To enable training, the authors build OmniAction: 140k episodes, 5k+ speakers, 2.4k event sounds, 640 backgrounds, and six contextual-instruction types. A LIBERO-based variant (OmniAction-LIBERO) is also released.

Results

  • In simulation and real-world experiments, RoboOmni surpasses text- and ASR-based baselines on success rate, inference speed, intention recognition, and proactive assistance.
  • ASR-pipeline baselines (speech → text → policy) are outperformed by the end-to-end omni-modal approach.

Significance

Reframes manipulation as proactive, context-driven assistance rather than literal command-following, and contributes a large omni-modal dataset (OmniAction) that combines speech, environmental audio, and vision — a step toward robots that infer what you want from natural multimodal context.

Links

Related pages

← Back to ICLR-2026