ICLR 2026 OmniEVA - Heungwoo/research GitHub Wiki

OmniEVA — Embodied Versatile Planner with Adaptive 3D Grounding

Venue: ICLR 2026 Category: Embodied reasoning · Planning Trend tag: 3D grounding · Embodiment-aware planning Authors: Huawei Noah's Ark Lab. Yuecheng Liu, Dafeng Chi, Shiguang Wu, Zhanguang Zhang (equal); corresponding Yuzheng Zhuang.

Approach diagram

flowchart LR
  subgraph Inputs[Inputs]
    Img[RGB images / video<br/>16 frames train · 32 inference] 
    Depth[Depth maps + intrinsics K + extrinsics M_i]
    Text[Instruction T]
  end
  Text --> ST[all-MiniLM-L6-v2<br/>sentence transformer · d_st=384]
  Img --> ViT[InternVL3-8B ViT encoder<br/>448×448 patches]
  Depth --> P3D[3D world coords<br/>0.1 m voxel · sinusoidal patch PE]
  ViT --> AvgP[Avg-pool → V_avg^I · 1024-d]
  ST --> Concat[Concat<br/>1024+384]
  AvgP --> Concat
  Concat --> MLP[2-layer MLP · hidden 256] --> Logits[V^g ∈ R²]
  Logits --> Gumbel["Gumbel-Softmax · hard gating g ∈ {0,1}"]
  ViT --> Hybrid[V_hybrid = V^I + g · V^p]
  P3D --> Hybrid
  Hybrid --> LLM[InternVL3-8B LLM]
  Text --> LLM
  LLM --> Resp[Response: NL · 2D/3D pts · 3D bbox]
  LLM -.SFT then GRPO.- TEGRPO[TE-GRPO: r_format + r_task + λ_t · r_embod]
Loading

Problem

Two well-defined gaps in MLLM-based embodied planners:

  1. Geometric Adaptability Gap. 2D-only models (SpatialVLM, RoboPoint, RoboBrain) fail at tasks requiring depth, occlusion handling, or stacking; 3D-LLMs (3D-LLM, LLaVA-3D, 3DRS) hard-code 3D injection regardless of task — wasting compute and adding noise when 3D inputs are incomplete or irrelevant.
  2. Embodiment Constraint Gap. Models trained on web-scale or rule-based-synthetic data often produce plans that are theoretically valid but kinematically infeasible — they ignore object affordances, workspace limits, and arm reachability.

OmniEVA addresses both with a single architecture.

Detailed Method

Backbone

  • Base MLLM: InternVL3-8B (Zhu et al., 2025).
  • Sentence encoder for the gate: all-MiniLM-L6-v2 (Reimers & Face, 2021) — 384-d.
  • Image resolution: 448 × 448.
  • Video sampling: 16 frames train / 32 frames inference.
  • 3D voxel size: 0.1 m (coordinates discretized for both inputs and outputs).
  • Output formats: natural language; 2D points in [0, 1000]; 3D points/boxes in 0.1 m grid.

Task-Adaptive Gated Router (TAGR)

Patch-level 3D positional encoding. Each depth image Di ∈ RH×W is reprojected into world coordinates Pi ∈ RH×W×3 using camera intrinsics. Coords are partitioned into ViT-aligned patches and averaged → P′i ∈ RHp×Wp×3. A sinusoidal encoding maps each patch to a dv-dimensional vector → Vp.

Two conditioning signals to the gate:

  • Task condition: VT = SentenceTransformer(T) ∈ R384
  • Scene condition: VavgI = AvgPool(VI) ∈ R1024

Gate logits: Vg = MLPψ(Concat([VT, VavgI])) ∈ R²

Hard gating via Gumbel-Softmax (the paper argues against soft gating because soft-weighted sinusoidal PEs distort numerical stability):

g = GumbelSoftmax(Vg, τ) ∈ {0, 1} VhybridI = VI + g · Vp

This is mathematically equivalent to a 2-expert Mixture-of-Experts between pure-vision and 3D-augmented tokens. A KL term against Bernoulli(0.5) regularizes gating:

Ltotalψ,θ = LCE(olabel, o) + α · LKLψ(Vg ∥ Pprior)

Three-stage training pipeline

Stage 1 — TAGR Pretraining. ViT frozen; sentence transformer frozen; MLP gate trained with lr=1e-4; LLM updated with low lr=5e-7. Datasets include depth-aware spatial reasoning data from ScanNet, Matterport3D, 3RScan, ArkitScenes. After pretraining, TAGR parameters are frozen and the fine-tuned LLM is discarded to avoid interference. Gumbel temperature scheduled exponentially: τ = τinit·exp(−4.5·steps / max_steps), τinit=1.0 → τmin=0.05.

Stage 2 — Supervised Fine-Tuning. Builds OmniEVA-Base. Hybrid dataset of (a) general embodied reasoning (2D images, video, 3D for spatial referring + scene captioning) and (b) custom embodied task data (navigation + manipulation). ViT frozen, LLM fine-tuned at lr=1e-5. Batch size 256; 1 epoch; cosine LR schedule.

Stage 3 — TE-GRPO RFT. Builds OmniEVA-ER. Builds on GRPO (Shao et al., 2024) with three rewards:

  • rformat — encourages "......" structure.
  • rtask = EvalTask(q, oi) ∈ [0,1] — semantic task completion (e.g., fraction of points inside a target region).
  • rembod = EvalExec(q, oi) ∈ {0, 1} — kinematic feasibility check via a motion planner in simulation (verifies reachability + workspace + collisions).

Progressive embodiment curriculum:

ri,tacc = ritask · (λt · riembod + (1 − λt))

λt increases over time: early training rewards task completion alone; later steps demand strict embodiment feasibility. Group size G = 8 candidates per training step; LLM at lr=1e-5; batch size 128; AdamW; bf16; weight decay 0.1; gradient clip 1.0.

Comprehensive Results

Ablation of 3D-integration design (Table 1)

Trained only on SQA3D + ScanQA + Scan2Cap + ScanRefer training splits, fair comparison vs Video-3D-LLM and 3DRS:

Method SQA3D ScanQA Scan2Cap ScanRefer Average
Cross-Attention (separate vision & 3D tokens) 55.1 27.5 43.3 4.5 32.6
Cross-Attention (interleaved) 55.8 27.5 42.0 3.6 32.2
Hard-coded 3D Integration 61.2 31.5 95.5 41.2 57.3
Without 3D Integration 61.2 30.7 75.5 4.3 42.9
Dynamic 3D Integration, Soft Gating 60.6 30.7 85.6 26.9 51.0
Dynamic 3D Integration, Hard Gating (Ours) 62.6 30.8 97.9 43.1 58.7

Cross-attention is catastrophic on Scan2Cap (-50 points). Soft gating consistently underperforms hard gating because soft weights distort the sinusoidal positional encodings.

TAGR activation per semantic category (Figure 4)

3D activation rate by prompt cluster: Shape 76.9%, Action/Activity 50.9%, Visibility/Occlusion 33.0%, Direction/Orientation 30.5%, Identification 27.5%, Color 24.2%, Comparison 21.4%, Material/Texture 19.4%, Arrangement Pattern 13.6%, Counting 9.0%, State Condition 6.3%. The router cleanly learns that geometric and dynamic prompts need 3D, while descriptive/numerical ones don't.

2D + in-house benchmarks (Table 2)

Model Where2Place VSI-bench PACO-LVIS RoboRefit Where2Go Where2Fit Where2Approach Where2Grasp
GPT-4o 20.41 43.60 2.09 9.96 50.72 37.15 0.17 6.38
Gemini-2.5-Pro 28.60 48.83 3.14 17.91 55.07 41.82 3.50 27.00
InternVL3-78B 21.74 48.48 3.49 21.48 51.69 41.16 1.04 11.80
Qwen2.5-VL-72B 39.92 39.41 4.06 32.58 49.76 41.49 0.00 30.50
RoboPoint 46.80 – 9.21 47.83 – 56.64 2.46 35.97
RoboBrain2.0-7B 63.59 36.10 11.38 62.74 38.64 32.99 2.85 63.24
RoboBrain2.0-32B 73.59 42.69 16.23 69.98 41.06 59.23 4.35 67.60
OmniEVA-Base (8B) 74.95 57.17 21.01 91.19 86.96 78.14 7.37 73.91

Compact 8B model surpasses 32B and 72B baselines; +10.45 average gain over RoboBrain2.0-32B. PACO-LVIS jump (16.2 → 21.0) and RoboRefit jump (70.0 → 91.2) are particularly large.

3D embodied reasoning (Table 3)

Model SQA3D EM ScanQA EM Scan2Cap CIDEr ScanRefer w/o.a
LLaVA-3D 55.6 27.0 79.2 54.1 (w.a)
Inst3D-LMM – 24.6 79.7 57.8 (w.a)
Video-3D-LLM 58.6 30.1 83.8 58.1 (w.a)
3DRS 60.6 30.3 86.1 62.9 (w.a)
OmniEVA-Base 62.9 30.6 94.6 55.8 (w/o.a)

Leads on three of four benchmarks. On ScanRefer (3D visual grounding) it slightly trails; the paper notes OmniEVA-Base reaches 55.8 [email protected] using only text I/O without external detectors or task-specific heads — vs prior text-only best of 44.4.

Object navigation (Table 4)

Method HM3D SR HM3D SPL MP3D SR MP3D SPL
LFG 68.9 36.0 – –
PIRLNav 70.4 34.1 – –
UniNavid 73.7 37.1 – –
OmniEVA-Base 74.2 42.5 59.1 26.2

+5.4 SPL on HM3D over UniNavid; first work to score >40% SPL on HM3D-OVON in this setting.

Embodiment-aware reasoning (Figure 5)

Six tasks comparing RoboBrain2.0-7B, OmniEVA-Base, OmniEVA-ER w/o rembod, OmniEVA-ER (full):

Task RoboBrain2.0-7B OmniEVA-Base OmniEVA-ER w/o rembod OmniEVA-ER
Where2Fit 32.99 55.56 84.51 90.50
Where2Approach 2.81 5.58 6.25 39.86
Where2Grasp 63.24 67.81 71.00 82.50
Mobile Placement (Easy) – 22.00 43.50 57.00
Mobile Placement (Hard) – 4.00 7.00 52.00
Mobile Pickup 38.67 47.50 52.00 58.67

Where2Approach goes from 6.25 → 39.86 (a +33.6 jump); Mobile Placement (Hard) jumps 7 → 52. Removing rembod systematically degrades performance.

Cross-embodiment generalization by arm length (Table 5)

Models trained on arm lengths {75, 88, 110 cm}; evaluated on seen and unseen lengths from 72 cm to 105 cm:

Avg SR Seen Avg Unseen Avg
RoboBrain2.0-7B 18.37 19.29 17.97
OmniEVA-Base 43.30 45.61 42.32
OmniEVA-ER w/o rembod 65.35 70.17 63.29
OmniEVA-ER 81.89 85.08 80.52

OmniEVA-ER attains 80.5% on unseen arm lengths — +17 points over OmniEVA-Base.

Real robot deployment (Table 6)

Two off-the-shelf wheeled mobile robots with different dual 6-DoF arm spans (75 cm vs Galaxea R1 at 70 cm), 10 trials per task:

Cluttered Pick (Emb1 / Emb2) Cluttered Place (Emb1 / Emb2) Constrained Nav (Emb1 / Emb2)
OmniEVA-Base 6/10 / 5/10 6/10 / 5/10 9/10 / 8/10
OmniEVA-ER 8/10 / 7/10 9/10 / 9/10 8/10 / 8/10

OmniEVA-ER particularly excels on the more-constrained 70 cm embodiment for placement.

Limitations (as stated in §5)

  1. Scene-level gating. TAGR makes one binary decision per (scene, task) — heterogeneous environments can lead to suboptimal 3D feature integration. The activation analysis shows unexpectedly low gate activation on some spatial-relationship prompts. Future work: patch-level gating for finer-grained 3D adaptation.
  2. Unmodeled physical constraints. Embodiment-aware planning currently models arm length; arm DoFs, installation height, joint torque limits are unmodeled.

Significance & Positioning

OmniEVA is a counterpoint to "always-3D" embodied planners and a complement to "always-2D" VLAs:

  • vs SpatialVLM / RoboPoint: they trained on 2D-only spatial QA; OmniEVA shows that adding 3D, but only when needed outperforms both alternatives — including hard-coded 3D fusion (Scan2Cap +2.4, ScanRefer +1.9, average +1.4).
  • vs 3D-LLM / LLaVA-3D / Video-3D-LLM / 3DRS: these always inject 3D; OmniEVA's gate matches or beats them with cleaner 2D performance.
  • vs RoboBrain 2.0 (7B/32B): RoboBrain has high-level planning + low-level pointing; OmniEVA-Base at 8B beats RoboBrain2.0-32B by +10.45 average across in-house embodied benchmarks while adding 3D awareness.
  • vs π-family / Embodied-R1: OmniEVA stays at the planning layer (predicts subgoals/points/3D boxes); it is complementary to low-level VLA controllers.
  • vs From Seeing to Doing: both push planners toward producing feasible (not just coherent) plans. OmniEVA does this through TE-GRPO with a kinematic-feasibility reward.

The TAGR + TE-GRPO combination — gate the 3D pathway dynamically, then RL-finetune on a feasibility reward — is a clean recipe other 3D-LLM efforts can adopt.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️