ICLR 2026 OmniEVA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Embodied reasoning · Planning Trend tag: 3D grounding · Embodiment-aware planning Authors: Huawei Noah's Ark Lab. Yuecheng Liu, Dafeng Chi, Shiguang Wu, Zhanguang Zhang (equal); corresponding Yuzheng Zhuang.
flowchart LR
subgraph Inputs[Inputs]
Img[RGB images / video<br/>16 frames train · 32 inference]
Depth[Depth maps + intrinsics K + extrinsics M_i]
Text[Instruction T]
end
Text --> ST[all-MiniLM-L6-v2<br/>sentence transformer · d_st=384]
Img --> ViT[InternVL3-8B ViT encoder<br/>448×448 patches]
Depth --> P3D[3D world coords<br/>0.1 m voxel · sinusoidal patch PE]
ViT --> AvgP[Avg-pool → V_avg^I · 1024-d]
ST --> Concat[Concat<br/>1024+384]
AvgP --> Concat
Concat --> MLP[2-layer MLP · hidden 256] --> Logits[V^g ∈ R²]
Logits --> Gumbel["Gumbel-Softmax · hard gating g ∈ {0,1}"]
ViT --> Hybrid[V_hybrid = V^I + g · V^p]
P3D --> Hybrid
Hybrid --> LLM[InternVL3-8B LLM]
Text --> LLM
LLM --> Resp[Response: NL · 2D/3D pts · 3D bbox]
LLM -.SFT then GRPO.- TEGRPO[TE-GRPO: r_format + r_task + λ_t · r_embod]
Two well-defined gaps in MLLM-based embodied planners:
- Geometric Adaptability Gap. 2D-only models (SpatialVLM, RoboPoint, RoboBrain) fail at tasks requiring depth, occlusion handling, or stacking; 3D-LLMs (3D-LLM, LLaVA-3D, 3DRS) hard-code 3D injection regardless of task — wasting compute and adding noise when 3D inputs are incomplete or irrelevant.
- Embodiment Constraint Gap. Models trained on web-scale or rule-based-synthetic data often produce plans that are theoretically valid but kinematically infeasible — they ignore object affordances, workspace limits, and arm reachability.
OmniEVA addresses both with a single architecture.
- Base MLLM: InternVL3-8B (Zhu et al., 2025).
- Sentence encoder for the gate: all-MiniLM-L6-v2 (Reimers & Face, 2021) — 384-d.
- Image resolution: 448 × 448.
- Video sampling: 16 frames train / 32 frames inference.
- 3D voxel size: 0.1 m (coordinates discretized for both inputs and outputs).
- Output formats: natural language; 2D points in [0, 1000]; 3D points/boxes in 0.1 m grid.
Patch-level 3D positional encoding. Each depth image Di ∈ RH×W is reprojected into world coordinates Pi ∈ RH×W×3 using camera intrinsics. Coords are partitioned into ViT-aligned patches and averaged → P′i ∈ RHp×Wp×3. A sinusoidal encoding maps each patch to a dv-dimensional vector → Vp.
Two conditioning signals to the gate:
- Task condition: VT = SentenceTransformer(T) ∈ R384
- Scene condition: VavgI = AvgPool(VI) ∈ R1024
Gate logits: Vg = MLPψ(Concat([VT, VavgI])) ∈ R²
Hard gating via Gumbel-Softmax (the paper argues against soft gating because soft-weighted sinusoidal PEs distort numerical stability):
g = GumbelSoftmax(Vg, τ) ∈ {0, 1} VhybridI = VI + g · Vp
This is mathematically equivalent to a 2-expert Mixture-of-Experts between pure-vision and 3D-augmented tokens. A KL term against Bernoulli(0.5) regularizes gating:
Ltotalψ,θ = LCE(olabel, o) + α · LKLψ(Vg ∥ Pprior)
Stage 1 — TAGR Pretraining. ViT frozen; sentence transformer frozen; MLP gate trained with lr=1e-4; LLM updated with low lr=5e-7. Datasets include depth-aware spatial reasoning data from ScanNet, Matterport3D, 3RScan, ArkitScenes. After pretraining, TAGR parameters are frozen and the fine-tuned LLM is discarded to avoid interference. Gumbel temperature scheduled exponentially: τ = τinit·exp(−4.5·steps / max_steps), τinit=1.0 → τmin=0.05.
Stage 2 — Supervised Fine-Tuning. Builds OmniEVA-Base. Hybrid dataset of (a) general embodied reasoning (2D images, video, 3D for spatial referring + scene captioning) and (b) custom embodied task data (navigation + manipulation). ViT frozen, LLM fine-tuned at lr=1e-5. Batch size 256; 1 epoch; cosine LR schedule.
Stage 3 — TE-GRPO RFT. Builds OmniEVA-ER. Builds on GRPO (Shao et al., 2024) with three rewards:
- rformat — encourages "......" structure.
- rtask = EvalTask(q, oi) ∈ [0,1] — semantic task completion (e.g., fraction of points inside a target region).
- rembod = EvalExec(q, oi) ∈ {0, 1} — kinematic feasibility check via a motion planner in simulation (verifies reachability + workspace + collisions).
Progressive embodiment curriculum:
ri,tacc = ritask · (λt · riembod + (1 − λt))
λt increases over time: early training rewards task completion alone; later steps demand strict embodiment feasibility. Group size G = 8 candidates per training step; LLM at lr=1e-5; batch size 128; AdamW; bf16; weight decay 0.1; gradient clip 1.0.
Trained only on SQA3D + ScanQA + Scan2Cap + ScanRefer training splits, fair comparison vs Video-3D-LLM and 3DRS:
| Method | SQA3D | ScanQA | Scan2Cap | ScanRefer | Average |
|---|---|---|---|---|---|
| Cross-Attention (separate vision & 3D tokens) | 55.1 | 27.5 | 43.3 | 4.5 | 32.6 |
| Cross-Attention (interleaved) | 55.8 | 27.5 | 42.0 | 3.6 | 32.2 |
| Hard-coded 3D Integration | 61.2 | 31.5 | 95.5 | 41.2 | 57.3 |
| Without 3D Integration | 61.2 | 30.7 | 75.5 | 4.3 | 42.9 |
| Dynamic 3D Integration, Soft Gating | 60.6 | 30.7 | 85.6 | 26.9 | 51.0 |
| Dynamic 3D Integration, Hard Gating (Ours) | 62.6 | 30.8 | 97.9 | 43.1 | 58.7 |
Cross-attention is catastrophic on Scan2Cap (-50 points). Soft gating consistently underperforms hard gating because soft weights distort the sinusoidal positional encodings.
3D activation rate by prompt cluster: Shape 76.9%, Action/Activity 50.9%, Visibility/Occlusion 33.0%, Direction/Orientation 30.5%, Identification 27.5%, Color 24.2%, Comparison 21.4%, Material/Texture 19.4%, Arrangement Pattern 13.6%, Counting 9.0%, State Condition 6.3%. The router cleanly learns that geometric and dynamic prompts need 3D, while descriptive/numerical ones don't.
| Model | Where2Place | VSI-bench | PACO-LVIS | RoboRefit | Where2Go | Where2Fit | Where2Approach | Where2Grasp |
|---|---|---|---|---|---|---|---|---|
| GPT-4o | 20.41 | 43.60 | 2.09 | 9.96 | 50.72 | 37.15 | 0.17 | 6.38 |
| Gemini-2.5-Pro | 28.60 | 48.83 | 3.14 | 17.91 | 55.07 | 41.82 | 3.50 | 27.00 |
| InternVL3-78B | 21.74 | 48.48 | 3.49 | 21.48 | 51.69 | 41.16 | 1.04 | 11.80 |
| Qwen2.5-VL-72B | 39.92 | 39.41 | 4.06 | 32.58 | 49.76 | 41.49 | 0.00 | 30.50 |
| RoboPoint | 46.80 | – | 9.21 | 47.83 | – | 56.64 | 2.46 | 35.97 |
| RoboBrain2.0-7B | 63.59 | 36.10 | 11.38 | 62.74 | 38.64 | 32.99 | 2.85 | 63.24 |
| RoboBrain2.0-32B | 73.59 | 42.69 | 16.23 | 69.98 | 41.06 | 59.23 | 4.35 | 67.60 |
| OmniEVA-Base (8B) | 74.95 | 57.17 | 21.01 | 91.19 | 86.96 | 78.14 | 7.37 | 73.91 |
Compact 8B model surpasses 32B and 72B baselines; +10.45 average gain over RoboBrain2.0-32B. PACO-LVIS jump (16.2 → 21.0) and RoboRefit jump (70.0 → 91.2) are particularly large.
| Model | SQA3D EM | ScanQA EM | Scan2Cap CIDEr | ScanRefer w/o.a |
|---|---|---|---|---|
| LLaVA-3D | 55.6 | 27.0 | 79.2 | 54.1 (w.a) |
| Inst3D-LMM | – | 24.6 | 79.7 | 57.8 (w.a) |
| Video-3D-LLM | 58.6 | 30.1 | 83.8 | 58.1 (w.a) |
| 3DRS | 60.6 | 30.3 | 86.1 | 62.9 (w.a) |
| OmniEVA-Base | 62.9 | 30.6 | 94.6 | 55.8 (w/o.a) |
Leads on three of four benchmarks. On ScanRefer (3D visual grounding) it slightly trails; the paper notes OmniEVA-Base reaches 55.8 [email protected] using only text I/O without external detectors or task-specific heads — vs prior text-only best of 44.4.
| Method | HM3D SR | HM3D SPL | MP3D SR | MP3D SPL |
|---|---|---|---|---|
| LFG | 68.9 | 36.0 | – | – |
| PIRLNav | 70.4 | 34.1 | – | – |
| UniNavid | 73.7 | 37.1 | – | – |
| OmniEVA-Base | 74.2 | 42.5 | 59.1 | 26.2 |
+5.4 SPL on HM3D over UniNavid; first work to score >40% SPL on HM3D-OVON in this setting.
Six tasks comparing RoboBrain2.0-7B, OmniEVA-Base, OmniEVA-ER w/o rembod, OmniEVA-ER (full):
| Task | RoboBrain2.0-7B | OmniEVA-Base | OmniEVA-ER w/o rembod | OmniEVA-ER |
|---|---|---|---|---|
| Where2Fit | 32.99 | 55.56 | 84.51 | 90.50 |
| Where2Approach | 2.81 | 5.58 | 6.25 | 39.86 |
| Where2Grasp | 63.24 | 67.81 | 71.00 | 82.50 |
| Mobile Placement (Easy) | – | 22.00 | 43.50 | 57.00 |
| Mobile Placement (Hard) | – | 4.00 | 7.00 | 52.00 |
| Mobile Pickup | 38.67 | 47.50 | 52.00 | 58.67 |
Where2Approach goes from 6.25 → 39.86 (a +33.6 jump); Mobile Placement (Hard) jumps 7 → 52. Removing rembod systematically degrades performance.
Models trained on arm lengths {75, 88, 110 cm}; evaluated on seen and unseen lengths from 72 cm to 105 cm:
| Avg SR | Seen Avg | Unseen Avg | |
|---|---|---|---|
| RoboBrain2.0-7B | 18.37 | 19.29 | 17.97 |
| OmniEVA-Base | 43.30 | 45.61 | 42.32 |
| OmniEVA-ER w/o rembod | 65.35 | 70.17 | 63.29 |
| OmniEVA-ER | 81.89 | 85.08 | 80.52 |
OmniEVA-ER attains 80.5% on unseen arm lengths — +17 points over OmniEVA-Base.
Two off-the-shelf wheeled mobile robots with different dual 6-DoF arm spans (75 cm vs Galaxea R1 at 70 cm), 10 trials per task:
| Cluttered Pick (Emb1 / Emb2) | Cluttered Place (Emb1 / Emb2) | Constrained Nav (Emb1 / Emb2) | |
|---|---|---|---|
| OmniEVA-Base | 6/10 / 5/10 | 6/10 / 5/10 | 9/10 / 8/10 |
| OmniEVA-ER | 8/10 / 7/10 | 9/10 / 9/10 | 8/10 / 8/10 |
OmniEVA-ER particularly excels on the more-constrained 70 cm embodiment for placement.
- Scene-level gating. TAGR makes one binary decision per (scene, task) — heterogeneous environments can lead to suboptimal 3D feature integration. The activation analysis shows unexpectedly low gate activation on some spatial-relationship prompts. Future work: patch-level gating for finer-grained 3D adaptation.
- Unmodeled physical constraints. Embodiment-aware planning currently models arm length; arm DoFs, installation height, joint torque limits are unmodeled.
OmniEVA is a counterpoint to "always-3D" embodied planners and a complement to "always-2D" VLAs:
- vs SpatialVLM / RoboPoint: they trained on 2D-only spatial QA; OmniEVA shows that adding 3D, but only when needed outperforms both alternatives — including hard-coded 3D fusion (Scan2Cap +2.4, ScanRefer +1.9, average +1.4).
- vs 3D-LLM / LLaVA-3D / Video-3D-LLM / 3DRS: these always inject 3D; OmniEVA's gate matches or beats them with cleaner 2D performance.
- vs RoboBrain 2.0 (7B/32B): RoboBrain has high-level planning + low-level pointing; OmniEVA-Base at 8B beats RoboBrain2.0-32B by +10.45 average across in-house embodied benchmarks while adding 3D awareness.
- vs π-family / Embodied-R1: OmniEVA stays at the planning layer (predicts subgoals/points/3D boxes); it is complementary to low-level VLA controllers.
- vs From Seeing to Doing: both push planners toward producing feasible (not just coherent) plans. OmniEVA does this through TE-GRPO with a kinematic-feasibility reward.
The TAGR + TE-GRPO combination — gate the 3D pathway dynamically, then RL-finetune on a feasibility reward — is a clean recipe other 3D-LLM efforts can adopt.
- OpenReview: https://openreview.net/forum?id=tkEmIJv1tB
- PDF: https://openreview.net/pdf?id=tkEmIJv1tB
← Back to ICLR-2026