ICRA 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

ICRA 2026 — VLA & Manipulation Survey

Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program (ras.papercept.net/.../ICRA26/program, May 31 2026 snapshot).

At a glance

  • 5,088 submissions — a record for ICRA (official CFP).
  • The technical program lists ~2,820 archival papers (+ 131 late-breaking results), counted from the program. The official acceptance rate is not yet published; ICRA 2025 was 38.7% (1,606 / 4,153) for context.
  • Manipulation is the largest single theme. Topic-keyword paper counts (a paper may carry several): Reinforcement Learning 254 · Imitation Learning 197 · Deep Learning in Grasping & Manipulation 146 · Learning from Demonstration 133 · Force & Tactile Sensing 86 · Dexterous Manipulation 75 · Grasping 67 · Manipulation Planning 65 · Mobile Manipulation 36 · Bimanual 36 · Multifingered Hands 27.
  • ≥ 46 papers carry "Vision-Language-Action" / "VLA" in the title — VLA is now a mainstream ICRA topic, not a niche. The center of gravity differs from ICLR/CVPR: ICRA's VLA work is deployment- and sensor-centric (force/tactile, efficiency, streaming, edge, real-robot practicality), reflecting the venue's systems orientation.

How ICRA 2026 differs from the ML-venue VLA wave

Where ICLR 2026 emphasized architecture and CVPR 2026 emphasized perception & reasoning, ICRA 2026 is where VLA meets the robot:

  1. Sensor-grounded VLA — adding force, tactile, audio, and depth to the VLM→action stack (often without new hardware).
  2. Deployment efficiency — token pruning, streaming execution, edge inference, latency-aware design.
  3. RL / preference fine-tuning of continuous-action VLAs — making flow-matching and diffusion policies improvable from reward.
  4. Practicality benchmarks — moving past LIBERO toward latency / data-efficiency / robustness on real platforms.
  5. Navigation VLA as its own cluster, distinct from manipulation.

Detailed topic analyses — the full 728-paper corpus

Every manipulation/VLA paper in the ICRA 2026 program is grouped by technical topic and analyzed in a dedicated page (sub-trends + standout deep-dives + a complete paper table). The 728 papers partition as:

Topic Papers Detailed analysis
Vision-Language-Action models 45 [VLA](/Heungwoo/research/wiki/ICRA-2026-Topic-VLA)
Tactile & force-based manipulation 92 [Tactile & Force](/Heungwoo/research/wiki/ICRA-2026-Topic-Tactile-Force)
Dexterous & in-hand manipulation 75 [Dexterous](/Heungwoo/research/wiki/ICRA-2026-Topic-Dexterous)
Bimanual & dual-arm manipulation 50 [Bimanual](/Heungwoo/research/wiki/ICRA-2026-Topic-Bimanual)
Mobile manipulation 34 [Mobile](/Heungwoo/research/wiki/ICRA-2026-Topic-Mobile)
Assembly & contact-rich manipulation 58 [Assembly](/Heungwoo/research/wiki/ICRA-2026-Topic-Assembly)
Grasp synthesis & grippers 41 [Grasping](/Heungwoo/research/wiki/ICRA-2026-Topic-Grasping)
Perception for manipulation 82 [Perception](/Heungwoo/research/wiki/ICRA-2026-Topic-Perception)
Manipulation planning & TAMP 50 [Planning](/Heungwoo/research/wiki/ICRA-2026-Topic-Planning)
Diffusion & flow-matching policies 46 [Diffusion Policy](/Heungwoo/research/wiki/ICRA-2026-Topic-Diffusion-Policy)
Imitation learning & LfD (non-diffusion) 116 [Imitation Learning](/Heungwoo/research/wiki/ICRA-2026-Topic-Imitation)
RL · datasets · sim · representation 39 [RL/Data/Repr](/Heungwoo/research/wiki/ICRA-2026-Topic-RL-Data)
Total 728

(Partition is by primary keyword: VLA-titled papers first, then manipulation type, then learning paradigm. The curated reading list below highlights the most distinctive papers; the topic pages give exhaustive coverage.)

Themed reading list

1. Sensor-grounded VLA (force · tactile · audio · multi-sensor)

  • FD-VLA — distills a force token from vision+state into the VLM, giving force-aware contact-rich manipulation without a force/torque sensor (and the distilled token beats real sensor readings).
  • OmniVLA (manipulation; arXiv 2511.01210, Princeton/UCLA/MSRA) — physically-grounded multi-sensor (IR/radar/audio) fusion. (Name-collides with the navigation OmniVLA below.)
  • Also: CRAFT (force-aware curriculum fine-tuning), Audio-VLA (contact audio), Enhancing VLA Precision via FiLM-based Force/Torque-Vision Integration.

2. Spatial / 3D / depth grounding

  • DepthVLA (depth-aware spatial reasoning), AugVLA-3D (depth-driven feature augmentation), InSpire (intrinsic spatial reasoning), RetoVLA (reusing register tokens for spatial reasoning), Seeing Space and Motion (geometric + dynamic latent actions), ActiveVLA (active perception, arXiv 2601.08325), SG-VLA (spatially-grounded mobile manipulation, arXiv 2603.22760).

3. Reasoning · verification · human collaboration

  • VLA-Reasoner — plug-and-play online MCTS with a world model + value net; lifts OpenVLA real-world 22% → 41%.
  • CollabVLA (self-reflective, solicits human help), Do What You Say (runtime reasoning↔action alignment verification), Fast ECoT (7.5× faster embodied chain-of-thought), INSIGHT (inference-time help triggers).

4. Efficiency · streaming · edge deployment

  • LightVLA (differentiable token pruning) — differentiable, parameter-free visual-token pruning; −59% FLOPs / −38% latency while +2.9% success on OpenVLA-OFT (accuracy improves while pruning).
  • Also: Stream-To-Act (ROS 2-native token streaming for continuous execution), Robust Unknown Object Detection & Tracking for VLA on Edge Devices.

5. RL & preference fine-tuning of VLA

  • Flow Policy Optimization (FPO) — likelihood-free RL fine-tuning of flow-matching VLAs (π0); a different design point from ReinFlow (learnable-noise log-probs) and Qwen-VLA (ODE→SDE analytic log-prob).
  • Also: GRAPE (trajectory-level preference alignment, +51.8% in-domain / +58.2% unseen), Offline Reinforced Finetuning for chunk-based VLA, Toward Human Preference Optimization for VLA.

6. Data · pretraining · curation

  • Galaxea Open-World Dataset + G0 — 500 h / 100 K-trajectory open-world dataset + a dual-system VLA (Qwen2.5-VL planner + PaliGemma-3B flow-matching actor).
  • Also: Scalable VLA Pretraining from Real-Life Human Activity Videos, Developing VLA from Egocentric Videos (EgoScaler, +20% over from-scratch), ObjectVLA (synthetic multi-modal data → 100-object generalization), SCIZOR (self-supervised transition-level data curation, +15.4% avg).

7. Dexterous & bimanual VLA

  • Dexora — open-source 36-DoF dual-arm/dual-hand VLA; 66.7% dexterous-task success vs GR00T N1's 51.7%.
  • Also: DexGrasp AI Copilot (shared-autonomy arm-hand data collection, ~90% on 50+ objects).

8. Memory & retrieval

  • MAP-VLA — memory-augmented prompting (retrieved task-stage soft prompts) over a frozen π0; +7% sim / +25% real on long-horizon tasks. See VLA Memory review.
  • Also: ExpReS-VLA (experience replay + retrieval, arXiv 2511.06202).

9. Goal / world-model conditioning & zero-shot generalization

  • Goal-VLA — image-generative VLMs as object-centric world models; 59.9% on 8 RLBench tasks zero-shot (vs OpenVLA 0.2%, π0 0.0%). See Goal-Image-Conditioning review.
  • Also: Open-World Object Manipulation with VLA via Synthetic Multi-Modal Data.

10. Navigation VLA (distinct cluster)

  • OmniVLA (navigation) — omni-modal goal conditioning (2D pose / image / language) on an ~8.26B backbone; this is the cloud model used by AsyncVLA.
  • Also: UrbanVLA (urban micromobility), TrackVLA++ (embodied visual tracking with reasoning + memory).

11. Benchmarks · platforms · robustness · security

  • Rethinking VLA Practicality — CEBench + a 0.5B LLaVA-VLA baseline that matches a 7B model on CALVIN while training on a single RTX 4090. See RobustVLA, LIBERO-Plus.
  • Also: RealMirror (open-source VLA platform), Universal Adversarial Attacks on VLA (security), clutter-based evaluation protocols.

Cross-venue lineage

CoRL 2025 (π0.5; hierarchy + open-world)
  → NeurIPS 2025 (Knowledge Insulation, Real-Time Chunking; recipe + latency)
    → ICLR 2026 (architecture: discrete-diffusion, MoE, cross-embodiment)
      → CVPR 2026 (perception & reasoning: action-CoT, visual sketch, egocentric)
        → ICRA 2026 (DEPLOYMENT: sensor-grounded · efficient · RL-finetuned · practicality-benchmarked)

Related pages

← Back to ICRA-2026