IROS 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

IROS 2026 β€” VLA & Manipulation Survey

Venue: IEEE/RSJ IROS 2026 Β· Pittsburgh, USA Β· Sep 27 – Oct 1, 2026 Β· David L. Lawrence Convention Center. Source: the official IROS 2026 program (2026.ieee-iros.org/program, public, login-free). Paper decisions announced Jun 16, 2026; 1,900+ contributed papers, each presented as both a talk and a poster across 15 focused + 8 lightning rooms. Sourcing note. Compiled from the official program before the conference (Sep 2026). The program's searchable index exposes titles + authors + affiliations + abstracts (searches run locally in-browser), so summaries here are abstract-informed and program-verified. Final IEEE Xplore versions publish at/after the conference; treat paper IDs (#…) as program IDs.


1. At a glance

  • 1,900+ papers total; manipulation, dexterity, imitation/policy learning, and VLA/LLM are a major slice of the technical program.
  • The program is organized into named focused sessions (the venue's topical taxonomy) β€” Β§2 maps the in-scope ones. This is the most reliable pre-conference signal of what IROS 2026 emphasizes.
  • Theme of the keynotes: "Open Problems and Perspective Shifts," plus a plenary panel on the Specialist/Generalist debate β€” directly relevant to the VLA generalist-policy thread.
  • The manipulation center of gravity (from session names): policy learning at scale, VLA robustness/efficiency, in-context & demonstration-efficient imitation, dexterous grasping, and LLM/language in the control loop.

2. The in-scope session taxonomy (official focused sessions)

IROS 2026 names its sessions around specific topics; these are the manipulation/VLA-relevant ones (β‰ˆ40 focused + lightning sessions):

A. Learning & policy for manipulation Shaping the Learning Problem for Robot Control Β· Making Imitation Learning Faster and Leaner Β· RL for Contact-Rich Manipulation Skills (Γ—2) Β· Planning Manipulation for Assembly and Deformable Objects Β· Humans in the Robot Learning Loop Β· Building Robustness into Learned Robot Control Β· Generalization and Robustness in Learned Manipulation Β· Learning and Stress-Testing Robot Skill Models Β· Force and Motion in Learning from Demonstration Β· Data-Efficient Policy Learning for Manipulation Β· Representations for Generalizable Manipulation Policies Β· Learning Smarter Policies for Robot Manipulation Β· Scaling Up Robot Policy Learning Β· Advancing Policy Learning for Robotic Manipulation Β· Making RL Practical for Manipulation Β· Refining Learned Policies for Precision Contact Tasks

B. VLA / LLM / foundation models Extending Vision-Language-Action Model Capabilities Β· Building Memory into Vision-Language Robot Agents Β· Putting Language Models in the Control Loop Β· LLM Agents Reasoning Through Robot Tasks Β· Constraint-Guided Reasoning for Long-Horizon Manipulation Β· Verification and Safety for LLM-Guided Planning

C. Dexterous & grasping Generalizable Dexterous Grasping Β· Building Dexterous Manipulation Skills Β· Contact Sensing and Design for Dexterous Hands Β· Sensing Touch for Smarter Grasping Β· Sensing Contact for Steadier Grasping Β· Perception and Staging for In-Hand Manipulation Β· Sharpening Perception for Robotic Grasping Β· Grounding Manipulation Policies in 3D Perception Β· Dexterous Hands and Wrists for Assistive Robotics Β· Building Better Interfaces for Dexterous Teleoperation

D. Contact-rich & perception-for-manipulation Planning Contact-Rich Manipulation Sequences Β· Perception and Learning for Robot/Robust Manipulation Β· Sensing Contact for Smarter Manipulation Β· Planning Manipulation Under Uncertainty and Constraints

Out of scope β€” navigation & mobility. IROS 2026 also runs a large navigation cluster (Learning Robust Control for Visual Navigation Β· Safe and Predictive Robot Navigation Β· Modeling Motion and Environment for Navigation Β· Understanding and Navigating Crowded Social Spaces Β· Anticipating and Navigating Around Humans Β· New Signals and Maps for Drone Navigation, plus VLN/social-nav and world-model-for-navigation work like NavThinker). This survey is manipulation-scoped; the manipulation↔navigation crossover (mobile manipulation, loco-manipulation, joint nav+manip planning) is where the two meet β€” see the navigation links in Β§7.


3. Sampled papers by theme (extracted from the official program; titles + authors + abstracts available)

3.1 Vision-Language-Action models β€” "Extending VLA Capabilities" session

  • AnyCamVLA (#359) β€” zero-shot camera/viewpoint adaptation for viewpoint-robust VLAs.
  • LangGap (#522) β€” diagnosing and closing the language gap in VLAs (the instruction-following pathology β€” cf. Multi-Task VLA, LangForce).
  • OG-VLA (#3972) β€” orthographic image generation for a 3D-aware VLA.
  • BFA++ (#215) β€” hierarchical best-feature-aware token pruning for multi-view VLA (efficiency).
  • Safe-Night VLA (#2899) β€” thermal-perceptive VLA for safety-critical low-light tasks.
  • From Code to Action (#2158) β€” hierarchical learning of diffusion-VLM policies.
  • AtomVLA (#33) πŸ“„ β€” subtask-aware VLA + latent world-model scoring for offline GRPO post-training (LIBERO 97.0%). (full-paper analysis)
  • (efficiency cross-listed) IMLE-VLA (#4624) πŸ“„ β€” single-step action head (55 Hz, LIBERO 98.0%) Β· Fast Enough to Act β€” spatio-temporal visual-token merging for low-latency VLAs.

3.2 Scaling & RL for policy learning β€” "Scaling Up Robot Policy Learning"

  • VLA-RL (#4707) πŸ“„ β€” masterful, general manipulation via scalable reinforcement learning on a VLA (OpenVLA-7B +4.5% LIBERO; real 60β†’90%). (full-paper analysis)
  • Trajectory-Consistent Flow Matching (#3959) β€” robust visuomotor flow-matching policy learning.
  • LAR-MoE (#1956) β€” latent-aligned routing for Mixture-of-Experts in robotic imitation (cf. Multi-Task VLA MoE cluster).
  • Compositional Average-Velocity Modeling (#4500) β€” efficient diffusion offline RL.

3.3 In-context & demonstration-efficient imitation β€” "Making Imitation Learning Faster and Leaner" + scaling

  • RoboSSM (#3133) β€” scalable in-context imitation learning via state-space models (cf. In-Context Imitation).
  • ICLR: In-Context Imitation Learning with Visual Reasoning β€” in-context demo-following with visual reasoning.
  • IMLE-VLA β€” fast single-step action generation for VLA policies (real-time).
  • ESPADA β€” semantics-aware demonstration downsampling for faster imitation.
  • What Can Robot Foundation Models Learn from Limited Human Demonstrations? β€” fine-tuning with human task priors.

3.4 Manipulation policy β€” "Advancing Policy Learning for Robotic Manipulation"

  • MaskVLA β€” visual masking against trajectory overfitting of a VLA.
  • SynthLA β€” synthetic language-action multi-task policies for zero-shot real-world manipulation.
  • Jointly Learning Predicates and Actions β€” zero-shot skill composition.
  • RobustFuzz β€” semantic-disturbance benchmark for learned embodied policies.
  • Prior Reinforce β€” goal-conditioned dynamic manipulation with limited trials.

3.5 Language / reasoning in the loop β€” "Putting Language Models in the Control Loop" / LLM-agents

  • NL2SpaTiaL β€” geometric spatio-temporal logic specs from natural language for manipulation.
  • AURORA β€” dual-verification with adaptive rollback for LLM agents.
  • (session also covers) Constraint-Guided Reasoning for Long-Horizon Manipulation Β· Verification & Safety for LLM-Guided Planning.

3.6 ⭐ Representative papers by technical theme (full-coverage selection)

Selection criterion: one technically representative paper per distinct technical theme across the in-scope program β€” chosen for methodological distinctiveness (a clearly different mechanism), not by lab/vendor prestige or citation. The aim is to cover the technical space of IROS 2026 manipulation; each has an in-depth full-paper analysis (πŸ“„).

Technical theme Representative paper Core mechanism
VLA efficiency / real-time [IMLE-VLA](/Heungwoo/research/wiki/IROS-2026-IMLE-VLA) πŸ“„ single-step cIMLE action head (55 Hz)
Scalable online RL for VLA [VLA-RL](/Heungwoo/research/wiki/IROS-2026-VLA-RL) πŸ“„ trajectory-as-conversation + VLM process-reward
WM as offline critic (post-training) [AtomVLA](/Heungwoo/research/wiki/IROS-2026-AtomVLA) πŸ“„ latent-WM scoring β†’ offline GRPO
WM as cross-embodiment interface [Scaling Cross-Embodiment WM](/Heungwoo/research/wiki/IROS-2026-Cross-Embodiment-WM) πŸ“„ particle-based shared WM + MPC
WM as humanoid distillation scaffold [DreamMimic](/Heungwoo/research/wiki/IROS-2026-DreamMimic) πŸ“„ RSSM latent dynamics as distillation supervision
Bimanual β€” unified 3D policy [3D FlowMatch Actor](/Heungwoo/research/wiki/IROS-2026-3D-FlowMatch-Actor) πŸ“„ 3D-scene-token flow matching, single+dual-arm
Bimanual β€” symmetry prior [EquiBim](/Heungwoo/research/wiki/IROS-2026-EquiBim) πŸ“„ bilateral-equivariant policy
In-context imitation β€” long context [RoboSSM](/Heungwoo/research/wiki/IROS-2026-RoboSSM) πŸ“„ state-space model (linear, extrapolates)
In-context imitation β€” intent [ICLR: Visual Reasoning](/Heungwoo/research/wiki/IROS-2026-ICLR-Visual-Reasoning) πŸ“„ image-space reasoning traces
Long-horizon memory (retrofit) [TempoFit](/Heungwoo/research/wiki/IROS-2026-TempoFit) πŸ“„ training-free KV-cache memory
Dexterous grasping β€” keypoint VLA [DexKP-VLA](/Heungwoo/research/wiki/IROS-2026-DexKP-VLA) πŸ“„ VLM semantics β†’ keypoints β†’ optimizer, zero-shot
Dexterous grasping β€” reasoning [VCoT-Grasp](/Heungwoo/research/wiki/IROS-2026-VCoT-Grasp) πŸ“„ end-to-end grasp FM + visual chain-of-thought
Deformable + tactile-in-loop [MrGrasp](/Heungwoo/research/wiki/IROS-2026-MrGrasp) πŸ“„ multi-rate VTLA, fast tactile loop under VLM planner
In-hand reorientation [Proprioceptive Transformer](/Heungwoo/research/wiki/IROS-2026-Proprioceptive-Transformer) πŸ“„ joint-sensing-only cube rotation (award finalist)
Sim-to-real + articulated objects [Articulated-Tools Sim2Real](/Heungwoo/research/wiki/IROS-2026-Articulated-Tools-Sim2Real) πŸ“„ sim base + hardware-demo tactile refinement
Force-from-demonstration [FILIC](/Heungwoo/research/wiki/IROS-2026-FILIC) πŸ“„ dual-loop force-guided IL + impedance, no F/T sensor

Coverage note: this set now spans 17 technical themes β€” efficiency Β· RL Β· WM-as-scaffold (3 modes) Β· bimanual (2) Β· in-context (2) Β· memory Β· dexterous grasping (2) Β· deformable+tactile Β· in-hand Β· sim-to-real Β· force-LfD. This covers the manipulation program's core technical space; the remaining sessions in Β§2 are variations within these themes.


4. What IROS 2026 signals (from the session structure)

  1. VLA has its own dedicated sessions (Extending VLA Capabilities; Building Memory into VL Robot Agents) β€” VLA is now a first-class IROS track, not scattered. The emphasis skews to robustness (viewpoint, language gap, adversarial), efficiency (token pruning/merging, single-step), and memory β€” maturity concerns, not just new backbones.
  2. Policy-learning at scale is the largest cluster (~16 sessions) β€” scaling, data-efficiency, generalization/robustness, RL-made-practical dominate over novel action-head architectures. VLA-RL and MoE-routing (LAR-MoE) echo the Multi-Task VLA fixes.
  3. In-context / demonstration-efficient imitation is a visible sub-theme (RoboSSM, ICLR-visual-reasoning, ESPADA, human-task-prior fine-tuning) β€” converging with the In-Context Imitation and egocentric-video pretraining threads.
  4. Dexterous manipulation stays hardware-and-contact-heavy (10 sessions) β€” grasping, tactile/contact sensing, in-hand, teleop interfaces β€” the applied counterpart to the Dexterous-Hand Data Pyramid.
  5. The Specialist/Generalist plenary debate frames IROS 2026's meta-question β€” exactly the single-checkpoint multi-robot / generalist-VLA tension.

Positioning vs other 2026 venues: where ICLR/ICML/RSS 2026 pushed new architectures & representations, IROS 2026 reads as the robustness/efficiency/deployment counterpart β€” the manipulation program is heavy on making learned & VLA policies work reliably, efficiently, and from little data.


5. Trend deep-dives (three requested lenses)

5.1 World-Action Models β€” WAMs move from generation to manipulation policy & data

IROS 2026 has a substantial world-model manipulation cluster, and the center of gravity has shifted from pretty video generation toward using the WAM as policy, evaluator, or data engine:

  • WAM as VLA post-training / policy: AtomVLA πŸ“„ β€” scalable post-training for manipulation via predictive latent world models used as a critic for offline GRPO (reconstruction-free, deployable corner β€” cf. Ο‰-0, DYNA-2); RoDyn β€” an interactive robot-dynamic 2.5D world model for manipulation; Generalizable Robotic Insertion with World Models (contact-rich). (AtomVLA has a full-paper analysis)
  • WAM Γ— cross-embodiment / dexterity: Scaling Cross-Embodiment World Models πŸ“„ β€” a particle-based WM as a shared geometric interface across human/robot hands + MPC (cf. Cross-Embodiment).
  • WAM Γ— humanoid: DreamMimic πŸ“„ β€” whole-body loco-manipulation via a WM used as distillation supervision + representation (RSSM repurposed, not for planning); Learning Versatile Humanoid Manipulation with Touch Dreaming β€” a tactile world model.
  • WAM as data factory / evaluator: RoboDream β€” compositional world models for scalable robot data synthesis; GrndCtrl β€” grounding world models via self-supervised reward alignment; plus surgical-policy online evaluation via a world foundation model.

Insight: IROS 2026 confirms the World Models thesis at the systems level β€” the live questions are latent-vs-pixel (AtomVLA's latent WAM), cross-embodiment scaling, and WAM-as-data/eval, not whether WAMs help. The reconstruction-free latent WAM (deployable, low-latency) is the ascendant flavor, echoing Ο‰-0/DYNA-2.

πŸ†• WAM-vs-VLA trend update (from the full-coverage set). A sharper pattern emerges once you look across all the IROS 2026 WAM papers: the world model is being pushed out of the inference loop and into the training/representation scaffold. Not one of the representative WAM papers runs the WM as the deployed policy β€”

  • AtomVLA: WM = offline RL critic (scores chunks, absent at deploy);
  • DreamMimic: WM = distillation supervisor + representation (RSSM repurposed away from planning; student is a plain vision policy);
  • Scaling Cross-Embodiment WM: WM = shared cross-embodiment interface feeding classical MPC, not an end-to-end VLA.

This extends (not overturns) the convergence story in VLA Hybrid Architectures and WAM vs VLA Robustness: the 2025 debate was "WAM-backbone (imagineβ†’act, in-path, 3–4Γ— latency) vs VLA." IROS 2026 answers by dissolving the versus β€” the WAM survives as critic / distillation-teacher / shared-representation / MPC-model, i.e. a training-time or representational component, while the deployed controller stays a reactive VLA/policy. So the honest read is not "WAM beats VLA" but "the WAM is being absorbed as scaffolding around a reactive VLA" β€” the in-path generative WAM (DreamZero/Cosmos-Policy) is now the minority position, and the latency objection (Review-WAM-vs-VLA-Robustness) is sidestepped by construction.

5.2 Memory & In-Context Learning β€” a first-class IROS track

IROS 2026 gives memory its own focused session ("Building Memory into Vision-Language Robot Agents") plus an in-context-imitation thread β€” the applied face of VLA Memory, RoboMME, and In-Context Imitation:

  • Long-horizon memory representations: GaussMemory β€” task-driven 3D Gaussian scene memory; PROMPT β€” probabilistic memory trees for efficient skill reuse; TempoFit πŸ“„ β€” plug-and-play layer-wise temporal KV memory for long-horizon VLA (the KV-cache-as-memory route, cf. RoboTTT).
  • In-context imitation (watch-a-demo, no per-task FT): RoboSSM πŸ“„ β€” scalable in-context imitation via state-space models (linear-cost long context); ICLR (USC) πŸ“„ β€” in-context imitation with visual reasoning traces (jointly generate anticipated image-space trajectories + actions), improving unseen-task generalization. (full-paper analysis)
  • Demonstration efficiency: IMLE-VLA πŸ“„ β€” single-step action generation (55 Hz, LIBERO 98.0%), ESPADA β€” semantics-aware demo downsampling (~2Γ— speedup), "What Can Robot FMs Learn from Limited Human Demonstrations?" (Toyota) β€” human-task-prior fine-tuning.

Insight: the memory field is bifurcating into explicit structured memory (3D-Gaussian scene, memory trees) vs parametric/context memory (KV-cache, state-space, fast-weights) β€” exactly the In-Context Imitation Β§2 taxonomy, now with hardware-facing instances. State-space models (RoboSSM) and KV-memory (TempoFit) are the efficiency answer to the long-context latency wall that RoboTTT attacks with fast weights.

5.3 Bimanual & Humanoid manipulation β€” whole-body + two-hand coordination goes mainstream

Two-arm and humanoid manipulation are a large, distinct cluster:

  • Bimanual policy & coordination: EquiBim πŸ“„ β€” symmetry-equivariant bimanual policy; 3D FlowMatch Actor πŸ“„ (CMU/NVIDIA) β€” one 3D policy for single and dual-arm, +41.4% on PerAct2, ~30Γ— faster; Unified Learning of Temporal Task Structure and Action Timing for bimanual; Grasp, Handover, Rotate β€” bimanual reorientation via compositional diffusion + energy-based models; HapticVLA β€” contact-rich bimanual VLA that needs no inference-time tactile. (EquiBim & 3DFA have full-paper analyses)
  • Bimanual planning / language: Grounded Vision-Language Interpreter for Long-Horizon Bimanual TAMP; Prompt-To-Product β€” generative assembly via bimanual manipulation.
  • Bimanual data & retargeting: Scalable Multi-Task Data Generation via RL for Language-Conditioned Bimanual Dexterous manipulation; SPIDER β€” scalable physics-informed dexterous retargeting (cf. Dexterous-Hand Data Pyramid L4).
  • Humanoid whole-body loco-manipulation: ULTRA β€” unified multimodal whole-body control; CEER β€” compliant end-effector + root as a unified hierarchical interface; OmniDP β€” beyond-FOV large-workspace humanoid manipulation via omnidirectional 3D perception; AGILE (workflow), SteadyTray (residual-RL balancing); Responsibility-Induced Specialized Experts β€” a MoE-VLA for humanoid loco-manipulation (cf. Multi-Task VLA, HiMoE-VLA).

Insight: the frontier is coordination structure, not just adding a second arm β€” symmetry/equivariance (EquiBim), unified single/dual-arm policies (3D FlowMatch), explicit temporal-timing models, and whole-body interfaces that fold locomotion + two-hand manipulation into one controller (ULTRA/CEER). Data is the bottleneck β€” hence RL-based bimanual data generation and physics-informed retargeting (SPIDER). This is the hardware-grounded counterpart to Ξ¨β‚€/Humanoid VLA, and MoE-for-humanoid-VLA shows the multi-task interference fixes reaching whole-body control.


6. Caveats & how to extend this page

  • Titles + authors + abstracts are program-verified (the program's local search exposes them); final camera-ready versions land on IEEE Xplore at/after the conference. Numbers like #359 are program paper IDs.
  • Coverage: Β§3 and Β§5 sample the most VLA/manipulation-central sessions and the three requested lenses (WAM, memory/in-context, bimanual/humanoid); the full in-scope set is ~40 sessions (Β§2). This page can be upgraded to a full per-paper index (like RSS 2026 Papers) with abstract summaries per paper.
  • Intermediates saved during extraction: iros26/sessions.md (scratchpad).

7. Links

← Back to Reviews Β· Home