IROS 2025 VLA Manipulation Survey - Heungwoo/research GitHub Wiki
Compiled April 2026 (in retrospect). Focus: IROS 2025 VLA / manipulation papers, grouped by approach, with a CoRL 2025 β IROS 2025 β NeurIPS 2025 β ICLR 2026 positioning map. IROS is an IEEE systems venue β expect more real-robot deployment, hardware co-design, and systems-integration papers than at ML venues.
IROS 2025 (Hangzhou, Oct 19β25) accepted 1,991 papers (β46.2% from 4,306 submissions, a new record), plus 777 RA-L / T-RO / T-ASE / T-MECH journal transfers. The conference hosted 5 parallel "Award Finalists" sessions (~45 papers total across all category tracks) and a dense workshop program on Generative AI for Robotics, Tactile Sensing, AI-Meets-Autonomy, Fine Manipulation, and Embodied AI for Science.
Five distinctive trends compared to the ML venues (CoRL/NeurIPS/ICLR):
- VLAs become vertical-domain products β RoboNurse-VLA (surgical scrub nurse), CapsDT (capsule endoscopy), VLIN-RL (task planning), RoboDexVLM (dexterous manipulator task planning). IROS is the venue where VLA meets a real application domain with hardware.
- Real-robot systems integration, not architecture theory β HACTS (bilateral teleoperation hardware), Bunny-VisionPro (Vision Pro bimanual), MoMa-Teleop (zero-added-cost whole-body teleop), iDP3-humanoid (25-DoF cart-mounted humanoid). Every paper has a hardware demo.
- VLA acceleration and deployment β PD-VLA parallel decoding (2.52Γ faster), TinyVLA (<1B params), ReBot (real-to-sim-to-real for VLA adaptation), Refined Policy Distillation (VLA β compact RL expert). The "how do we actually ship this" papers.
- Dataset & platform papers at industrial scale β AgiBot World Colosseo (1M+ trajectories, 217 tasks, 100 robots, 5 domains) β Best Paper Finalist and the largest public manipulation dataset to date.
- Tactile-first manipulation is a whole track β VTAO-BiManip (visuo-tactile-action pretraining), Tactile-VLA workshop best paper, Adaptive Visuo-Tactile Fusion, RoTipBot. IROS is where tactile research concentrates.
IROS 2025 is the deployment-and-hardware companion to the ML-venue trio of CoRL 2025 / NeurIPS 2025 / ICLR 2026. Most IROS 2025 papers had arXiv preprints from late 2024 / early 2025 and target RA-L acceptance (presented at IROS by policy); the ML venues rarely cite them directly, but the hardware interfaces (HACTS, Bunny-VisionPro, iDP3-humanoid) and datasets (AgiBot World) are the substrate that ML work is built on.
flowchart LR
Site[4000 mΒ² 5-domain facility<br/>domestic Β· retail Β· industrial Β· restaurant Β· office] --> HW[100 real robots<br/>bimanual + dexterous]
HW --> Tele[Standardized teleop<br/>+ human-in-the-loop QC]
Tele --> Data[1M+ trajectories<br/>217 tasks]
Data --> GO1[Genie Operator-1<br/>generalist policy w/ latent actions]
GO1 --> Scale[Predictable scaling<br/>with data volume]
AgiBot World Colosseo (OpenDriveLab + AgiBot, Best Paper Award Finalist, arXiv 2503.06669 β RA-L β IROS 2025, also TRO 2026) is the IROS 2025 flagship for this wiki's purposes. It dwarfs every previous manipulation dataset (Open X-Embodiment, DROID, BridgeData V2) by an order of magnitude β 1M+ trajectories from 100 real bimanual robots across 217 real-world tasks in 5 deployment domains. The paper introduces Genie Operator-1 (GO-1), a latent-action generalist policy, and demonstrates predictable performance scaling with data β policies pre-trained on AgiBot World average +30% over Open X-Embodiment pretraining on the same downstream tasks. Downstream, AgiBot World is the reference training corpus for several 2026 papers on scaling laws and cross-embodiment generalization.
- VLA Architecture (IROS-specific spins)
- Real-Robot Deployment & Systems
- Dexterous Manipulation
- Mobile Manipulation
- Humanoid & Whole-Body
- Data & Benchmarks
- Teleoperation & Imitation-Data Interfaces
- Awards & Recognition
- Workshops
- Trends Specific to IROS
- Lineage to Other Venues
- Reading List
IROS VLA papers tend to target either a vertical application (medical, surgical, capsule) or a deployment optimization (acceleration, compression, RL distillation). Architectural novelty is lighter than at NeurIPS/ICLR, but real-robot results are heavier.
| Paper | Primary Affiliation | arXiv | Novelty |
|---|---|---|---|
| RoboNurse-VLA | CUHK MSMRC | 2409.19590 | SAM 2 + Llama 2 for surgical-instrument grasping/handover in a scrub-nurse system; voice-driven, real-time, robust to unseen tools. First VLA deployed end-to-end in a mock-OR setting. |
| CapsDT | CUHK (Ren Lab) | β | Diffusion-Transformer VLA for magnetically-actuated capsule endoscopy in the stomach; 26.25% real-world success on 4 endoscopy task levels. First VLA for capsule robots. |
| TinyVLA | Midea + ECNU | 2409.12514 | Compact <1B-param VLA (MLM backbone + diffusion head); no pretraining stage, 5.5Γ fewer params, +25.7% real success over OpenVLA. RA-L β IROS. |
| PD-VLA | HKUST(GZ) + PolyU | 2503.02310 | Reformulates AR VLA decoding as a parallel fixed-point iteration β training-free 2.52Γ inference speedup, works with action chunking. |
| CLAP | South China Univ. Tech. | β | Closed-loop diffusion-transformer action foundation model; critic module at inference forms a closed loop around a VLA base. |
| ReBot | UNC Chapel Hill + UT Austin | 2503.14526 | Real-to-sim-to-real video synthesis for VLA adaptation: replay real trajectories in sim to diversify objects, inpaint real backgrounds. +20% Franka real-world for OpenVLA. |
| Refined Policy Distillation (RPD) | Univ. Tech. Nuremberg + Freiburg | 2503.05833 | Distills a generalist VLA (OpenVLA / Octo) into a compact RL expert via on-policy RL + BC; student beats teacher, robust to camera changes. |
Where IROS really distinguishes itself from the ML venues: every paper targets a working hardware demo, typically with a video and a project page rather than a headline benchmark number.
| Paper | Primary Affiliation | arXiv | Systems Contribution |
|---|---|---|---|
| HACTS | BIT + Midea | 2503.24070 | Bilateral leader-follower teleop with 3D-printed low-cost hardware; enables human-as-copilot HITL-RL on VLAs. |
| MoMa-Teleop | Freiburg | 2409.15095 | Zero-added-cost whole-body teleop for mobile manipulators: existing interfaces (joystick, hand-guidance) send end-effector commands; RL policy handles the base. 5-demo imitation transfer. |
| Bunny-VisionPro | UCSD (Wang) + HKU | 2407.03162 | Real-time bimanual dexterous teleop via Apple Vision Pro + custom haptic devices; +11% success, -45% completion time vs Telekinesis. Collision + singularity avoidance built in. |
| iDP3 Humanoid | Stanford (Wu) | 2410.10803 | Whole-upper-body teleop + 25-DoF humanoid platform (height-adjustable cart + 3D LiDAR) + improved 3D diffusion policy. Single-scene data β generalizes to novel real-world scenes; 2000+ real rollouts. |
IROS dexterous track is enormous (3 "Dexterous Manipulation" sessions + grasping sessions). Representative standouts with learning angles:
| Paper | Primary Affiliation | arXiv | Novelty |
|---|---|---|---|
| VTAO-BiManip | Zhejiang Univ. + ChingMu | 2501.03606 | Masked Visual-Tactile-Action pretraining (MAE-style) with object-state prediction; +20% over prior visuo-tactile pretraining on bottle-cap unscrewing. |
| LDexMM | SJTU (Lu) | β (IEEE Xplore) | Two-phase: language β segmentation β functional grasp pose; RL refinement with language-object-grounding constraints. 31β73% success on 3β10 tasks with a single policy. |
| RoboDexVLM | HKUST(GZ) | 2503.01616 | VLM-based task planner with task-level recovery + language-guided dexterous grasp perception for long-horizon dexterous sequences. |
| DexDiffuser | KTH (Kragic) | 2402.02989 | Conditional diffusion grasp sampler + learned evaluator for dexterous multi-finger grasps; +9β19% over FFHNet. RA-L β IROS. |
| FoundationGrasp | β | 2404.10399 | LLM-generated semantic + VLM-generated geometric descriptions fused by a task-oriented grasp evaluator. LaViA-TaskGrasp dataset. |
| Adaptive Visuo-Tactile Fusion with Predictive Force Attention | β | β | Predictive force attention fuses vision + tactile for contact-rich dexterous manipulation. Dexterous-Manip-2 session. |
Mobile manipulation is where IROS 2025 has the clearest hardware-first identity relative to ML venues.
| Paper | Primary Affiliation | arXiv | Novelty |
|---|---|---|---|
| GeT-USE | Stanford (Fei-Fei) + UT Austin (MartΓn-MartΓn) | 2510.25754 | Learns simulated embodiment extensions (hypothetical end-effectors) to identify useful tool geometries, then transfers to real-world tool selection. 22-DoF bimanual mobile robot, +30β60% on 3 tasks. |
| MORE | Freiburg | 2505.03035 | Scene-graph + active-filtering LLM planner for zero-shot rearrangement; first approach to solve a significant fraction of BEHAVIOR-1K. |
| Env-Mani | β | β | Quadrupedal loco-manipulation with environment-in-the-loop β uses environment geometry as compliant fixture. |
| AC-DiT (cross-venue) | PKU | β | Mobile-to-body conditioning + 2D/3D perception for mobile manipulation diffusion policies (see NeurIPS 2025 listing). |
| Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller | HKUST(GZ) + HITSZ | β | Best Paper Award Finalist. Learned arm-as-pusher controller enables legged robots to clear doors/curtains during navigation. |
| Paper | Primary Affiliation | arXiv | Novelty |
|---|---|---|---|
| iDP3 Humanoid (flagship) | Stanford | 2410.10803 | See Deployment section. 25-DoF full-sized humanoid. |
| Efficient Learning of a Unified Policy for Whole-Body Manipulation and Locomotion Skills | β | 2507.04229 | Best Paper Award Finalist. Explicit kinematic model of manipulator embedded in RL reward; zero-shot transfer to Aliengo/X20 + Unitree Z1. |
| HumanRobot Intrinsic Skill Transfer and Programming by Demonstration System (I) | β | β | Invited journal paper: skill transfer from human demonstration to humanoid. |
| Learning Accurate Whole-Body Throwing with High-Frequency Residual Policy and Pullback Tube Acceleration | β | β | Whole-body throwing policy for bipedal/humanoid robots with residual correction. |
| HiFAR | β | β | Multi-stage curriculum for high-dynamics humanoid fall recovery. |
| SHIELD | Caltech / UT Austin | β | Safety on humanoids via CBFs in expectation on learned dynamics. Robot Safety 1 session. |
| Paper | Primary Affiliation | arXiv | Scale / Focus |
|---|---|---|---|
| AgiBot World Colosseo (flagship) | AgiBot + OpenDriveLab | 2503.06669 | Best Paper Award Finalist (also IEEE T-RO 2026). 1M+ trajectories, 217 tasks, 100+ bimanual robots, 5 domains. GO-1 latent-action generalist policy; +30% over Open X-Embodiment pretraining. |
| Ξ» (Lambda) | β | β | Benchmark for data-efficiency in long-horizon indoor mobile manipulation. |
| Benchmarking Long-Horizon Mobile Manipulation in Multi-Room Dynamic Environments | Tsinghua (Ma) | β | Multi-room + dynamic-object benchmark. |
| DG16M | β | β | 16M-grasp dataset for dual-arm grasping with force-optimized labels. |
| BookBot | β | β | Voice-driven cluttered-book grasping benchmark. |
IROS is the home venue for teleoperation hardware papers β the "how to actually collect the data" layer that ML-venue VLAs assume as free infrastructure.
| Paper | Interface | arXiv |
|---|---|---|
| HACTS | Bilateral 3D-printed leader-follower; HITL-RL on VLAs | 2503.24070 |
| Bunny-VisionPro | Apple Vision Pro + custom haptics, bimanual dex | 2407.03162 |
| MoMa-Teleop | Zero-added-cost whole-body; joystick or hand-guidance + RL base | 2409.15095 |
| iDP3 whole-upper-body teleop | Custom humanoid cart; enables iDP3 training data | 2410.10803 |
| Haptic-ACT | Immersive VR with compliant-control feedback | β |
| DARt Vinci | Egocentric surgical-robot data collection | β |
| A Hybrid Mapping Method: Balancing Efficiency and Intuitiveness in Lateral Teleoperation | Lateral-motion teleop mapping (Award Finalist 5) | β |
IROS has ~40 main-track award categories plus workshop awards. For 2025:
| Award | Paper | Affiliation |
|---|---|---|
| Best Conference Paper | Interdigitated Electrodes for Selective Stimulation of Skeletal Muscle Actuators in Biosyncretic Robots | Shenyang Inst. of Automation (CAS) |
| Best Student Paper | Neural MP: A Generalist Neural Motion Planner (Dalal et al.) | CMU β arXiv 2409.05864 |
| Best Paper on Cognitive Robotics | (category exists) | β |
| Best Paper on Robot Mechanisms & Design (finalist) | TFRR: Tensegrity-Based Fracture Reduction Robot | PRISMA Lab, UniNA |
Across the 5 "Award Finalists" sessions, the VLA/manipulation-relevant finalists include:
- Neural MP (CMU β won Best Student Paper) β reactive generalist motion planner that distills expert MP data into a policy; used downstream in ManipGen.
- Interactive Navigation for Legged Manipulators with Learned Arm-Pushing Controller (HKUST(GZ) + HITSZ).
- AgiBot World Colosseo (AgiBot + OpenDriveLab) β 1M+ trajectories.
- Efficient Learning of a Unified Policy for Whole-Body Manipulation and Locomotion Skills (Hou et al.).
- Octopi-X: Large Tactile-Vision-Language Model for Physical Property Inference β Tactile Sensing workshop Best Paper Finalist.
- RoTipBot β Tactile workshop Best Paper.
(IROS does not publish a single consolidated Outstanding Paper list the way NeurIPS/ICLR do β awards are distributed across ~40 category-specific tracks. See ieee-ras.org conference awards.)
Relevant VLA / manipulation / embodied workshops at IROS 2025:
| Workshop | Focus |
|---|---|
| Generative AI for Robotics and Smart Manufacturing | ECoT reasoning in VLAs; generative data for robotics. |
| AI Meets Autonomy (CMU VLA Challenge) | Generalist embodied agents, VLN, world models, foundation models + the CMU VLA Challenge results. |
| Robotic Fine Manipulation: Tactile + Visual + Intelligent Control | Sergey Levine keynote. RoTipBot (Best Paper), Octopi-X, InvariantCloud (finalists). Tactile-VLA paper (2507.09160) featured. |
| AI Robot for Science | Embodied AI for autonomous scientific research. |
| Robotic Data Generation and Evaluation (RoDGE) | Synthetic + collected data pipelines for robot learning. |
| SASA Teleoperation | Teleoperation interfaces and shared autonomy. |
| Active Perception | activep-ws.github.io. |
| PPNIV | iros25-ppniv.github.io. |
What IROS 2025 VLA/manipulation looks like compared to NeurIPS 2025 / ICLR 2026:
- Hardware-first framing. Every paper has a robot, not just a benchmark. Systems contributions (teleop hardware, sensor integration, control-stack deployment) get comparable credit to ML novelty. Contrast NeurIPS 2025 where the flagship is Knowledge Insulation β a training recipe with no new hardware.
- VLAs are products in a vertical. Surgical (RoboNurse-VLA), capsule-endoscopic (CapsDT), agricultural, industrial. NeurIPS/ICLR VLA papers optimize generality; IROS VLA papers optimize a domain.
- Acceleration, compression, distillation dominate the VLA track. TinyVLA, PD-VLA, RPD, ReBot, Fast Policy. All target the gap between "a 7B VLA exists in a paper" and "a 7B VLA runs on the robot." Contrast CoRL (architectural novelty) and ICLR (RL / world-model theory).
- Tactile is a first-class research track. A dedicated Tactile Sensing workshop with its own best paper awards; dedicated Dexterous Manipulation sessions with tactile emphasis (VTAO-BiManip, Adaptive Visuo-Tactile Fusion, RoTipBot). NeurIPS/ICLR have 0 or 1 tactile papers.
- Industrial and medical robotics as a continuous spectrum. Medical Robots & Systems sessions 1β7 are ~70 papers; many integrate VLMs/VLAs. This volume does not exist at ML venues.
- Papers appear as "arXiv 2024/early-2025 β RA-L β presented at IROS 2025." The RA-L journal-to-IROS pipeline means many IROS 2025 papers pre-date the ML venue submissions that cite them. This explains why, chronologically, IROS 2025 (Oct) can host papers whose ideas already appear in NeurIPS 2025 (Dec) and ICLR 2026 (Apr) citations β the RA-L version was already public.
IROS 2025 (Oct) sits between CoRL 2025 (Sept) and NeurIPS 2025 (Dec). Because IROS accepts primarily via RA-L, most papers existed as arXiv 2024 / early-2025 preprints long before the physical conference. Direct lineage:
| CoRL 2025 / pre-IROS seed | IROS 2025 descendant |
|---|---|
| DexUMI (dex-hand data interface) | Bunny-VisionPro + HACTS β bimanual data-collection hardware in the same design-pattern lineage |
| Streaming Flow Policy (low-latency action streaming) | PD-VLA β parallel decoding is the IROS systems-engineering analog |
| Visual Imitation β Humanoid | iDP3 Humanoid (IROS 2025) β same 3D-diffusion-policy backbone, different data-collection approach |
IROS 2025 papers are rarely cited as IROS papers by ML venues β they are usually cited by their arXiv IDs from months earlier. But the ideas thread forward:
| IROS 2025 idea | ML-venue successor |
|---|---|
| TinyVLA (small-efficient VLA) | β SmolVLA + FLOWER + MiniVLA family at ICLR 2026 |
| PD-VLA (parallel AR decoding) | β FASTER, Discrete Diffusion VLA, DIVA & Fast-dVLA, Spec-VLA (speculative decoding) |
| ReBot (real-to-sim-to-real for VLA) | β X-Sim / DexFlyWheel (NeurIPS 2025) lineage |
| Refined Policy Distillation (VLA β RL expert) | β SimpleVLA-RL, VLA-RFT, Stage-Aware RL |
| AgiBot World Colosseo (1M-traj dataset) | β RoboCasa365, RoboArena β, EgoDex; substrate for X-VLA / UniVLA / GO-1 descendants |
| HACTS / MoMa-Teleop / Bunny-VisionPro (teleop hardware) | β PLD, Hybrid Training β HITL-RL + residual-correction interfaces at ICLR 2026 |
| Tactile-VLA (workshop paper) | β No ICLR 2026 1:1 mirror; likely surfaces at ICLR 2027 / RSS 2026 |
| iDP3 + whole-upper-body teleop | β WholeBodyVLA, UniHM humanoid VLA stack |
No single IROS 2025 paper has the citation density of Knowledge Insulation (NeurIPS Spotlight) or Ο0.5 (CoRL Oral), but three are genuine hubs in the ICLR 2026 bibliography:
- AgiBot World Colosseo β cited by ~every ICLR 2026 scaling/cross-embodiment paper as "largest public manipulation dataset."
- Bunny-VisionPro β cited as the canonical Vision Pro bimanual teleop reference.
- Neural MP (Best Student Paper) β cited by long-horizon manipulation papers (e.g., ManipGen, which uses Neural MP as its low-level planner).
| Priority | Papers |
|---|---|
| Must read | AgiBot World Colosseo (2503.06669), Neural MP (2409.05864 β Best Student), iDP3 Humanoid (2410.10803) |
| VLA architecture | TinyVLA (2409.12514), PD-VLA (2503.02310), RoboNurse-VLA (2409.19590), CapsDT |
| Deployment / systems | ReBot (2503.14526), RPD (2503.05833), HACTS (2503.24070) |
| Dexterous | VTAO-BiManip (2501.03606), RoboDexVLM (2503.01616), LDexMM, DexDiffuser (2402.02989) |
| Mobile manipulation | GeT-USE (2510.25754), MORE (2505.03035), MoMa-Teleop (2409.15095) |
| Humanoid | iDP3 (2410.10803), Unified Whole-Body Policy (2507.04229) |
| Teleop hardware | Bunny-VisionPro (2407.03162), HACTS (2503.24070), MoMa-Teleop (2409.15095) |
| Workshop-adjacent | Tactile-VLA (2507.09160), ManiGaussian++ (2506.19842) |
- IROS 2025 official site Β· Conference Digest PDF
- IROS 2025 submission stats (Twitter/X) β 4,306 submissions
- Paper list (DoongLi mirror) β full accepted-paper CSV with abstracts
- Paper Copilot IROS statistics
- Best Conference Paper announcement (SIA / CAS)
- Neural MP Best Student Paper (Deepak Pathak on X) Β· project
- AgiBot World (OpenDriveLab GitHub)
- IEEE RAS conference awards
Per-paper arXiv links appear inline in the tables above. Some RA-L / IEEE Xplore papers (LDexMM, CLAP, CapsDT, some dex-manip / mobile-manip finalists) do not have a public arXiv preprint β DOIs are on the IEEE Xplore listings linked from the paper list.
- IROS vs RA-L timing. The "RA-L β presented at IROS" policy means roughly half of IROS 2025 papers had arXiv preprints in 2024 / early 2025, so they chronologically precede CoRL 2025 / NeurIPS 2025 ideas they share lineage with. Date-based lineage claims need care.
- Harmonic Mobile Manipulation (Yang et al., UCSD/AI2) is the IROS 2024 Best Paper on Mobile Manipulation, often cited alongside IROS 2025 due to ongoing follow-up work. Not an IROS 2025 paper.
- Neural MP arXiv (2409.05864) is from Sept 2024; won Best Student Paper at IROS 2025 (presented Oct 2025). This is normal for the RA-L pipeline.
- Tactile-VLA (2507.09160) was a workshop paper at the Tactile-Manipulation workshop, not a main-track IROS paper.
- AgiBot World is listed as a Best Paper Finalist in official announcements and as T-RO 2026 in GitHub metadata β both refer to the same work; the conference track at IROS 2025 is the short form.