ICRA 2026 Topic VLA - Heungwoo/research GitHub Wiki

ICRA 2026 β€” Vision-Language-Action Models (Topic Analysis)

This page analyzes the 45 papers at ICRA 2026 that carry "Vision-Language-Action" / "VLA" in the title or are core VLA contributions. The character of the ICRA cohort is distinct from the ML-venue VLA wave (ICLR 2026, CVPR 2026): it is deployment- and sensor-centric. Where CVPR/ICLR optimize architecture and reasoning benchmarks, ICRA's VLA work overwhelmingly asks how the model meets the robot — adding force/tactile/audio/depth to the VLM→action stack (usually without new hardware), pruning and streaming tokens for edge latency, making continuous-action (flow/diffusion) policies improvable from reward, and stress-testing practicality, robustness, and security on real platforms. Notable cross-cutting signals: (1) sensor grounding is the single largest cluster — vision-only VLA is now treated as insufficient for contact-rich work; (2) flow-matching policies have become the default continuous-action head, and several papers attack their two open problems (RL fine-tuning, action-coherence/jitter); (3) inference-time methods (MCTS search, introspection, affordance learning, memory prompting) are proliferating because they upgrade frozen pretrained VLAs cheaply; and (4) navigation VLA has split into its own coherent sub-field with multi-modal goal conditioning.

Sub-trends

1. Sensor-grounded VLA (force Β· tactile Β· audio Β· multi-sensor)

The largest theme. The shared premise is that RGB-only VLMs lack the contact and dynamic-process signals needed for precise, contact-rich manipulation, and the shared trick is to inject the missing modality without demanding new sensors at deployment. FD-VLA distills a force token from vision+state into the VLM so a force/torque sensor is needed only at training time; CRAFT instead adapts a VLA to contact tasks via a force-aware curriculum fine-tuning schedule; and Enhancing VLA Precision via FiLM-based Force/Torque-Vision Integration fuses F/T into the visual stream through FiLM conditioning for insertion-grade alignment. Audio-VLA adds contact-audio perception (AudioCLIP encoder, LoRA-fine-tuned Llama2 backbone) to perceive collision/contact events that vision misses, and introduces a Task-Completion-Rate metric for dynamic processes. The manipulation-side OmniVLA (arXiv 2511.01210 β€” note the name collision with the navigation OmniVLA in Β§7) generalizes this to unified multi-sensor perception, fusing IR/radar/audio into a physically-grounded policy. The pattern across all five: a modality-specific encoder feeds tokens or FiLM parameters into an otherwise-standard VLA, and the gains concentrate on alignment/insertion/deformable tasks where vision alone plateaus.

2. Spatial / 3D / depth grounding

A cluster attacking VLMs' weak 3D spatial reasoning, which it inherits from 2D-pretrained backbones. DepthVLA adds a pretrained depth transformer in a mixture-of-transformers design (VLM + depth expert + action expert, fully shared attention) and reports large real-world and Simpler-simulator gains. AugVLA-3D does depth-driven feature augmentation rather than adding a stream. InSpire takes a language route: it prepends a spatial-reasoning VQA ("in which direction is the [object] relative to the robot?") to redirect attention to task-relevant factors and break spurious visual correlations, needing no auxiliary data. RetoVLA reuses otherwise-discarded ViT register tokens as a free spatial-reasoning signal, framing it as compression-friendly. Seeing Space and Motion improves Latent Action Models by injecting geometric and dynamic awareness into the latent action encoder. The common insight: spatial competence can be added through an auxiliary expert, an auxiliary language task, or recovered from latent/register representations already present β€” and all three beat scaling action data alone.

3. Reasoning, verification & introspection (inference-time)

These methods upgrade frozen pretrained VLAs at inference time. VLA-Reasoner runs online Monte-Carlo Tree Search over candidate action chunks, scoring each branch by rolling it out through a 600M action-aware world model β€” treating imagined outcomes as "rationales." Do What You Say addresses a subtler failure: a reasoning VLA can emit a correct textual plan yet a misaligned low-level action, so it adds runtime reasoning-action alignment verification to steer execution back to the stated intent. INSIGHT uses token-level uncertainty as an introspective help-trigger signal so the policy can ask a human supervisor before failing. Learning Affordances at Inference-Time lets the VLA reflect on failed tries and revise affordances mid-deployment. CollabVLA turns a visuomotor policy into a self-reflective collaborator that "dreams together" with a human, targeting interpretability and latency. The unifying move: rather than retrain, add a search / verification / uncertainty / reflection loop on top β€” a cheap path to reliability that the ICRA cohort clearly favors.

4. Efficiency & deployment (pruning Β· streaming Β· edge Β· practicality)

ICRA's signature concern. LightVLA is a parameter-free, hyperparameter-free differentiable token pruner that cuts FLOPs/latency while raising success rate (it inverts the usual efficiency/accuracy trade-off). RetoVLA (also Β§2) doubles as a compression method. Stream-To-Act is a ROS 2-native token-streaming runtime that emits actions continuously for real-time control loops instead of blocking on full-sequence decoding. Robust Unknown Object Detection and Tracking on Edge Devices proposes a stepwise VLA for Jetson AGX Orin, sidestepping the memory blow-up of end-to-end VLA on edge hardware. EveryDayVLA pushes the hardware floor down with a ~$300 6-DOF manipulator and a matching policy. VLA Practicality ("Rethinking the Practicality of VLA") is the cluster's benchmark/baseline statement, arguing existing VLAs are over-parameterized and pre-training-hungry, and offering a leaner baseline. Together they form a coherent "make VLA actually deployable" agenda spanning compute, latency, hardware cost, and evaluation.

5. RL & preference fine-tuning of continuous-action policies

A focused, technically deep cluster around the open problem of improving flow-matching/diffusion VLAs from reward (their likelihoods are intractable, so standard RL/DPO does not transfer cleanly). FPO / Reinforcement Fine-Tuning of Flow-Matching Policies derives a flow-compatible policy-optimization objective with stable convergence under sparse reward, evaluated on LIBERO and ALOHA. Offline Reinforced Finetuning for Chunk-Based VLA distills a real-world RL policy into a chunk-based VLA via a vision-guided copilot, keeping the action-chunking interface. Toward Human Preference Optimization is a pilot study probing the limits of imitation learning on multi-step tasks (evaluating GR00T N1.6) and motivating preference-based signals. ACG (Action Coherence Guidance) targets a flow-specific pathology β€” high generative capacity makes flow policies jittery/noise-sensitive under imitation β€” with a guidance term enforcing temporal action coherence. The cluster reflects that flow-matching is now the dominant action head, and ICRA is where its reward-driven and stability fixes are being worked out.

6. Memory, continual learning & specialization

Pretrained VLAs generalize broadly but forget, lack long-horizon memory, and underperform on the specific deployment they are sent to. MAP-VLA adds memory-augmented prompting so a frozen VLA can carry context across a long-horizon task instead of relying only on the current frame. ExpReS-VLA specializes a generalist (e.g. OpenVLA) to a fixed task set via experience replay and retrieval, trading broad zero-shot for consistent in-distribution performance. Learning Affordances at Inference-Time (also Β§3) is continual-learning-flavored, accumulating affordance corrections across attempts. The shared diagnosis: deployment values reliable specialization with memory over broad-but-shallow generalization, and lightweight prompting/retrieval beats full retraining.

7. Navigation, tracking & world-model / goal conditioning

Navigation VLA is now its own sub-field. OmniVLA (nav) (UC Berkeley; ~9,500 hours across 10 platforms) is an omni-modal navigation policy that composes language, spatial-coordinate, and visual-reference goals in one model. UrbanVLA is a route-conditioned model for urban micromobility/delivery, aligning noisy map routes with egocentric vision and trained SFT-then-RFT (reports >55% improvement over baselines on MetaUrban; real >500m routes on a Unitree Go2). TrackVLA++ brings reasoning (a Polar-CoT spatial token) and a Target-Identification Memory to embodied visual tracking, for robustness under occlusion and distractors. On the manipulation side, world-model/goal conditioning appears in Goal-VLA, which uses image-generative VLMs as object-centric world models for zero-shot manipulation via reflection-through-synthesis, and in DAM-VLA, a dynamic action model bridging gross motion and precise manipulation in dynamic scenes. The common thread is goal flexibility: conditioning on coordinates, images, routes, or synthesized goal states rather than language alone.

8. Data, pretraining, platforms & dexterity

The supply-side and embodiment cluster. Galaxea / G0 pairs a 500-hour, 100K-trajectory single-embodiment open-world dataset with a dual-system VLA (the System-2 G0-VLM hits 83.3% on Table Bussing vs 32.0% for Gemini-2.5-pro). Dexora is a fully open-source stack (hardware + data + policy) for high-DoF (36-DoF) bimanual dexterity, a regime usually locked behind proprietary hardware. Two papers pretrain from human video instead of teleoperation β€” Developing VLA from Egocentric Videos and Scalable VLA Pretraining with Real-Life Human Activity Videos (treating the human hand as the end-effector) β€” a scalable-data strategy. Open-World Object Manipulation via Synthetic Multi-Modal Data generates synthetic data for object generalization. RealMirror is an open platform (3D-Gaussian-Splatting reconstruction + sim-to-real) for end-to-end VLA research without a real robot. DexGrasp AI Copilot learns end-to-end dexterous arm-hand policies with shared-autonomy teleoperation. The cluster's bet: scalable non-teleop data and open embodiments are the path past the data bottleneck.

9. Generalization, robustness & security (cross-cutting)

A smaller but important set. Exploiting Vulnerabilities demonstrates universal adversarial attacks on VLA models in robotics β€” a rare security-focused contribution showing physical-world VLAs are attackable. SVP (Dual Stochastic Visual Prompting) diagnoses "distracted attention" as shortcut learning in models like OpenVLA and fixes it with stochastic visual prompts rather than architecture changes. Toward Embodiment Equivariant VLA Policy attacks cross-embodiment generalization via equivariance rather than scale. Hierarchical LLM-VLA-Controller Integration layers an LLM planner over a VLA to combat memorization-over-semantics. Domain-specific deployments round out the cohort: ultrasound-guided needle insertion, NeuroVLA (endoscopic neurosurgery debulking), and TMR-VLA (magnetic control of a tri-leg silicone soft robot) β€” three medical/soft-robot VLAs proving the paradigm transfers well beyond tabletop grippers.

Standout deep-dives

  • VLA-Reasoner (arXiv 2509.22643) β€” Online MCTS over action chunks, scored by rolling out a 600M action-aware world model so imagined outcomes serve as rationales. Improves frozen policies on LIBERO: OpenVLA-SFT 76.0%β†’81.0% (+5.0 pp), Octo-Small 26.5%β†’37.3% (+10.8 pp), SpatialVLA 34.0%β†’41.8% (+7.8 pp). A clean demonstration that test-time search + a learned world model upgrades any pretrained VLA without retraining.

  • LightVLA (arXiv 2509.12594) β€” Differentiable, parameter-free and hyperparameter-free token pruning that retains ~78 visual tokens on average. Reports βˆ’59.1% FLOPs and βˆ’38.2% latency while raising success rate to 97.4% avg vs OpenVLA-OFT 94.5% (+2.9 pp), and far above VLM-oriented pruners (FlashVLA 73.7%, SP-VLA 74.9%, VLA-Cache 74.7%) that collapse on VLA tasks. The headline is that performance-driven pruning breaks the usual efficiency/accuracy trade-off.

  • Goal-VLA (arXiv 2506.23919) β€” Uses image-generative VLMs as object-centric world models for zero-shot manipulation. On RLBench (8 tasks, 100 runs each) averages 59.9% vs MOKA 26.0%, MolmoAct 11.3%, Ο€0 0.0%; real-world (4 tasks) 60% vs MolmoAct 27.5%. Ablation: a 40.0% baseline rises to 83.8% with input enhancement + reflector and 88.8% with three reflection iterations, quantifying the reflection-through-synthesis loop.

  • FPO β€” RL Fine-Tuning of Flow-Matching Policies (arXiv 2510.09976) β€” A flow-compatible policy-optimization objective (CFM-based) that makes flow-matching VLAs improvable from sparse reward where standard RL/DPO does not transfer. LIBERO success: Spatial 97.2 Β· Object 97.3 Β· Goal 89.4 Β· Long 65.3 Β· avg 87.2; ALOHA Transfer Cube learning curve reaches ~65%. Directly addresses the dominant continuous-action head's biggest open problem.

  • Galaxea / G0 (arXiv 2509.00576) β€” Open-world dataset (500 hours, 100K trajectories, 150 task categories, 50 scenes, 11 sites, 1,600 objects) on a single 23-DoF mobile-bimanual embodiment (Galaxea R1 Lite), plus a dual-system VLA. Single-embodiment pretraining significantly beats no-pretraining in 20-trajectory few-shot settings, and the System-2 G0-VLM reaches 83.3% on Table Bussing vs 32.0% for Gemini-2.5-pro. A strong argument for embodiment-consistent data.

  • OmniVLA (navigation) (arXiv 2509.19480, UC Berkeley) β€” Omni-modal navigation VLA built on an OpenVLA-class backbone, trained on ~9,500 hours across 10 platforms, that flexibly composes language / spatial-coordinate / visual-reference goals and follows novel language instructions zero-shot. Name-collision warning: a different ICRA 2026 paper, "OmniVLA: Physically-Grounded Multimodal VLA …" (arXiv 2511.01210, Princeton/UCLA/MSRA), is a manipulation model with IR/radar/audio fusion β€” unrelated.

  • FD-VLA (arXiv 2602.02142) β€” Force-Distilled VLA: distills a force-awareness token from vision + proprioception into the VLM so a force/torque sensor is required only during training, not deployment. Targets the largest ICRA VLA theme (sensor grounding) with the most practical framing β€” contact-rich performance without contact-rich hardware.

Complete paper list (45)

Paper code Title arXiv
ThI1I.100 Learning End-To-End Dexterous Arm-Hand VLA Policies with Shared Autonomy: DexGrasp AI Copilot for Efficient Teleoperation 2511.00139
ThI1I.271 Learning Affordances at Inference-Time for Vision-Language-Action Models 2510.19752
ThI1I.301 Open-World Object Manipulation with Vision-Language-Action Models Via Synthetic Multi-Modal Data
ThI2I.128 Do What You Say: Steering Vision-Language-Action Models Via Runtime Reasoning-Action Alignment Verification 2510.16281
ThI2I.154 FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation 2602.02142
ThI2I.164 Offline Reinforced Finetuning for Chunk-Based VLA Via Real-World RL Policy Distillation with Vision-Guided Copilot
ThI2I.249 CRAFT: Adapting VLA Models to Contact-Rich Manipulation Via Force-Aware Curriculum Fine-Tuning 2602.12532
ThI2I.69 The Better You Learn, the Smarter You Prune: Towards Efficient Vision-Language-Action Models Via Differentiable Token Pruning (LightVLA) 2509.12594
ThI2I.82 Galaxy Open-World Dataset and G0 Dual-System VLA Model 2509.00576
ThI2LB.11 Robust Unknown Object Detection and Tracking for Vision-Language-Action Models on Edge Devices
TuAT3.2 DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation 2603.00926
TuI1I.167 Developing Vision-Language-Action Model from Egocentric Videos 2509.21986
TuI1I.179 RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI 2509.14687
TuI1I.185 Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models 2509.26251
TuI1I.332 RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models 2509.21243
TuI1I.84 EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation 2511.05397
TuI1LB.21 Toward Human Preference Optimization for Vision-Language-Action Models: A Pilot Study on the Limits of Imitation Learning
TuI1LB.22 Enhancing VLA Precision in Robotic Manipulation Via FiLM-Based Force/Torque-Vision Integration
TuI2I.104 DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning 2510.13375
TuI2I.128 VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning Via Online Monte Carlo Tree Search 2509.22643
TuI2I.277 A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking
TuI2I.287 NeuroVLA: Surgical Scenario-Aware Learning of Debulking Skills in Endoscopic Robotic Neurosurgery Via Vision-Language-Action Model
TuI2I.295 Scalable Vision-Language-Action Model Pretraining for Robotic Dexterous Manipulation with Real-Life Human Activity Videos 2510.21571
TuI2I.319 TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking 2510.07134
WeAT1.1 Dexora: Open-Source VLA for High-DoF Bimanual Dexterity 2605.18722
WeI1I.142 Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline 2602.22663
WeI1I.145 SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting
WeI1I.148 Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics
WeI1I.160 Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models (FPO) 2510.09976
WeI1I.238 Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation 2511.09958
WeI1I.248 ACG: Action Coherence Guidance for Flow-Based Vision-Language-Action Models 2510.22201
WeI1I.287 INSIGHT: INference-Time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models 2510.01389
WeI1I.296 Goal-VLA: Image-Generative VLMs As Object-Centric World Models Empowering Zero-Shot Robot Manipulation 2506.23919
WeI1I.311 TMR-VLA: Vision-Language-Action Model for Magnetic Motion Control of Tri-Leg Silicone-Based Soft Robot 2603.00420
WeI1I.87 Toward Embodiment Equivariant Vision-Language-Action Policy 2509.14630
WeI1LB.7 Hierarchical LLM-VLA-Controller Integration for Task Generalization
WeI2I.109 CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human 2509.14889
WeI2I.145 AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models 2602.10698
WeI2I.162 MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation 2511.09516
WeI2I.241 ExpReS-VLA: Specializing Vision-Language-Action Models through Experience Replay and Retrieval 2511.06202
WeI2I.283 OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation 2511.01210
WeI2I.301 Stream-To-Act: ROS 2 Native Token Streaming for Continuous Motion Execution of Vision-Language-Action Models
WeI2I.324 UrbanVLA: A Vision-Language-Action Model for Urban Micromobility 2510.23576
WeI2I.71 OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation 2509.19480
WeI2I.72 InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning 2505.13888

Related

← Back to ICRA-2026-VLA-Manipulation-Survey