ICRA 2026 Topic VLA - Heungwoo/research GitHub Wiki
ICRA 2026 β Vision-Language-Action Models (Topic Analysis)
This page analyzes the 45 papers at ICRA 2026 that carry "Vision-Language-Action" / "VLA" in the title or are core VLA contributions. The character of the ICRA cohort is distinct from the ML-venue VLA wave (ICLR 2026, CVPR 2026): it is deployment- and sensor-centric. Where CVPR/ICLR optimize architecture and reasoning benchmarks, ICRA's VLA work overwhelmingly asks how the model meets the robot β adding force/tactile/audio/depth to the VLMβaction stack (usually without new hardware), pruning and streaming tokens for edge latency, making continuous-action (flow/diffusion) policies improvable from reward, and stress-testing practicality, robustness, and security on real platforms. Notable cross-cutting signals: (1) sensor grounding is the single largest cluster β vision-only VLA is now treated as insufficient for contact-rich work; (2) flow-matching policies have become the default continuous-action head, and several papers attack their two open problems (RL fine-tuning, action-coherence/jitter); (3) inference-time methods (MCTS search, introspection, affordance learning, memory prompting) are proliferating because they upgrade frozen pretrained VLAs cheaply; and (4) navigation VLA has split into its own coherent sub-field with multi-modal goal conditioning.
Sub-trends
1. Sensor-grounded VLA (force Β· tactile Β· audio Β· multi-sensor)
The largest theme. The shared premise is that RGB-only VLMs lack the contact and dynamic-process signals needed for precise, contact-rich manipulation, and the shared trick is to inject the missing modality without demanding new sensors at deployment. FD-VLA distills a force token from vision+state into the VLM so a force/torque sensor is needed only at training time; CRAFT instead adapts a VLA to contact tasks via a force-aware curriculum fine-tuning schedule; and Enhancing VLA Precision via FiLM-based Force/Torque-Vision Integration fuses F/T into the visual stream through FiLM conditioning for insertion-grade alignment. Audio-VLA adds contact-audio perception (AudioCLIP encoder, LoRA-fine-tuned Llama2 backbone) to perceive collision/contact events that vision misses, and introduces a Task-Completion-Rate metric for dynamic processes. The manipulation-side OmniVLA (arXiv 2511.01210 β note the name collision with the navigation OmniVLA in Β§7) generalizes this to unified multi-sensor perception, fusing IR/radar/audio into a physically-grounded policy. The pattern across all five: a modality-specific encoder feeds tokens or FiLM parameters into an otherwise-standard VLA, and the gains concentrate on alignment/insertion/deformable tasks where vision alone plateaus.
2. Spatial / 3D / depth grounding
A cluster attacking VLMs' weak 3D spatial reasoning, which it inherits from 2D-pretrained backbones. DepthVLA adds a pretrained depth transformer in a mixture-of-transformers design (VLM + depth expert + action expert, fully shared attention) and reports large real-world and Simpler-simulator gains. AugVLA-3D does depth-driven feature augmentation rather than adding a stream. InSpire takes a language route: it prepends a spatial-reasoning VQA ("in which direction is the [object] relative to the robot?") to redirect attention to task-relevant factors and break spurious visual correlations, needing no auxiliary data. RetoVLA reuses otherwise-discarded ViT register tokens as a free spatial-reasoning signal, framing it as compression-friendly. Seeing Space and Motion improves Latent Action Models by injecting geometric and dynamic awareness into the latent action encoder. The common insight: spatial competence can be added through an auxiliary expert, an auxiliary language task, or recovered from latent/register representations already present β and all three beat scaling action data alone.
3. Reasoning, verification & introspection (inference-time)
These methods upgrade frozen pretrained VLAs at inference time. VLA-Reasoner runs online Monte-Carlo Tree Search over candidate action chunks, scoring each branch by rolling it out through a 600M action-aware world model β treating imagined outcomes as "rationales." Do What You Say addresses a subtler failure: a reasoning VLA can emit a correct textual plan yet a misaligned low-level action, so it adds runtime reasoning-action alignment verification to steer execution back to the stated intent. INSIGHT uses token-level uncertainty as an introspective help-trigger signal so the policy can ask a human supervisor before failing. Learning Affordances at Inference-Time lets the VLA reflect on failed tries and revise affordances mid-deployment. CollabVLA turns a visuomotor policy into a self-reflective collaborator that "dreams together" with a human, targeting interpretability and latency. The unifying move: rather than retrain, add a search / verification / uncertainty / reflection loop on top β a cheap path to reliability that the ICRA cohort clearly favors.
4. Efficiency & deployment (pruning Β· streaming Β· edge Β· practicality)
ICRA's signature concern. LightVLA is a parameter-free, hyperparameter-free differentiable token pruner that cuts FLOPs/latency while raising success rate (it inverts the usual efficiency/accuracy trade-off). RetoVLA (also Β§2) doubles as a compression method. Stream-To-Act is a ROS 2-native token-streaming runtime that emits actions continuously for real-time control loops instead of blocking on full-sequence decoding. Robust Unknown Object Detection and Tracking on Edge Devices proposes a stepwise VLA for Jetson AGX Orin, sidestepping the memory blow-up of end-to-end VLA on edge hardware. EveryDayVLA pushes the hardware floor down with a ~$300 6-DOF manipulator and a matching policy. VLA Practicality ("Rethinking the Practicality of VLA") is the cluster's benchmark/baseline statement, arguing existing VLAs are over-parameterized and pre-training-hungry, and offering a leaner baseline. Together they form a coherent "make VLA actually deployable" agenda spanning compute, latency, hardware cost, and evaluation.
5. RL & preference fine-tuning of continuous-action policies
A focused, technically deep cluster around the open problem of improving flow-matching/diffusion VLAs from reward (their likelihoods are intractable, so standard RL/DPO does not transfer cleanly). FPO / Reinforcement Fine-Tuning of Flow-Matching Policies derives a flow-compatible policy-optimization objective with stable convergence under sparse reward, evaluated on LIBERO and ALOHA. Offline Reinforced Finetuning for Chunk-Based VLA distills a real-world RL policy into a chunk-based VLA via a vision-guided copilot, keeping the action-chunking interface. Toward Human Preference Optimization is a pilot study probing the limits of imitation learning on multi-step tasks (evaluating GR00T N1.6) and motivating preference-based signals. ACG (Action Coherence Guidance) targets a flow-specific pathology β high generative capacity makes flow policies jittery/noise-sensitive under imitation β with a guidance term enforcing temporal action coherence. The cluster reflects that flow-matching is now the dominant action head, and ICRA is where its reward-driven and stability fixes are being worked out.
6. Memory, continual learning & specialization
Pretrained VLAs generalize broadly but forget, lack long-horizon memory, and underperform on the specific deployment they are sent to. MAP-VLA adds memory-augmented prompting so a frozen VLA can carry context across a long-horizon task instead of relying only on the current frame. ExpReS-VLA specializes a generalist (e.g. OpenVLA) to a fixed task set via experience replay and retrieval, trading broad zero-shot for consistent in-distribution performance. Learning Affordances at Inference-Time (also Β§3) is continual-learning-flavored, accumulating affordance corrections across attempts. The shared diagnosis: deployment values reliable specialization with memory over broad-but-shallow generalization, and lightweight prompting/retrieval beats full retraining.
7. Navigation, tracking & world-model / goal conditioning
Navigation VLA is now its own sub-field. OmniVLA (nav) (UC Berkeley; ~9,500 hours across 10 platforms) is an omni-modal navigation policy that composes language, spatial-coordinate, and visual-reference goals in one model. UrbanVLA is a route-conditioned model for urban micromobility/delivery, aligning noisy map routes with egocentric vision and trained SFT-then-RFT (reports >55% improvement over baselines on MetaUrban; real >500m routes on a Unitree Go2). TrackVLA++ brings reasoning (a Polar-CoT spatial token) and a Target-Identification Memory to embodied visual tracking, for robustness under occlusion and distractors. On the manipulation side, world-model/goal conditioning appears in Goal-VLA, which uses image-generative VLMs as object-centric world models for zero-shot manipulation via reflection-through-synthesis, and in DAM-VLA, a dynamic action model bridging gross motion and precise manipulation in dynamic scenes. The common thread is goal flexibility: conditioning on coordinates, images, routes, or synthesized goal states rather than language alone.
8. Data, pretraining, platforms & dexterity
The supply-side and embodiment cluster. Galaxea / G0 pairs a 500-hour, 100K-trajectory single-embodiment open-world dataset with a dual-system VLA (the System-2 G0-VLM hits 83.3% on Table Bussing vs 32.0% for Gemini-2.5-pro). Dexora is a fully open-source stack (hardware + data + policy) for high-DoF (36-DoF) bimanual dexterity, a regime usually locked behind proprietary hardware. Two papers pretrain from human video instead of teleoperation β Developing VLA from Egocentric Videos and Scalable VLA Pretraining with Real-Life Human Activity Videos (treating the human hand as the end-effector) β a scalable-data strategy. Open-World Object Manipulation via Synthetic Multi-Modal Data generates synthetic data for object generalization. RealMirror is an open platform (3D-Gaussian-Splatting reconstruction + sim-to-real) for end-to-end VLA research without a real robot. DexGrasp AI Copilot learns end-to-end dexterous arm-hand policies with shared-autonomy teleoperation. The cluster's bet: scalable non-teleop data and open embodiments are the path past the data bottleneck.
9. Generalization, robustness & security (cross-cutting)
A smaller but important set. Exploiting Vulnerabilities demonstrates universal adversarial attacks on VLA models in robotics β a rare security-focused contribution showing physical-world VLAs are attackable. SVP (Dual Stochastic Visual Prompting) diagnoses "distracted attention" as shortcut learning in models like OpenVLA and fixes it with stochastic visual prompts rather than architecture changes. Toward Embodiment Equivariant VLA Policy attacks cross-embodiment generalization via equivariance rather than scale. Hierarchical LLM-VLA-Controller Integration layers an LLM planner over a VLA to combat memorization-over-semantics. Domain-specific deployments round out the cohort: ultrasound-guided needle insertion, NeuroVLA (endoscopic neurosurgery debulking), and TMR-VLA (magnetic control of a tri-leg silicone soft robot) β three medical/soft-robot VLAs proving the paradigm transfers well beyond tabletop grippers.
Standout deep-dives
-
VLA-Reasoner (arXiv 2509.22643) β Online MCTS over action chunks, scored by rolling out a 600M action-aware world model so imagined outcomes serve as rationales. Improves frozen policies on LIBERO: OpenVLA-SFT 76.0%β81.0% (+5.0 pp), Octo-Small 26.5%β37.3% (+10.8 pp), SpatialVLA 34.0%β41.8% (+7.8 pp). A clean demonstration that test-time search + a learned world model upgrades any pretrained VLA without retraining.
-
LightVLA (arXiv 2509.12594) β Differentiable, parameter-free and hyperparameter-free token pruning that retains ~78 visual tokens on average. Reports β59.1% FLOPs and β38.2% latency while raising success rate to 97.4% avg vs OpenVLA-OFT 94.5% (+2.9 pp), and far above VLM-oriented pruners (FlashVLA 73.7%, SP-VLA 74.9%, VLA-Cache 74.7%) that collapse on VLA tasks. The headline is that performance-driven pruning breaks the usual efficiency/accuracy trade-off.
-
Goal-VLA (arXiv 2506.23919) β Uses image-generative VLMs as object-centric world models for zero-shot manipulation. On RLBench (8 tasks, 100 runs each) averages 59.9% vs MOKA 26.0%, MolmoAct 11.3%, Ο0 0.0%; real-world (4 tasks) 60% vs MolmoAct 27.5%. Ablation: a 40.0% baseline rises to 83.8% with input enhancement + reflector and 88.8% with three reflection iterations, quantifying the reflection-through-synthesis loop.
-
FPO β RL Fine-Tuning of Flow-Matching Policies (arXiv 2510.09976) β A flow-compatible policy-optimization objective (CFM-based) that makes flow-matching VLAs improvable from sparse reward where standard RL/DPO does not transfer. LIBERO success: Spatial 97.2 Β· Object 97.3 Β· Goal 89.4 Β· Long 65.3 Β· avg 87.2; ALOHA Transfer Cube learning curve reaches ~65%. Directly addresses the dominant continuous-action head's biggest open problem.
-
Galaxea / G0 (arXiv 2509.00576) β Open-world dataset (500 hours, 100K trajectories, 150 task categories, 50 scenes, 11 sites, 1,600 objects) on a single 23-DoF mobile-bimanual embodiment (Galaxea R1 Lite), plus a dual-system VLA. Single-embodiment pretraining significantly beats no-pretraining in 20-trajectory few-shot settings, and the System-2 G0-VLM reaches 83.3% on Table Bussing vs 32.0% for Gemini-2.5-pro. A strong argument for embodiment-consistent data.
-
OmniVLA (navigation) (arXiv 2509.19480, UC Berkeley) β Omni-modal navigation VLA built on an OpenVLA-class backbone, trained on ~9,500 hours across 10 platforms, that flexibly composes language / spatial-coordinate / visual-reference goals and follows novel language instructions zero-shot. Name-collision warning: a different ICRA 2026 paper, "OmniVLA: Physically-Grounded Multimodal VLA β¦" (arXiv 2511.01210, Princeton/UCLA/MSRA), is a manipulation model with IR/radar/audio fusion β unrelated.
-
FD-VLA (arXiv 2602.02142) β Force-Distilled VLA: distills a force-awareness token from vision + proprioception into the VLM so a force/torque sensor is required only during training, not deployment. Targets the largest ICRA VLA theme (sensor grounding) with the most practical framing β contact-rich performance without contact-rich hardware.
Complete paper list (45)
| Paper code | Title | arXiv |
|---|---|---|
| ThI1I.100 | Learning End-To-End Dexterous Arm-Hand VLA Policies with Shared Autonomy: DexGrasp AI Copilot for Efficient Teleoperation | 2511.00139 |
| ThI1I.271 | Learning Affordances at Inference-Time for Vision-Language-Action Models | 2510.19752 |
| ThI1I.301 | Open-World Object Manipulation with Vision-Language-Action Models Via Synthetic Multi-Modal Data | |
| ThI2I.128 | Do What You Say: Steering Vision-Language-Action Models Via Runtime Reasoning-Action Alignment Verification | 2510.16281 |
| ThI2I.154 | FD-VLA: Force-Distilled Vision-Language-Action Model for Contact-Rich Manipulation | 2602.02142 |
| ThI2I.164 | Offline Reinforced Finetuning for Chunk-Based VLA Via Real-World RL Policy Distillation with Vision-Guided Copilot | |
| ThI2I.249 | CRAFT: Adapting VLA Models to Contact-Rich Manipulation Via Force-Aware Curriculum Fine-Tuning | 2602.12532 |
| ThI2I.69 | The Better You Learn, the Smarter You Prune: Towards Efficient Vision-Language-Action Models Via Differentiable Token Pruning (LightVLA) | 2509.12594 |
| ThI2I.82 | Galaxy Open-World Dataset and G0 Dual-System VLA Model | 2509.00576 |
| ThI2LB.11 | Robust Unknown Object Detection and Tracking for Vision-Language-Action Models on Edge Devices | |
| TuAT3.2 | DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation | 2603.00926 |
| TuI1I.167 | Developing Vision-Language-Action Model from Egocentric Videos | 2509.21986 |
| TuI1I.179 | RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI | 2509.14687 |
| TuI1I.185 | Seeing Space and Motion: Enhancing Latent Actions with Geometric and Dynamic Awareness for Vision-Language-Action Models | 2509.26251 |
| TuI1I.332 | RetoVLA: Reusing Register Tokens for Spatial Reasoning in Vision-Language-Action Models | 2509.21243 |
| TuI1I.84 | EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation | 2511.05397 |
| TuI1LB.21 | Toward Human Preference Optimization for Vision-Language-Action Models: A Pilot Study on the Limits of Imitation Learning | |
| TuI1LB.22 | Enhancing VLA Precision in Robotic Manipulation Via FiLM-Based Force/Torque-Vision Integration | |
| TuI2I.104 | DepthVLA: Enhancing Vision-Language-Action Models with Depth-Aware Spatial Reasoning | 2510.13375 |
| TuI2I.128 | VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning Via Online Monte Carlo Tree Search | 2509.22643 |
| TuI2I.277 | A Vision-Language-Action Model for Adaptive Ultrasound-Guided Needle Insertion and Needle Tracking | |
| TuI2I.287 | NeuroVLA: Surgical Scenario-Aware Learning of Debulking Skills in Endoscopic Robotic Neurosurgery Via Vision-Language-Action Model | |
| TuI2I.295 | Scalable Vision-Language-Action Model Pretraining for Robotic Dexterous Manipulation with Real-Life Human Activity Videos | 2510.21571 |
| TuI2I.319 | TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking | 2510.07134 |
| WeAT1.1 | Dexora: Open-Source VLA for High-DoF Bimanual Dexterity | 2605.18722 |
| WeI1I.142 | Rethinking the Practicality of Vision-Language-Action Model: A Comprehensive Benchmark and an Improved Baseline | 2602.22663 |
| WeI1I.145 | SVP: Improving Vision-Language-Action Models with Dual Stochastic Visual Prompting | |
| WeI1I.148 | Exploiting Vulnerabilities: Universal Adversarial Attacks on Vision-Language-Action Models in Robotics | |
| WeI1I.160 | Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models (FPO) | 2510.09976 |
| WeI1I.238 | Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation | 2511.09958 |
| WeI1I.248 | ACG: Action Coherence Guidance for Flow-Based Vision-Language-Action Models | 2510.22201 |
| WeI1I.287 | INSIGHT: INference-Time Sequence Introspection for Generating Help Triggers in Vision-Language-Action Models | 2510.01389 |
| WeI1I.296 | Goal-VLA: Image-Generative VLMs As Object-Centric World Models Empowering Zero-Shot Robot Manipulation | 2506.23919 |
| WeI1I.311 | TMR-VLA: Vision-Language-Action Model for Magnetic Motion Control of Tri-Leg Silicone-Based Soft Robot | 2603.00420 |
| WeI1I.87 | Toward Embodiment Equivariant Vision-Language-Action Policy | 2509.14630 |
| WeI1LB.7 | Hierarchical LLM-VLA-Controller Integration for Task Generalization | |
| WeI2I.109 | CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human | 2509.14889 |
| WeI2I.145 | AugVLA-3D: Depth-Driven Feature Augmentation for Vision-Language-Action Models | 2602.10698 |
| WeI2I.162 | MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation | 2511.09516 |
| WeI2I.241 | ExpReS-VLA: Specializing Vision-Language-Action Models through Experience Replay and Retrieval | 2511.06202 |
| WeI2I.283 | OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation | 2511.01210 |
| WeI2I.301 | Stream-To-Act: ROS 2 Native Token Streaming for Continuous Motion Execution of Vision-Language-Action Models | |
| WeI2I.324 | UrbanVLA: A Vision-Language-Action Model for Urban Micromobility | 2510.23576 |
| WeI2I.71 | OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation | 2509.19480 |
| WeI2I.72 | InSpire: Vision-Language-Action Models with Intrinsic Spatial Reasoning | 2505.13888 |
Related
β Back to ICRA-2026-VLA-Manipulation-Survey