ICML 2026 - Heungwoo/research GitHub Wiki
ICML 2026
International Conference on Machine Learning 2026 β COEX Convention & Exhibition Center, Seoul, South Korea, July 6β11, 2026 (Jul 6 tutorials/expo Β· Jul 7β9 main conference Β· Jul 10β11 workshops).
Scale: 23,918 submissions β ~6,352 accepted (β26.6%) per the public submission statistics; the official virtual proceedings list 6,636 distinct accepted papers (6,804 schedule events, of which 168 are Orals and the rest Posters). Spotlight/Oral tier β the top 0.7% of submissions.
Robotics / manipulation share: a keyword sweep of the official accepted list surfaced ~220 robotics / embodied-AI candidates, of which 99 are genuinely about robot manipulation (the rest are autonomous-driving VLAs, pure navigation, world-model/RL learning theory, or non-robotic uses of "manipulation"). That makes manipulation β 1.5% of all ICML 2026 papers β smaller than at CVPR/ICLR by share, but the largest manipulation presence ICML has ever had, and notably ML-methodology-heavy (systematic studies, recipes, theory) rather than systems-paper-heavy.
Source & method. This index was built against the official virtual data file
https://icml.cc/static/virtual/data/icml-2026-orals-posters.json(2026-06 snapshot, 6,804 events). Every paper below was filtered by title+keyword match, then each abstract was fetched from itsicml.cc/virtual/2026/poster/<id>page and summarized individually. Quantitative figures in back-ticks are quoted from the paper's own abstract; where an abstract states no number, none is shown (no estimated/invented metrics). OpenReview blocks guest access to ICML 2026 notes, so author institutions are not yet attached. Treat one-line summaries as abstract-derived, pending full-text review.
Surveys hosted in this wiki
- Per-paper in-depth pages now exist for all 99 manipulation papers (linked from the category list below) β each with Problem Β· Method Β· Results Β· Significance and, where the paper has a preprint, the paper's own figures embedded. A few highest-value papers also have long-form reviews: [Review-RoboMME]] and [Action Space: EEF vs Joint.
Orals among the manipulation papers (5)
- XR-1 β Unified Vision-Motion Codes (dual-branch VQ-VAE), 12,000+ real rollouts across 6 embodiments. (also tagged ICLR 2026 in our wiki β verify which venue is canonical)
- From Pixels to Tokens β systematic study of latent-action supervision for VLAs; discrete latent action tokens win.
- From Abstraction to Instantiation (BehaviorVLA) β causal Mamba behavior encoder + phase-conditioned decoder for robustness under shift.
- Pretrained VLAs are Surprisingly Resistant to Forgetting in Continual Learning β big pretrained VLAs barely forget; simple Experience Replay can hit zero forgetting.
- RoboMME β standardized benchmark for memory in robotic generalist policies (16 tasks, 14 memory-augmented Ο0.5 variants).
π Top-20 technically notable papers (synthesized ranking)
Synthesis of editorial judgment + quantitative signals (GitHub β + arXiv-preprint citations via Semantic Scholar + Oral status), snapshot 2026-06-09. Pure simulators / data-generators are excluded (RoboTwin 2.0, VLA-Arena, CaP-X, OXE-AugE, SoMA, DLO-Lab, FlatLab, ManiSoft, SafeLab, AIR-VLA) β but insight / evaluation papers are kept (RoboMME, From Pixels to Tokens, Demystifying Action Space, Pretrained-VLA-Forgetting). β /citation are early preprint-era signals that favor early code-releasers and undercount work with no public repo/arXiv.
| # | Paper | Category | β | cite | Oral | Affiliations |
|---|---|---|---|---|---|---|
| 1 | RDT2 ΒΉ | Foundation/scaling | 775 | 15 | Tsinghua University (THU-ML) | |
| 2 | DreamDojo | World model | 923 | 42 | NVIDIA; HKUST; UC Berkeley; UW; Stanford; KAIST | |
| 3 | XR-1 ΒΉ | Architecture | 174 | 11 | β | Beijing Innovation Center of Humanoid Robotics (X-Humanoid); Beihang; PKU |
| 4 | Discrete Diffusion VLA ΒΉ | Architecture | 65 | 64 | HKU; Shanghai AI Lab; SJTU; Huawei | |
| 5 | Being-H0 ΒΉ | Human-video pretrain | 48 | 69 | Peking University; Renmin University; BeingBeyond | |
| 6 | DexMachina | Dexterous | 225 | 31 | Stanford University; NVIDIA | |
| 7 | VLAC (Progress Critic) | RL for VLA | 306 | β | Shanghai AI Laboratory (InternRobotics) | |
| 8 | Latent Reasoning VLA | Reasoning | 67 | 7 | Tsinghua; PKU; USTC | |
| 9 | LangForce | Analysis/method | 65 | 10 | ZGC-EmbodyAI (Zhongguancun Academy) | |
| 10 | HALO | Reasoning | β | 2 | HKUST | |
| 11 | Dual-Stream Diffusion | World model | β | 13 | KAIST (RLWRLD) | |
| 12 | DECO | Dexterous/tactile | 28 | 0 | BAAI; TU Munich | |
| 13 | See What Matters | Efficiency | 62 | 1 | University of Sydney | |
| 14 | SpecPrune-VLA | Efficiency | β | 24 | SJTU; Infinigence-AI; Shanghai Innovation Institute | |
| 15 | BehaviorVLA | Representation | β | 0 | β | HIT (Shenzhen); Sun Yat-sen University |
| 16 | VLANeXt | Recipe/insight | 196 | 5 | S-Lab NTU; SYSU; ACE Robotics | |
| 17 | From Pixels to Tokens | Insight | 28 | 0 | β | Renmin University (KBReasoning) |
| 18 | Pretrained VLAs Resist Forgetting | Insight | β | 4 | β | UT Austin; KAIST; Microsoft |
| 19 | Demystifying Action Space | Insight | β | 1 | Tsinghua (IAIR/Wuxi) | |
| 20 | RoboMME | Memory insight | 111 | 6 | β | University of Michigan; Stanford; Figure AI |
ΒΉ Already has a wiki page under another 2026 venue (CVPR/ICLR); these titles are also in the official ICML 2026 accepted list, so venue-of-record needs reconciliation β link points to the existing page.
How to read it: the top tier (RDT2, DreamDojo, XR-1, DexMachina) is strong on both axes. Citations elevate Discrete Diffusion VLA (64) and Being-H0 (69) β the most-cited methods. Editorial judgment retains low-traction-but-important work: the Oral insight papers (From Pixels to Tokens, Pretrained-VLA-Forgetting, RoboMME) and the latent-reasoning / efficiency methods. 7 of 20 have industry involvement (NVIDIA Γ2, X-Humanoid, BeingBeyond, Figure AI, Huawei, Infinigence-AI).
Each paper links to a detail page (Problem Β· Method Β· Results Β· Significance Β· Links); page metrics use the same 2026-06-09 snapshot.
Manipulation papers by category
99 papers. Headline numbers in back-ticks are quoted from the abstract.
VLA architecture & backbones
- From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models (Oral) β A systematic comparison of image-based versus action-based latent-action supervision strategies for VLAs under a unified baseline, finding that directly supervising the VLM with discrete latent action tokens is most effective.
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations (Oral) β A VLA model that learns Unified Vision-Motion Codes via a dual-branch VQ-VAE jointly encoding visual dynamics and robotic motion, trained in three stages to bridge cross-embodiment and human-demonstration domain gaps. β
validated through over 12,000 real-world rollouts across six robot embodiments and 120+ manipulation tasks - Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds β Introduces Any3D-VLA, which merges simulator-generated, sensor-derived, and model-estimated point clouds into a unified training framework to learn domain-agnostic 3D representations fused with 2D inputs for more robust VLA spatial understanding.
- Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting β Introduces Visual Attentive Prompting (VAP), a training-free perceptual adapter that grounds personalized objects via open-vocabulary detection and injects visual prompts to give frozen VLAs top-down selective attention for personalized manipulation commands.
- Contrastive Representation Regularization for Vision-Language-Action Models β RS-CL is a robot-state-aware contrastive representation regularizer that aligns VLA representations with proprioceptive states using inter-state distances as soft supervision to improve manipulation performance. β
69.7% on RoboCasa-Kitchen; real-robot success 45.0% to 58.3% - Demystifying Action Space Design for Robotic Manipulation Policies β A large-scale empirical study dissecting action space design along temporal and spatial axes (absolute vs. delta, joint-space vs. task-space) for imitation-based manipulation policies and its effect on learnability and control stability. β
Based on 13,000+ real-world rollouts on a bimanual robot and evaluation on 500+ trained models over four scenarios - EnsembleVLA: Ensemble Learning for Vision-Language Action Models β EnsembleVLA is an energy-based framework that formulates diffusion- and flow-based VLA models as energy-based models to compose multiple pretrained policies with learnable weights and confidence-aware gating.
- FOCA: Future-Oriented Conditioning for Data-Efficient Vision-Language-Action Adaptation β Proposes FOCA, a VLA adaptation method that predicts future interaction embeddings and aligns to future goal observations (with optional co-training on world-model synthetic videos) for data-efficient few-shot imitation. β
95.7% success with 20 demonstrations on LIBERO - Fourier Features Let Agents Learn High Precision Policies with Imitation Learning β The paper maps point clouds into high-dimensional Fourier space via a parametric projection to overcome neural networks' spectral bias, improving high-precision point-cloud imitation learning.
- GeoMoLa: Geometry-Aware Motion Latents for Learning Robust Manipulation Policies β Introduces GeoMoLa, which learns discrete motion latent codes by predicting how point clouds evolve during manipulation (a 4D spatial-temporal objective) rather than reconstructing visual observations, using only single-view RGB-D input.
- LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries β LangForce uses a dual-branch Bayesian decomposition with learnable Latent Action Queries to maximize conditional pointwise mutual information between actions and instructions, countering the vision shortcut in VLA training. β
11.3% improvement on the OOD SimplerEnv benchmark - Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs β Task-Agnostic Pretraining (TAP) pretrains a VLA on cheap off-task/play trajectories via an inverse-dynamics objective, then aligns the learned physical priors with language using minimal expert data. β
25% vs. 0% under camera perturbations - Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation β A VLA framework that decouples manipulation into coarse 'move' and contact-critical 'operate' phases using a dual-expert policy routed by a learnable phase selector with MLLM-generated phase labels. β
average success rate of 68.9%, outperforming the monolithic Pi0 baseline by +24% - N2M: Bridging Navigation and Manipulation by Learning Pose Preference from Rollout β A mobile-manipulation base-positioning method (N2M) that learns policy-aware, viewpoint-invariant pose preferences from policy rollouts using a viewpoint augmentation strategy, avoiding pre-built scene reconstruction.
- NeurVLA: Unleashing Failure-Handling Capability of Vision-Language-Action Models via Neural-Symbolic Reasoning β NeurVLA is a neural-symbolic framework that jointly handles failure correction and prevention via reasoning and internalizes these failure-handling capabilities into VLA models for robotic manipulation.
- Neural Implicit Action Fields: From Discrete Waypoints to Continuous Functions for Vision-Language-Action Models β Reformulates VLA action prediction from discrete waypoints to continuous-time action function regression using an MLLM as a spectral modulator over a learnable motion prior, enabling differentiable velocity/acceleration/jerk supervision. β
achieves state-of-the-art results on CALVIN and LIBERO benchmarks - RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation β A retrieval-augmented VLA framework for training-free test-time adaptation that combines behavior-aligned context retrieval with a grounded execution pipeline to overcome in-context imitation learning's adaptation bottleneck.
- RDT2: Exploring the Scaling Limit of UMI Data Towards Zero-Shot Cross-Embodiment Generalization β RDT2 is a 7B-parameter robotic foundation model trained on over 10,000 hours of UMI-collected data with a three-stage RVQ/flow-matching/distillation recipe for zero-shot cross-embodiment manipulation. β
over 10,000 hours of demonstrations - Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action β A spatial memory framework (SOMA) that equips VLA models with persistent spatial-semantic memory from multi-view head-camera observations, enabling manipulation of objects outside the current camera view.
- UniCoD: Enhancing Robot Policy via Unified Continuous and Discrete Representation Learning β UniCoD learns a unified continuous-and-discrete representation by pretraining on over 1M internet manipulation videos and fine-tuning on robot data to map predictive representations to action tokens. β
9% and 12% over baselines in sim and real-world OOD tasks - VLANeXt: Recipes for Building Strong VLA Models β A unified empirical study dissecting VLA design choices across foundational components, perception, and action modeling, distilling 12 findings into a recipe and a model (VLANeXt). β
outperforms prior state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks
Efficiency Β· pruning Β· deployment
- Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment β Provides a systematic model-hardware co-characterization framework for low-cost VLA edge deployment, identifying a two-phase (compute-bound VLM, memory-bound action expert) inference pattern and proposing DP-Cache and V-AEFusion to reduce diffusion redundancy and enable asynchronous pipeline parallelism. β
up to 2.9x speedup on GPUs and 3.3x on edge NPUs with only marginal success degradation - EcoVLA: Environment-Aware Adaptive Pruning with Interleaved Inference Orchestration for Vision-Language-Action Models β EcoVLA is a training-free adaptive channel-pruning framework with environment-aware sparsity updates and interleaved inference scheduling that accelerates VLA inference with minimal accuracy loss. β
up to 1.60x speedup with only 0.4% drop in success rate (2.18x combined with token pruning) - Reflex: Real-Time Vision-Language-Action Control through Streaming Inference β Reflex enables real-time streaming inference for flow-matching VLAs by exploiting timestep-invariance to partition attention into static/sliding/dynamic regions for O(1) cache updates, plus AdaRMSNorm and an async pipeline. β
achieves a 2.58x inference speedup and 50Hz stable streaming, reducing reaction latency by up to 54% - STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction β A lightweight spatiotemporal consistency prediction mechanism that constructs high-quality warm-start actions plus a velocity-aware perturbation injection scheme to accelerate diffusion-policy inference without sacrificing action quality. β
STEP with 2 steps can achieve an average 21.6% and 27.5% higher success rate than BRIDGER and DDIM on the RoboMimic benchmark and real-world tasks, respectively - See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model β Proposes the Differentiable Grid Sampler (GridS), a plug-and-play module that performs task-aware continuous resampling of visual tokens via differentiable interpolation to drastically compress VLA visual tokens while preserving spatial information. β
76% reduction in FLOPs with no degradation in the success rate (fewer than 10% original visual tokens) - Sparse ActionGen: Accelerating Diffusion Policy with Real-time Pruning β Accelerates diffusion-policy action generation with a rollout-adaptive, observation-conditioned prune-then-reuse mechanism that caches and substitutes redundant computations across timesteps and blocks. β
achieves up to 4x generation speedup without sacrificing performance - SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning β SpecPrune-VLA is a training-free two-level (action-level static and layer-level dynamic) token pruning method with an action-aware controller that accelerates VLA inference using global context plus local attention. β
1.57x speedup in LIBERO simulation and 1.70x on real-world tasks - Speedup Patch: Learning a Plug-and-Play Policy to Accelerate Embodied Manipulation β Proposes a policy-agnostic offline-RL scheduler that downsamples action chunks under a Constrained MDP, using a learned world model as a safety-constraint surrogate to accelerate embodied policies without retraining. β
approximately 1.8x execution speedup across diverse policies while preserving original success rates - Think Less, Act Early: Reinforced Latent Reasoning with Early Exit in Vision-Language-Action Models β Proposes AVA-VLA, a latent-reasoning VLA that models reasoning as unobservable latent variables refined by an RL-based denoising mechanism and adaptively terminated via a confidence-based early-exit strategy to cut inference latency.
Reasoning Β· CoT Β· test-time scaling
- Decompose and Recompose: Reasoning New Skills from Existing Abilities for Cross-Task Robotic Manipulation β A skill-reasoning in-context-learning framework that decomposes seen demonstrations into atomic skill-action pairs and recomposes them for unseen tasks via compositional reasoning, using task-adaptive dynamic and coverage-aware static demonstration libraries.
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning β A unified VLA model using a Mixture-of-Transformers architecture to perform embodied multimodal chain-of-thought reasoning by sequentially combining textual task reasoning, visual subgoal prediction, and action prediction. β
surpassing baseline policy Pi0 by 34.1% on RoboTwin benchmark - LaST_0: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Model β Proposes a VLA framework that reasons before acting in a token-efficient latent spatio-temporal chain-of-thought space (future visual dynamics, 3D structure, proprioception), with a Mixture-of-Transformers dual-system design separating low-frequency reasoning and high-frequency action experts. β
improves mean success rates by 13%, 14% and 14% over prior SOTA VLA methods - Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models β Internalizes multi-modal Chain-of-Thought reasoning into continuous latent representations via a curriculum that transitions from explicit textual/visual CoT supervision to latent reasoning, removing explicit CoT generation at inference for efficient action control. β
up to a 90% reduction in inference latency compared to explicit CoT-based VLA approaches - SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models β Proposes a training-free, verifier-free single-forward-pass inference strategy that jointly modulates visual perception and action based on self-uncertainty, exploring more when uncertain and exploiting when confident.
- Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery β Introduces Sentinel-VLA, a metacognitive VLA with an active sentinel module for real-time status monitoring that triggers on-demand reasoning or error recovery, plus a self-evolving continual learning algorithm and orthogonal adapter to prevent forgetting. β
boosts the task success rate by over 30% compared to the SOTA model, PI0 - TapSampling: Inference-Time Sampling with a Task-Progress-Understanding Verifier for Robotic Manipulation β Proposes a policy-agnostic inference-time sampling framework that draws candidate actions from an Action-VAE latent space and selects among them with a verifier trained to predict task-progress outcomes from sequential robot data.
- Temporal Difference Calibration in Sequential Tasks: Application to Vision-Language-Action Models β Formulates sequential calibration for episodic VLA tasks via a sequential Brier score whose risk minimizer equals the policy's value function, enabling temporal-difference value estimation as a principled calibration mechanism over partial trajectories.
- VLA-ATTC: Adaptive Test-Time Compute for VLA Models with Relative Action Critic Model β Proposes a VLA framework with uncertainty-triggered adaptive test-time compute, using a Relative Action Critic that selects the best candidate action via pairwise comparisons instead of unstable absolute value estimation. β
On LIBERO-LONG, reduces the failure rate of the SOTA model PI0.5 by over 50%
World models for manipulation
- Cross-Embodiment Robot Foundation World Models with Latent Actions β A Latent Action Conditioned Robot World Model (LAC-WM) that operates in a learned unified latent action space shared across embodiments, enabling better adaptation to unseen robots than explicit-action-conditioned baselines. β
LAC-WM achieves up to a 46.7% improvement in performance over EAC-WM - DreamDojo: A Real-Time Robot World Model from Large-Scale Human Videos β DreamDojo is a foundation robot world model pretrained on 44,000 hours of egocentric human video using continuous latent actions, distilled to run in real time after robot fine-tuning. β
real-time performance at 10.93 FPS - Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model β A world-model-augmented VLA (DUST) using a multimodal diffusion transformer with separate modality streams, decoupled flow-matching loss, and asynchronous action/vision sampling to jointly predict states and actions. β
up to 6% gains over state-of-the-art VLA and world-modeling baselines, with inference-time scaling providing an additional 2-5% improvement; real-world Franka Research 3 outperforms baselines by 10% in success rate - From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation β MoLA converts imagined future videos into executable actions by using multiple modality-aware pretrained inverse dynamics models (semantic, depth, flow) to infer a mixture of latent actions bridging video imagination and policy execution.
- MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation β MVISTA-4D is an embodied 4D world model that generates geometrically consistent arbitrary-view RGBD from a single-view input and infers actions via test-time trajectory-latent optimization plus a residual inverse dynamics model.
- Predicting What Matters: Robust Generalist Robot Policy Learning via Future Semantic Mask β The Mask World Model predicts the evolution of semantic masks instead of RGB pixels using video diffusion, imposing a geometric information bottleneck, and integrates this with a diffusion policy head for robust end-to-end control.
- RoboFlow4D: A Lightweight Flow World Model Toward Real-Time Flow-Guided Robotic Manipulation β Introduces a lightweight end-to-end flow world model that predicts multi-frame 3D flows from images and instructions to guide action generation via slow-fast collaboration for real-time, resource-efficient manipulation.
- Scaling Real-World Robot Policy Evaluation via Discrete Diffusion World Model β Proposes dWorldEval, an action-centric discrete-diffusion world model that treats actions as first-class tokens in a unified token space with sparse keyframe memory and progress-as-text, enabling reliable automatic policy evaluation.
- SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-Body Manipulation β SoMA is a 3D Gaussian-Splat neural simulator that couples deformable dynamics, environmental forces, and robot joint actions in a unified latent space for end-to-end real-to-sim soft-body manipulation. β
improves resimulation accuracy and generalization by 20% - Structured 4D Latent World Model for Robot Planning β A world model that predicts the evolution of a scene's 3D structure in a structured latent space conditioned on observations and text, decodable into 3D formats and used as a planner whose futures are converted to actions by a goal-conditioned inverse dynamics module.
- Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models β Proposes VLA-MBPO, a model-based RL framework for VLA finetuning that adapts unified multimodal models for data-efficient world modeling, uses interleaved view decoding for multi-view consistency, and chunk-level branched rollout to mitigate compounding errors.
- VLAW: Iterative Co-Improvement of Vision-Language-Action Policy and World Model β VLAW iteratively improves a VLA policy and an action-conditioned video world model together, using real rollouts to refine the world model which then generates synthetic data to improve the policy. β
39.2% absolute success rate improvement over the base policy
Diffusion Β· flow-matching policies
- DADP: Domain Adaptive Diffusion Policy β Proposes a diffusion policy that disentangles static domain representations from transient dynamics via lagged-context dynamical prediction and injects them through domain-aware adjustment of the diffusion prior and target for zero-shot adaptation to unseen transition dynamics.
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies β Discrete Diffusion VLA is a unified-transformer policy that decodes discretized action chunks via discrete diffusion inside the VLM backbone, with adaptive decoding order and re-masking for error correction. β
96.5% avg. success on LIBERO - FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor Policy β Proposes FocalPolicy, a visuomotor policy combining frequency-optimized chunking with locally anchored flow matching and a foresight composite objective to improve inter-chunk coherence and long-horizon action smoothness.
- From Noise to Control: Parameterized Diffusion Policies β Parameterized Diffusion Policy learns a diffusion policy over a smooth continuous latent manifold where distances reflect trajectory similarity, enabling controllable interpolation and generalization to novel constraints without weight updates.
- From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges β Proposes ResVLA, which reframes action generation as refinement-from-intent by spectrally decomposing motion into deterministic low-frequency anchoring and stochastic high-frequency residuals refined via residual diffusion.
- OMP: One-step Meanflow Policy with Directional Alignment β OMP is a one-step MeanFlow manipulation policy that adds directional velocity alignment and a differential approximation of the JVP operator to enable high-fidelity, real-time single-step generation.
- Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation β Proposes BridgePolicy, which integrates observations directly into the diffusion process via a diffusion-bridge formulation (with multi-modal fusion and a semantic aligner) so sampling starts from an informative observation-conditioned prior rather than random noise. β
52 simulation tasks on three benchmarks and 5 real-world tasks - The Lie We Tell: Correcting the Euclidean Fallacy in Vision Language Action Policies via Score Matching on Tangent Space β Introduces Lie Diffuser Actor, a diffusion VLA policy operating intrinsically on the SE(3) manifold via left-invariant SDEs, tangent-space score prediction, and exponential-map retraction to avoid manifold drift and guarantee equivariance and geodesic optimality. β
On CALVIN ABC->D, LDA improves average task length from 3.06 to 3.30 (+7.8%)
RL for VLA Β· manipulation
- From Abstraction to Instantiation: Learning Behavioral Representation for Vision-Language-Action Model (Oral) β Proposes BehaviorVLA, which learns generalized behavior representations for VLA models using a causal Mamba-based Visuomotor Behavior Encoder and a Phase-conditioned Behavior Decoder to improve robustness under distribution shift.
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual Learning (Oral) β An empirical study finding that large-scale pretrained VLA models resist catastrophic forgetting far better than small policies trained from scratch, with simple Experience Replay sometimes achieving zero forgetting at small replay sizes.
- A Generalist Pair-wise Progress Critic Model for Vision-Language-Action Robots β Proposes VLAC, a unified autoregressive vision-language action-critic model that predicts pair-wise task-progress deltas to provide dense intrinsic rewards and robust actions for real-world reinforcement learning.
- DyGRO-VLA: Cross-Task Scaling of Vision-Language-Action Models via Dynamic Grouped Residual Optimization β A two-stage RL optimization framework for VLAs that captures cross-task latent representations via information-theoretic principles and refines policy optimization through a mixture-of-RL-residuals to improve generalization.
- Focus-Then-Contact: Speeding Up Robotic Contact-Rich Task Learning with Affordance-Guided Real-World Residual Reinforcement Learning β Proposes Focus Then Contact (FTC), a lightweight method combining residual RL base actions with an affordance-guided reward to accelerate convergence of human-in-the-loop real-world RL for contact-rich manipulation tasks.
- HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control β A hierarchical embodied memory framework decoupling control into a high-frequency Executor, a Sentry for working memory, and a Planner for long-term strategy, with a dynamic cross-modal knowledge system supporting add/update/delete operations for long-horizon tasks.
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation β Proposes LAGEA, which converts natural-language failure reflections from vision-language models into temporally grounded, decaying shaped rewards for reinforcement learning in robotic manipulation. β
improvements of 9.0% on random goals, 5.3% on fixed goals, and 17% on fetch tasks - LARA: Latent Action Representation Alignment for Vision-Language-Action Models β Jointly optimizes a latent action model and a VLA model via representation alignment, so action trajectories prevent spurious visual changes while forward-dynamics regularization reduces VLA hallucinations. β
approximately 10% improvement in pre-training scenarios, 5% enhancement for post-training, and 15% gains in latent action model refinement - ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation β ReLAM learns an anticipation model that proposes keypoint-based subgoals from action-free videos to automatically generate dense rewards for hierarchical RL on long-horizon visual manipulation.
- Scaling by Diversified Experience for Vision-Language-Action Models β Introduces SyVLA, which uses an Intention Decoupling algorithm to isolate control features from reasoning context and a similar-sample-guided RL pipeline to stabilize policy updates and improve out-of-distribution generalization.
- SkillNet: Hierarchical Skill Modeling for Compositional Generalization in Vision-Language Action Models β Models skill attributes hierarchically using motion code and the VerbNet framework and regulates a mixture-of-experts structure with transferable skill embeddings as soft constraints to enable compositional generalization across tasks. β
achieves an improvement of performance by 16.0% and 23.9% [zero-shot and few-shot transfer] - Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited Demonstrations β QUEST, a model-based RL-from-demonstrations framework that adaptively switches between exploration and exploitation guided by uncertainty, using intrinsic rewards, ensemble-dynamics planning, and hybrid sampling for sparse-reward manipulation. β
outperforms state-of-the-art methods by 17% on average, with gains increasing to 60% on difficult tasks
Robustness Β· safety Β· interpretability
- Can VLMs Diagnose and Recover from VLA Manipulation Faults? β Introduces the VLA-FixBench fault dataset and FaultEval framework to benchmark 20 VLMs on diagnosing perception/planning/control failures, plus a VLM-VLA collaboration mechanism that localizes deviations and rolls back execution for targeted recovery. β
an idealized feedback loop can improve task success rates by 13% on LIBERO and 35% on real-world robots - Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts β A diagnostic benchmark (LIBERO-Gen) that restructures VLA evaluation into in-distribution, compositional, and domain-generalization tiers to expose spurious invariance and brittleness masked by standard metrics. β
identifies Pi0.5 as the top performer (64.0% in Spatial-CG; 21.2% in Task-CG) - Drift is a Sampling Error: SNR-Aware Power Distributions for Long-Horizon Robotic Planning β CAPS is a training-free inference-time framework that uses power-distribution sampling and SNR-triggered adaptive MCMC search to mitigate instruction drift in long-horizon VLA manipulation.
- Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models β An interpretability study introducing the Interventional Significance Score and Nuisance Mass Ratio metrics to quantify causal influence of visual regions on VLA actions and predict generalization under distribution shift.
- PACT: Self-Evolving Physical Safety Alignment for Diffusion Policies in Embodied Manipulation β A post-training framework (PACT) that aligns pretrained diffusion policies with constraint-feasible regions by distilling constraint gradients via reverse-KL optimization with a constraint-tightening curriculum, without demonstrations or task rewards. β
reduces safety violations by 31.0% on average while improving task success by 30.7% - StableVLA: Towards Robust Vision-Language-Action Models without Extra Data β Proposes an information-bottleneck adapter that filters visual noise to make VLA models robust to unseen real-world visual disturbances without extra data or augmentation, adding fewer than 10M parameters. β
IB-Adapter consistently improves over the baseline by an average of 30% - TRAP: Hijacking VLA CoT-Reasoning via Adversarial Patches β Proposes TRAP, the first targeted adversarial attack framework for CoT-reasoning VLA models, using a physical adversarial patch to corrupt intermediate chain-of-thought reasoning and hijack the robot's manipulation output toward adversary-defined behaviors.
Dexterous Β· bimanual Β· tactile Β· force
- Cross-Tactile Sensor Representation Learning β CTSRL learns sensor-agnostic visuo-tactile representations via a Cross-Sensor Modulator and a two-stage synthetic-then-real self-supervised paradigm for generalization to unseen tactile sensors.
- DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter β Proposes DECO, a decoupled multimodal diffusion transformer that disentangles vision, proprioception, and tactile signals through specialized conditioning pathways with a lightweight tactile adapter, and releases the 50-hour DECO-50 bimanual dexterous tactile dataset. β
72.25% average success rate and a 21% improvement over the baseline - DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation β A curriculum-based algorithm for functional retargeting that learns bimanual dexterous policies to track object states from human hand-object demonstrations using virtual object controllers with decaying strength, plus a simulation benchmark.
- EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video β EgoTactile is a benchmark and diffusion-based method (EgoPressureDiff) that estimates full-hand grasp pressure for everyday objects from egocentric video using a pretrained video diffusion backbone.
- Joint-Space Empowerment as a Theory of Dexterous Motor Coordination β Introduces joint-space empowerment, an information-theoretic principle quantifying an agent's control over its body, used to discover low-dimensional high-empowerment action manifolds for overactuated musculoskeletal systems that improve dexterity and sample efficiency.
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation β A scalable closed-loop simulation framework that uses MLLMs with simulation-in-the-loop verification and five-axis domain randomization to generate diverse synthetic data and a unified evaluation benchmark for dual-arm manipulation. β
3.6x improvement in few-shot real-world transfer (over a 10-demo baseline) and a 2.2x gain in zero-shot generalization - Tabero: Learning Gentle Manipulation with Closed-Loop Force Feedback from Vision, Touch, and Language β Tabero is a vision-tactile-language benchmark and VTLA model with a decoupled force-position command interface executed by a hybrid controller for gentle, force-aware language-conditioned manipulation. β
reduces average grip force by over 70% under gentle instructions - Vision-Language-Action Pretraining from Large-Scale Human Videos β The paper proposes physical instruction tuning that pretrains a VLA from large-scale human-hand videos with perspective spatial alignment and part-level motion tokenization to transfer to dexterous robot manipulation. β
millimeter-level reconstruction accuracy
Benchmarks Β· data Β· simulation
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies (Oral) β A standardized benchmark for evaluating memory in VLA models on long-horizon, history-dependent manipulation, with a taxonomy of temporal/spatial/object/procedural memory and memory-augmented variants on a pi0.5 backbone. β
16 manipulation tasks ... 14 memory-augmented VLA variants built on the pi0.5 backbone - AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation β Introduces AIR-VLA, the first VLA benchmark tailored for aerial manipulation systems, featuring physics-based simulation and a dataset of 3,000 teleoperated demonstrations covering manipulation, object understanding, semantic reasoning, and planning. β
dataset of 3,000 teleoperated demonstrations - CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation β A framework (CaP-Gym + CaP-Bench) for benchmarking code-as-policy coding agents on robot manipulation, plus training-free (CaP-Agent0) and RL-with-verifiable-rewards (CaP-RL) methods that improve performance via test-time computation. β
evaluation across 12 models on 7 simulation tasks - DLO-Lab: Benchmarking Deformable Linear Object Manipulations with Differentiable Physics β Introduces a differentiable simulator and benchmark suite for deformable linear object manipulation that models extensibility, elasticity, bending plasticity, and interactions, plus a DLO agent handling topological complexity via strategic grasping and task decomposition.
- Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation β Identifies a coverage-density trade-off in budget-constrained VLA adaptation and proposes Anchor-Centric Adaptation, a two-stage method that stabilizes a policy skeleton with repeated anchor demonstrations then selectively expands coverage to high-risk boundaries.
- FlatLab: A Unified Methodology Framework and Simulation-Based Benchmark for Robotic Manipulation of Flat Objects β FlatLab pairs a strategy-generator/action-execution framework for manipulating flat objects with a high-fidelity simulation benchmark for diverse rigid and deformable flat objects.
- ManiSoft: Towards Vision-Language Manipulation for Soft Robotics β Introduces ManiSoft, a benchmark and simulator for vision-language manipulation with soft robotic arms, featuring soft-body dynamics, four deformable-control tasks, and 6,300 scenes with expert trajectories. β
6,300 diverse scenes with expert trajectories - OXE-AugE: A Large-Scale Robot Augmentation of OXE for Scaling Cross-Embodiment Policy Learning β Introduces AugE-Toolkit and the OXE-AugE dataset, augmenting Open X-Embodiment with 9 robot embodiments and over 4.4 million trajectories to study how robot augmentation improves cross-embodiment generalist policy learning. β
improving success rates by 24-45% on previously unseen robot-gripper combinations across four real-world manipulation tasks - SafeLab: An Interactive High-Fidelity Benchmark for Embodied Safety in Scientific Robotics β A generative simulation benchmark grounded in a high-fidelity chemistry lab that integrates an LLM task-synthesis engine, an automated expert, and an interactive RL environment to evaluate and improve embodied safety in precision-critical manipulation. β
RL post-training pipeline improves success rates by 37% - Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation β Presents an efficient video augmentation framework that converts simulated VLA videos into realistic training videos via structured-condition extraction, caption rewriting, and conditional video transfer, with diffusion feature-reuse and coreset sampling for scalability. β
improves RDT-1B by 8% on RobotWin 2.0 and boosts pi0 by 5.1% on LIBERO-Plus - VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models β An open-source VLA benchmarking framework structuring tasks across task structure, language, and visual observation dimensions with Safety/Distractor/Extrapolation/Long-Horizon categories and perturbation-based diagnostics. β
11 task suites ... containing 170 total tasks at three difficulty levels (L0-L2)
Specialty (mobile-manip Β· aerial Β· affordance)
- Joint Navigation and Manipulation Planning with 3D Interaction Chains β Introduces 3D Interaction Chains, a unified open-vocabulary mobile-manipulation framework that couples navigation and manipulation planning over a shared 3D feature map, scoring multi-stage waypoint chains by VLM-based feasibility and transition cost.
- Learning AttributeβAffordance Hierarchies in Hyperbolic Space for Open-Vocabulary 3D Object Affordance Grounding β An Attribute-Affordance Hierarchies framework localizes affordance regions on 3D objects from images or text by modeling hierarchical attribute-affordance relationships with hypergraphs and hyperbolic-space concept embeddings.
Trends β what's distinctive about ICML 2026 manipulation
-
VLA is now the default manipulation paradigm at ICML too. 21 of 99 papers are core VLA architecture/backbone work, and VLA framing pervades the efficiency, reasoning, RL, world-model, and benchmark clusters. ICML β historically a methods venue β has fully absorbed the VLA agenda that CoRL/ICLR/CVPR drove over 2024β2025.
-
Efficiency & on-robot deployment is the second-largest cluster (9+ papers). SpecPrune-VLA, EcoVLA, See-What-Matters (GridS), Speedup-Patch, Sparse-ActionGen, Reflex (50 Hz streaming), STEP, and a modelβhardware XPU characterization paper all target real-time/edge inference. Reported speedups cluster around 1.5β4Γ. This is the strongest "make VLAs actually runnable" wave in any 2026 venue so far.
-
Reasoning is going latent and test-time, not textual-CoT. Latent-Reasoning-VLA (
β90% inference latencyvs explicit CoT), LaSTβ, AVA-VLA (early-exit), VLA-ATTC and SCALE (uncertainty-gated test-time compute), Sentinel-VLA (metacognitive monitoring). The explicit language-CoT VLA of 2025 is being replaced by latent reasoning + adaptive compute β same efficiency pressure as trend #2. -
World-model-augmented VLA is a mature cluster (12 papers). Latent-action world models (LAC-WM, MoLA), co-improvement loops (VLAW, model-based RL VLA), structured-4D / semantic-mask prediction, and real-to-sim neural simulators (SoMA, RoboFlow4D). Continues the ICLR-2026 "world model as policy/evaluator" thread, now with explicit VLA coupling.
-
A robustness/skepticism backlash, mirroring CVPR's LIBERO-Plus. Diagnostic benchmarks (LIBERO-Gen "Dismantling the Illusion", VLA-Arena perturbations, RoboMME memory), safety alignment (PACT), the first targeted adversarial attack on VLA chain-of-thought (TRAP), and causal interpretability metrics. The field is auditing the headline-number inflation it produced.
-
Ο0 / Ο0.5 is the universal baseline. HALO, Move-Then-Operate, Sentinel-VLA, VLA-ATTC, RoboMME, "Dismantling the Illusion" and others benchmark directly against Physical Intelligence's Ο0/Ο0.5 β it is now the de-facto reference policy the way OpenVLA was in 2024.
-
ICML's signature: empirical "demystifying" papers. VLANeXt (12 findings β a recipe), From Pixels to Tokens (latent-action supervision study), Demystifying Action Space Design (13,000+ rollouts, 500+ models), Pretrained VLAs Resist Forgetting. Where CoRL ships robots and CVPR ships perception, ICML ships controlled studies of VLA design choices β the most rigorous methodology cluster of the 2026 venues.
-
New benchmark glut (11 papers). RoboTwin 2.0 (bimanual), VLA-Arena, RoboMME (memory), ManiSoft (soft robots), DLO-Lab (deformable linear objects), FlatLab (flat objects), SafeLab (chemistry-lab safety), AIR-VLA (aerial manip), CaP-X (coding agents), OXE-AugE (4.4M-traj OXE augmentation). Benchmarks now specialize by object physics and embodiment rather than competing as general suites.
Cross-venue lineage
ICML 2026 (July) lands after CVPR 2026 (June) and the rebuilt ICLR 2026 (May). Threads that carry through:
- World-model-as-policy/evaluator (ICLR β CVPR GigaBrain/CoWVLA β ICML's 12-paper world-model cluster).
- Latent action representations (ICLR villa-X/UniVLA β ICML XR-1, LARA, LAC-WM, MoLA).
- Efficiency/pruning (ICLR FASTER/SP-VLA β ICML's deployment cluster, now hardware-aware).
- Robustness audits (CVPR LIBERO-Plus β ICML LIBERO-Gen / VLA-Arena / RoboMME).
- RDT2 appears in both the CVPR-2026 and ICML-2026 lists in this wiki β confirm the canonical venue.
β οΈ De-duplication note. Several titles (XR-1, RDT2, Discrete Diffusion VLA, Human-Video VLA pretraining) also have wiki pages tagged to ICLR 2026 or CVPR 2026. ICML, ICLR, and CVPR 2026 overlap in time and some are distinct papers with similar names; others may be the same work. These need a venue-of-record reconciliation pass before per-paper pages are created.
Methodology of this index
- Source of truth:
https://icml.cc/static/virtual/data/icml-2026-orals-posters.json(6,804 events; 6,636 distinct papers; 168 Orals). - Filter: title + keyword match on
vision-language-action / VLA / manipulation / dexterous / grasp / bimanual / humanoid / tactile / affordance / visuomotor / diffusion-policy / flow-policy / imitation / cross-embodiment / world-model+robot, with negative filters for off-topic "manipulation" (image forgery, market/LLM-judge manipulation, graph reasoning) and exclusion of autonomous-driving-only VLAs, pure navigation, and pure learning-theory. - Per-paper verification: each candidate's abstract was fetched from its
icml.cc/virtual/2026/poster/<id>page; summaries and quoted numbers are abstract-derived. 220 candidates β 99 confirmed manipulation papers. - Known gaps: author institutions unavailable (OpenReview guest-blocked); abstracts not yet cross-checked against arXiv full text; venue-of-record overlaps with CVPR/ICLR 2026 not yet reconciled.