ICLR 2026 - Heungwoo/research GitHub Wiki
ICLR 2026
ICLR 2026 received 19,525 valid submissions and accepted 5,355 papers (225 orals; 27.4% acceptance β a three-year low). 779 were desk-rejected and 5,042 withdrawn, so 13,763 submissions received an accept/reject decision (8,408 rejected). VLA / robot manipulation is one of the largest growth areas β ~210 robot- or VLA-related accepted papers by keyword count, vs. <100 in 2025.
This index was rebuilt against the official accepted list at https://iclr.cc/virtual/2026/papers.html (data file: /static/virtual/data/iclr-2026-orals-posters.json, 2026-05-09 snapshot). Items listed under "Confirmed at ICLR 2026" have been individually verified against the OpenReview decision; companion technical reports and arXiv-only preprints that were previously cross-listed are now relocated to a separate section.
Surveys hosted in this wiki
- VLA & Manipulation Survey β five categories, 8 trend analyses, ~37 paper summaries.
Confirmed at ICLR 2026 β wiki pages
These wiki pages map 1-to-1 to a paper in the official ICLR 2026 accepted list. The bullet shows the official accepted title when it differs from our shorthand.
Architecture Β· Diffusion / discrete-diffusion action heads
- Unified Diffusion VLA β Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
Architecture Β· Memory Β· Cross-embodiment Β· MoE
- HAMLET β HAMLET: Switch Your Vision-Language-Action Model into a History-Aware Policy
- MemoryVLA β MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- X-VLA β X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- UniVLA β Unified Vision-Language-Action Model
- WholeBodyVLA β WholeBodyVLA: Towards Unified Latent VLA for Whole-body Loco-manipulation Control
- Cosmos Policy β Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- VLM4VLA β VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
Tokenizers Β· Quantization Β· Efficiency
- FASTER β FASTer: Toward Powerful and Efficient Autoregressive VisionβLanguageβAction Models with Learnable Action Tokenizer and Block-wise Decoding
- AutoQVLA β QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
Data Β· Benchmarks Β· Simulation
- EgoDex β EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- RoboCasa365 β RoboCasa365: A Large-Scale Simulation Framework for Training and Benchmarking Generalist Robots
- RoboArena β β RobotArena β: Scalable Robot Benchmarking via Real-to-Sim Translation
- WorldGym β WorldGym: World Model as An Environment for Policy Evaluation
Training approaches Β· Co-training Β· Reasoning Β· Composition
- Actions as Language β Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- InstructVLA β Vision-Language-Action Instruction Tuning: From Understanding to Manipulation
- Hybrid Training β Hybrid Training for Vision-Language-Action Models
- Embodied-R1 β Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Compose Your Policies β Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
- Ctrl-World β Ctrl-World: A Controllable Generative World Model for Robot Manipulation
Dexterous manipulation
- DexNDM β DexNDM: Closing the Reality Gap for Dexterous In-Hand Rotation via Joint-Wise Neural Dynamics Model
- UniHM β UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
- RFS β RFS: Reinforcement learning with Residual flow steering for dexterous manipulation
RL for VLA β see also RL
- PLD β Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- Stage-Aware RL β SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
- VITA β VITA: Vision-to-Action Flow Matching Policy
- SimpleVLA-RL β SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Confirmed at ICLR 2026 β newly added wiki pages (May 2026 batch)
These pages were created in the May 2026 wiki rebuild, one per accepted paper. Each page follows the standard template (mermaid pipeline Β· Problem Β· Method Β· Results Β· Significance Β· Links).
Architecture Β· VLA backbones
- OneTwoVLA β Unified VLA with Adaptive Reasoning (System-1/System-2 switching inside one model)
- HybridVLA β Collaborative Diffusion and Autoregression in a Unified VLA
- Vlaser β VLA with Synergistic Embodied Reasoning (Vlaser-6M reasoning corpus)
- villa-X β Enhancing Latent Action Modeling in VLA (Vision-Language-Latent-Action)
- MetaVLA β Unified Meta Co-Training for Efficient Embodied Adaptation (post-training, backbone-agnostic)
- PixelVLA β Advancing Pixel-level Understanding in VLA (multiscale pixel-aware encoder + visual prompting)
- TwinVLA β Twin Single-Arm VLAs for Bimanual Manipulation (no bimanual pre-training)
- VER β Vision Expert Transformer via Foundation Distillation and Dynamic Routing (<0.4%-param router)
Spatial / 3D grounding for VLA
- Spatial Forcing β Implicit alignment to a 3D foundation model
- Spatially Guided Training β SP-VLA two-stage spatial pretraining (SimplerEnv 66 β 84%)
- From Spatial to Actions (FALCON) β 3D tokens injected at the action head
- EquAct β SE(3)-equivariant transformer + iFiLM, 18 RLBench tasks
- PA3FF β Part-aware dense 3D feature field + Part-Aware Diffusion Policy
- Seeing Across Views β MV-RoboBench (1.7k QA, 8 subtasks)
Efficiency Β· Pruning Β· Acceleration
- SP-VLA β Joint scheduling + token pruning, 1.5Γ/2.4Γ speedup
- Action-aware Dynamic Pruning β Phase-adaptive token pruning, 1.35Γ on OpenVLA-OFT
- Verifier-free Test-Time Sampling β MG-Select via KL-to-masked-reference
- Block-wise Adaptive Caching for Accelerating Diffusion Policy
- Real-Time Robot Execution with Masked Action Chunking
Robustness Β· Adaptation
- On Robustness of VLA β RobustVLA (17 perturbations / 4 modalities, +12.6% on Οβ)
- Robust Fine-tuning via Parameter Merging β pretrained β finetuned weight interpolation
- Align-Then-Steer β VAE-aligned latents, +9.8% sim / +32% real cross-embodiment
- When would Vision-Proprioception Policies Fail in Robotic Manipulation?
World models Β· Video for actions
- WMPO β World Model-based Policy Optimization for VLA
- Genie Envisioner β Unified World Foundation Platform
- Vid2World β Video Diffusion β Interactive World Models
- ViPRA β Video Prediction for Robot Actions (+16% SIMPLER, +13% real, 22 Hz)
- Geometry-aware 4D Video Generation β 4D video gen for manipulation
- Sim2Real VLA β zero-shot synthesized-skills transfer
- Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
- Sparse Imagination for Efficient Visual World Model Planning
- Test-Time Mixture of World Models for Embodied Agents
Embodied reasoning Β· Planning Β· CoT
- From Seeing to Doing β Bridging reasoning and decision (8 spatial benchmarks)
- OmniEVA β Embodied Versatile Planner via Task-Adaptive 3D-Grounded reasoning
- VLMgineer β VLMs as Robotic Toolsmiths
- Self-Refining VLM (ARMOR) β Robotic failure detection and reasoning
- Cortical Policy β Dual-Stream View Transformer for manipulation
- RoboInter β Holistic Intermediate Representation Suite
- Policy Contrastive Decoding for Robotic Foundation Models
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Self-Improving Loops for Visual Robotic Planning
Bimanual Β· Mobile Β· Dexterous
- VLBiMan β Vision-Language Anchored Bimanual from one-shot demos
- MoMaGen β Bimanual mobile manipulation demo generation under constraints
- DemoGrasp β Universal dexterous grasping from a single demonstration
- DexMove β Tactile-Guided Non-Prehensile Manipulation with dexterous hands
- D-REX β Differentiable Real-to-Sim-to-Real Engine for Dexterous Grasping
- RoboPARA β RoboPARA: Dual-Arm Robot Planning with Parallel Allocation and Recomposition Across Tasks (LLM dependency-graph dual-arm parallelism, X-DAPT dataset)
- CLAP (Coarse-to-Fine Keypoints) β Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints (VLM 3D-keypoint grounding; +12% on GemBench with 1/5 trajectories)
- SMP (MoE Diffusion Skills) β Abstracting Robot Manipulation Skills via Mixture-of-Experts Diffusion Policies (orthogonal skill basis + sticky routing)
- PF-DAG β Primary-Fine Decoupling for Action Generation in Robotic Imitation (discrete-mode + mode-conditioned MeanFlow; 56 tasks)
- OmniReset (Emergent Dexterity) β Emergent Dexterity via Diverse Resets and Large-Scale RL (programmatic diverse resets enable on-policy RL with no demos/curriculum; zero-shot real transfer)
- House of Dextra β Cross-Embodied Co-Design for Dexterous Hands (joint morphology+policy co-design; design-to-deploy a new hand in <24h)
Humanoid Β· Whole-body Β· Loco-manipulation
- BFM-Zero β Promptable Behavioral Foundation Model for Humanoid Control via unsupervised RL
- HWC-Loco β Hierarchical Whole-Body Control for robust humanoid locomotion
- From Language to Locomotion β Retargeting-free humanoid control via motion latent guidance
- LIFT (Pretrain-Finetune Humanoid) β Bridging Large-Scale Pretraining and Efficient Finetuning for Humanoid Control (off-policy SAC pretrain + physics-informed world-model finetune)
- HVD β Hierarchical Value-Decomposed Offline RL for Whole-Body Control (kinematic-tree Q-value decomposition; WB-50 dataset)
Navigation Β· Embodied agents
- Embodied Navigation Foundation Model (NavFoM) β 7-benchmark navigation FM
- REI-Bench β Can embodied agents understand vague instructions? (36.9% degradation finding)
- Ground Slow, Move Fast Β· JanusVLN Β· AutoFly Β· CompassNav Β· CE-Nav Β· From Seeing to Experiencing Β· Towards Physically Executable 3D Gaussian for Embodied Navigation
RL Β· Flow policies
- Flow Matching Policy Gradients (FPO) β RL recipe for flow-matching policies
- Translating Flow to Policy (HinFlow) β Hindsight online imitation, >2Γ improvement
- Guided Flow Policy β Offline RL with high-value flow guidance (144 tasks)
- Batch Online RL β What matters for self-improving robot learning from autonomous data (Q-functions + implicit extraction + diffusion)
- Learning from Constrained Demonstrators β Robot learns a better policy than its constrained (e.g. joystick) demonstrator
Data Β· Imitation Β· Benchmarks
- DataMIL β Datamodel-based selection for imitation learning
- Reflective DDVLA β Discrete Diffusion for Reflective VLA in Autonomous Driving β the only "discrete-diffusion VLA" actually at ICLR 2026 (autonomous-driving variant; the manipulation DDVLA is NeurIPS 2025 / arXiv)
- ManipEvalAgent Β· AutoBio Β· ArtVIP Β· Action Chunking & Exploratory Data Β· Demystifying Diffusion Policies Β· MIKASA (Memory, Benchmark & Robots) Β· Cross-Embodiment Offline RL Β· DeFI (Disentangled Fwd/Inv) Β· MemER (Experience Retrieval)
Sensing Β· Tactile
- AnyTouch 2 β General optical tactile representation for dynamic, force-aware perception (ToucHD dataset) Β· SpikePingpong β Spike-vision fast-slow system for high-precision robot table tennis
Companion preprints / technical reports β referenced in the wiki, not at ICLR 2026
These wiki pages are still useful as standalone analyses, but they describe technical reports, blog posts, or arXiv-only preprints that are not in the ICLR 2026 proceedings. Renamed from "ICLR 2026" framing.
Physical Intelligence (PI) β production reference series
- Ο0.6 β PI technical report (Nov 2025)
- Ο0.7 β PI technical report (Apr 2026)
- Ο*0.6 + RECAP β PI technical report (RL from experience)
- RL Tokens (RLT) β PI blog post (online RL on contact-rich tasks)
- π Ο series evolution β Ο0 β Ο0.7 model/data/training side-by-side
Toyota Research Institute / Levine lab / Berkeley β Feb 2026 arXiv preprints
- LBM Co-training Study (TRI) β A Systematic Study of Data Modalities and Strategies for Co-training Large Behavior Models (arXiv 2602.01067, Feb 2026 β submitted post-ICLR-2026 deadline)
- AsyncVLA (Levine lab) β Hirose / Glossop / Shah / Levine, Feb 2026 arXiv
- Steerable Policies (Berkeley+PI) β Chen / Glossop / Driess / Pertsch / Levine et al., Feb 2026 arXiv
Wiki pages with venue still being verified
These pages exist in the wiki but the corresponding paper title was not found in the May 2026 ICLR-2026 accepted list. They may be at NeurIPS 2025, CVPR 2026, IROS 2026, or arXiv-only. Treat as preprint-quality references until confirmed:
- Discrete Diffusion VLA β only an autonomous-driving discrete-diffusion VLA was accepted; the manipulation DDVLA is likely NeurIPS 2025 / arXiv
- dVLA
- DIVA & Fast-dVLA
- HiMoE-VLA
- HyperVLA
- Human-Video Pretraining
- OmniSAT
- VLA-RFT
- XR-1
Methodology of this index
- Source of truth:
https://iclr.cc/static/virtual/data/iclr-2026-orals-posters.json(the JSON the official virtual site uses to populate/virtual/2026/papers.html). - Filter: keyword match on
name,keywords, andtopicforVLA,vision-language-action,manipulation,dexterous,humanoid,whole-body,bimanual,embodied,cross-embodiment,gripper,grasp,imitation learning,behavior cloning,flow-matching policy,diffusion policy,world model + robot/manipulation/action, etc., with negative filters for off-topic uses (image manipulation,language-model fine-tuning, etc.). - Coverage: 211 papers identified as VLA / robot-manipulation / embodied AI / dexterous / humanoid / navigation. The "we have not yet reviewed" section above lists ~75 of the most architecturally or empirically relevant of these.
- Verification: every wiki page in this index was individually substring-matched against the JSON title set.
β Back to ICLR