ICML 2026 - Heungwoo/research GitHub Wiki

ICML 2026

International Conference on Machine Learning 2026 β€” COEX Convention & Exhibition Center, Seoul, South Korea, July 6–11, 2026 (Jul 6 tutorials/expo Β· Jul 7–9 main conference Β· Jul 10–11 workshops).

Scale: 23,918 submissions β†’ ~6,352 accepted (β‰ˆ26.6%) per the public submission statistics; the official virtual proceedings list 6,636 distinct accepted papers (6,804 schedule events, of which 168 are Orals and the rest Posters). Spotlight/Oral tier β‰ˆ the top 0.7% of submissions.

Robotics / manipulation share: a keyword sweep of the official accepted list surfaced ~220 robotics / embodied-AI candidates, of which 99 are genuinely about robot manipulation (the rest are autonomous-driving VLAs, pure navigation, world-model/RL learning theory, or non-robotic uses of "manipulation"). That makes manipulation β‰ˆ 1.5% of all ICML 2026 papers β€” smaller than at CVPR/ICLR by share, but the largest manipulation presence ICML has ever had, and notably ML-methodology-heavy (systematic studies, recipes, theory) rather than systems-paper-heavy.

Source & method. This index was built against the official virtual data file https://icml.cc/static/virtual/data/icml-2026-orals-posters.json (2026-06 snapshot, 6,804 events). Every paper below was filtered by title+keyword match, then each abstract was fetched from its icml.cc/virtual/2026/poster/<id> page and summarized individually. Quantitative figures in back-ticks are quoted from the paper's own abstract; where an abstract states no number, none is shown (no estimated/invented metrics). OpenReview blocks guest access to ICML 2026 notes, so author institutions are not yet attached. Treat one-line summaries as abstract-derived, pending full-text review.

Surveys hosted in this wiki

  • Per-paper in-depth pages now exist for all 99 manipulation papers (linked from the category list below) β€” each with Problem Β· Method Β· Results Β· Significance and, where the paper has a preprint, the paper's own figures embedded. A few highest-value papers also have long-form reviews: [Review-RoboMME]] and [Action Space: EEF vs Joint.

Orals among the manipulation papers (5)

  • XR-1 β€” Unified Vision-Motion Codes (dual-branch VQ-VAE), 12,000+ real rollouts across 6 embodiments. (also tagged ICLR 2026 in our wiki β€” verify which venue is canonical)
  • From Pixels to Tokens β€” systematic study of latent-action supervision for VLAs; discrete latent action tokens win.
  • From Abstraction to Instantiation (BehaviorVLA) β€” causal Mamba behavior encoder + phase-conditioned decoder for robustness under shift.
  • Pretrained VLAs are Surprisingly Resistant to Forgetting in Continual Learning β€” big pretrained VLAs barely forget; simple Experience Replay can hit zero forgetting.
  • RoboMME β€” standardized benchmark for memory in robotic generalist policies (16 tasks, 14 memory-augmented Ο€0.5 variants).

πŸ… Top-20 technically notable papers (synthesized ranking)

Synthesis of editorial judgment + quantitative signals (GitHub β˜… + arXiv-preprint citations via Semantic Scholar + Oral status), snapshot 2026-06-09. Pure simulators / data-generators are excluded (RoboTwin 2.0, VLA-Arena, CaP-X, OXE-AugE, SoMA, DLO-Lab, FlatLab, ManiSoft, SafeLab, AIR-VLA) β€” but insight / evaluation papers are kept (RoboMME, From Pixels to Tokens, Demystifying Action Space, Pretrained-VLA-Forgetting). β˜…/citation are early preprint-era signals that favor early code-releasers and undercount work with no public repo/arXiv.

# Paper Category β˜… cite Oral Affiliations
1 RDT2 ΒΉ Foundation/scaling 775 15 Tsinghua University (THU-ML)
2 DreamDojo World model 923 42 NVIDIA; HKUST; UC Berkeley; UW; Stanford; KAIST
3 XR-1 ΒΉ Architecture 174 11 β˜… Beijing Innovation Center of Humanoid Robotics (X-Humanoid); Beihang; PKU
4 Discrete Diffusion VLA ΒΉ Architecture 65 64 HKU; Shanghai AI Lab; SJTU; Huawei
5 Being-H0 ΒΉ Human-video pretrain 48 69 Peking University; Renmin University; BeingBeyond
6 DexMachina Dexterous 225 31 Stanford University; NVIDIA
7 VLAC (Progress Critic) RL for VLA 306 β€” Shanghai AI Laboratory (InternRobotics)
8 Latent Reasoning VLA Reasoning 67 7 Tsinghua; PKU; USTC
9 LangForce Analysis/method 65 10 ZGC-EmbodyAI (Zhongguancun Academy)
10 HALO Reasoning β€” 2 HKUST
11 Dual-Stream Diffusion World model β€” 13 KAIST (RLWRLD)
12 DECO Dexterous/tactile 28 0 BAAI; TU Munich
13 See What Matters Efficiency 62 1 University of Sydney
14 SpecPrune-VLA Efficiency β€” 24 SJTU; Infinigence-AI; Shanghai Innovation Institute
15 BehaviorVLA Representation β€” 0 β˜… HIT (Shenzhen); Sun Yat-sen University
16 VLANeXt Recipe/insight 196 5 S-Lab NTU; SYSU; ACE Robotics
17 From Pixels to Tokens Insight 28 0 β˜… Renmin University (KBReasoning)
18 Pretrained VLAs Resist Forgetting Insight β€” 4 β˜… UT Austin; KAIST; Microsoft
19 Demystifying Action Space Insight β€” 1 Tsinghua (IAIR/Wuxi)
20 RoboMME Memory insight 111 6 β˜… University of Michigan; Stanford; Figure AI

ΒΉ Already has a wiki page under another 2026 venue (CVPR/ICLR); these titles are also in the official ICML 2026 accepted list, so venue-of-record needs reconciliation β€” link points to the existing page.

How to read it: the top tier (RDT2, DreamDojo, XR-1, DexMachina) is strong on both axes. Citations elevate Discrete Diffusion VLA (64) and Being-H0 (69) β€” the most-cited methods. Editorial judgment retains low-traction-but-important work: the Oral insight papers (From Pixels to Tokens, Pretrained-VLA-Forgetting, RoboMME) and the latent-reasoning / efficiency methods. 7 of 20 have industry involvement (NVIDIA Γ—2, X-Humanoid, BeingBeyond, Figure AI, Huawei, Infinigence-AI).

Each paper links to a detail page (Problem Β· Method Β· Results Β· Significance Β· Links); page metrics use the same 2026-06-09 snapshot.


Manipulation papers by category

99 papers. Headline numbers in back-ticks are quoted from the abstract.

VLA architecture & backbones

Efficiency Β· pruning Β· deployment

Reasoning Β· CoT Β· test-time scaling

World models for manipulation

Diffusion Β· flow-matching policies

RL for VLA Β· manipulation

Robustness Β· safety Β· interpretability

Dexterous Β· bimanual Β· tactile Β· force

Benchmarks Β· data Β· simulation

Specialty (mobile-manip Β· aerial Β· affordance)


Trends β€” what's distinctive about ICML 2026 manipulation

  1. VLA is now the default manipulation paradigm at ICML too. 21 of 99 papers are core VLA architecture/backbone work, and VLA framing pervades the efficiency, reasoning, RL, world-model, and benchmark clusters. ICML β€” historically a methods venue β€” has fully absorbed the VLA agenda that CoRL/ICLR/CVPR drove over 2024–2025.

  2. Efficiency & on-robot deployment is the second-largest cluster (9+ papers). SpecPrune-VLA, EcoVLA, See-What-Matters (GridS), Speedup-Patch, Sparse-ActionGen, Reflex (50 Hz streaming), STEP, and a model–hardware XPU characterization paper all target real-time/edge inference. Reported speedups cluster around 1.5–4Γ—. This is the strongest "make VLAs actually runnable" wave in any 2026 venue so far.

  3. Reasoning is going latent and test-time, not textual-CoT. Latent-Reasoning-VLA (βˆ’90% inference latency vs explicit CoT), LaSTβ‚€, AVA-VLA (early-exit), VLA-ATTC and SCALE (uncertainty-gated test-time compute), Sentinel-VLA (metacognitive monitoring). The explicit language-CoT VLA of 2025 is being replaced by latent reasoning + adaptive compute β€” same efficiency pressure as trend #2.

  4. World-model-augmented VLA is a mature cluster (12 papers). Latent-action world models (LAC-WM, MoLA), co-improvement loops (VLAW, model-based RL VLA), structured-4D / semantic-mask prediction, and real-to-sim neural simulators (SoMA, RoboFlow4D). Continues the ICLR-2026 "world model as policy/evaluator" thread, now with explicit VLA coupling.

  5. A robustness/skepticism backlash, mirroring CVPR's LIBERO-Plus. Diagnostic benchmarks (LIBERO-Gen "Dismantling the Illusion", VLA-Arena perturbations, RoboMME memory), safety alignment (PACT), the first targeted adversarial attack on VLA chain-of-thought (TRAP), and causal interpretability metrics. The field is auditing the headline-number inflation it produced.

  6. Ο€0 / Ο€0.5 is the universal baseline. HALO, Move-Then-Operate, Sentinel-VLA, VLA-ATTC, RoboMME, "Dismantling the Illusion" and others benchmark directly against Physical Intelligence's Ο€0/Ο€0.5 β€” it is now the de-facto reference policy the way OpenVLA was in 2024.

  7. ICML's signature: empirical "demystifying" papers. VLANeXt (12 findings β†’ a recipe), From Pixels to Tokens (latent-action supervision study), Demystifying Action Space Design (13,000+ rollouts, 500+ models), Pretrained VLAs Resist Forgetting. Where CoRL ships robots and CVPR ships perception, ICML ships controlled studies of VLA design choices β€” the most rigorous methodology cluster of the 2026 venues.

  8. New benchmark glut (11 papers). RoboTwin 2.0 (bimanual), VLA-Arena, RoboMME (memory), ManiSoft (soft robots), DLO-Lab (deformable linear objects), FlatLab (flat objects), SafeLab (chemistry-lab safety), AIR-VLA (aerial manip), CaP-X (coding agents), OXE-AugE (4.4M-traj OXE augmentation). Benchmarks now specialize by object physics and embodiment rather than competing as general suites.

Cross-venue lineage

ICML 2026 (July) lands after CVPR 2026 (June) and the rebuilt ICLR 2026 (May). Threads that carry through:

  • World-model-as-policy/evaluator (ICLR β†’ CVPR GigaBrain/CoWVLA β†’ ICML's 12-paper world-model cluster).
  • Latent action representations (ICLR villa-X/UniVLA β†’ ICML XR-1, LARA, LAC-WM, MoLA).
  • Efficiency/pruning (ICLR FASTER/SP-VLA β†’ ICML's deployment cluster, now hardware-aware).
  • Robustness audits (CVPR LIBERO-Plus β†’ ICML LIBERO-Gen / VLA-Arena / RoboMME).
  • RDT2 appears in both the CVPR-2026 and ICML-2026 lists in this wiki β€” confirm the canonical venue.

⚠️ De-duplication note. Several titles (XR-1, RDT2, Discrete Diffusion VLA, Human-Video VLA pretraining) also have wiki pages tagged to ICLR 2026 or CVPR 2026. ICML, ICLR, and CVPR 2026 overlap in time and some are distinct papers with similar names; others may be the same work. These need a venue-of-record reconciliation pass before per-paper pages are created.

Methodology of this index

  • Source of truth: https://icml.cc/static/virtual/data/icml-2026-orals-posters.json (6,804 events; 6,636 distinct papers; 168 Orals).
  • Filter: title + keyword match on vision-language-action / VLA / manipulation / dexterous / grasp / bimanual / humanoid / tactile / affordance / visuomotor / diffusion-policy / flow-policy / imitation / cross-embodiment / world-model+robot, with negative filters for off-topic "manipulation" (image forgery, market/LLM-judge manipulation, graph reasoning) and exclusion of autonomous-driving-only VLAs, pure navigation, and pure learning-theory.
  • Per-paper verification: each candidate's abstract was fetched from its icml.cc/virtual/2026/poster/<id> page; summaries and quoted numbers are abstract-derived. 220 candidates β†’ 99 confirmed manipulation papers.
  • Known gaps: author institutions unavailable (OpenReview guest-blocked); abstracts not yet cross-checked against arXiv full text; venue-of-record overlaps with CVPR/ICLR 2026 not yet reconciled.

← Back to ICML Β· Home