CoRL 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki

CoRL 2026 β€” VLA & Manipulation Survey (preliminary, community-sourced)

Venue: 10th Conference on Robot Learning (CoRL) 2026 Β· Austin, TX (JW Marriott) Β· Nov 9–12, 2026 (workshops Nov 9, main Nov 10–12). Scale: 687 papers accepted Β· 32.8% acceptance Β· 2,094 active submissions (decisions announced; official per-paper program on OpenReview not yet public). ⚠️ Sourcing caveat. There is no official accepted-papers list yet. This page is a preliminary, community-sourced selection of robot-manipulation papers whose CoRL 2026 acceptance is either explicitly stated by the authors (βœ…) or reported/announced but not yet independently verified (⚠️). Every listed paper is a real arXiv preprint; the acceptance is what carries the caveat. To be replaced by a full survey once the official program publishes. Legend: βœ… = author-confirmed CoRL 2026 Β· ⚠️ = reported/likely, unverified.


1. At a glance

  • 687 accepted / 2,094 submitted (32.8%) β€” CoRL stays selective and single-track. The manipulation/VLA/humanoid/WAM slice is large but not yet enumerable pre-program.
  • ~47 author-confirmed in-scope papers (Β§2) across eight themes β€” World-Action / latent-dynamics (SG-WAM, Sensitivity-Shaping, K-UBM, DAP, PhysCoRe, WHIRL, StressDream), VLA (FiberTune, VLA-Feedback, S2, CounterAlign, Real-Time-AR, AFP, MessyMem, Efficient-VLA, MolmoAct2, MolmoBOT, SAE-VLA), dexterous & visuo-tactile (Dex-X, Touch2Trace, DexFLEX, HUGS, TeleDexter, MiTaS, FTP-1), whole-body humanoid (OmniContact, AdaPT, HiPHI, Perceptive-BFM, MeshMimic, CHIP), demonstration-efficient RL (FlashRFCL, SDPG, UniLab, TOPReward), policy adaptation / sim-to-real (EmbodiSteer, FlowDAgger, TAM, VLS), data-gen, benchmarks & human-video (HuRo, RoboReel, SPARC, WireCraft, LUCID, WIYH), and safety/adversarial (Trajectory-Redirection, BarrierFormer) β€” plus reported (⚠️) StellaVLA/Choice-Policies/Weave/Look-Before-You-Move/Disentangling/MAGMA-GEN and adjacent driving/HRI/locomotion papers.
  • What they signal (Β§3): all the 2026 threads the wiki tracks are present, and several sharpen: (a) prediction is being used to monitor execution, not just generate it (K-UBM's flow discrepancy β†’ replan; WHIRL predicts human takeover β†’ avoid; DAP's bidirectional obs↔action coupling); (b) VLAs are going reactive inside the action chunk (VLA-Feedback's final-step correction; Real-Time-AR); (c) the visual interface is being narrowed (S2 evidence budgets, AFP anti-shortcut foveation, SDPG end-to-end pixels); and (d) human video is being robotized at scale (HuRo 630K episodes) alongside a humanoid mocap substrate (HiPHI).
  • Not in scope here: this is manipulation-centric; CoRL's pure locomotion/navigation/driving/theory tracks are excluded (SafeDriveVLA, BridgeSim, VLAlert, BPF, AI-Coaching, Mind-the-Phase are listed only as adjacent reference in Β§2.9).

2. Papers by theme (βœ… confirmed Β· ⚠️ reported)

2.1 World-Action Models / world-model & latent-dynamics policies

  • βœ… SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space (2608.01397) β€” aligns world modeling with action generation by predicting future dynamics in a geometry-aware, policy-derived representation space (not pixels). A 0.9B model with no large-scale embodied pretraining hits LIBERO 98.5% / LIBERO-Plus 73% and beats baselines in ID and OOD real-world eval. β†’ the latent/representation-space WAM corner (World Models, VLA Hybrid Architectures).
  • βœ… Sensitivity Shaping for Latent Modeling (2606.14585, UCSD β€” Yu, Li, Zhang, Christensen, Gao) β€” generative dynamics models enable planning, but safe deployment needs to detect policy-induced OOD transitions; existing surrogates fail when dynamics are locally insensitive to critical action choices. Introduces support-conditioned control-sensitivity regularization so the learned dynamics respond sensitively to control changes in high-support regions β†’ a world-model reliability contribution (planning-with-learned-dynamics safety), cf. World Models Β§6 hallucination/OOD.
  • βœ… K-UBM: Going with the Flow β€” Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity (2602.07413, Georgia Tech β€” Yunhai Han, … Harish Ravichandar) β€” treats manipulation as the coupled evolution of robot action + environmental visual flow in a shared latent space governed by a structured (Koopman) linear system. Instead of repeatedly predicting short action chunks, it rolls out the latent dynamical system as a "pseudo planner" for coherent long-horizon behavior + very fast inference. Crucially it predicts future visual flow and uses the predicted-vs-observed flow discrepancy as a runtime monitor β†’ event-triggered replanning. 7 sim + 4 real dexterous tasks: competitive success with smoother execution, occlusion robustness, efficient inference. β†’ prediction-as-execution-monitor (WAM used to supervise rather than generate the action path), cf. World Models, Dexterous Manipulation.
  • βœ… DAP: Dynamics-Aligned Flow Matching Policy (2510.27114 β€” Cho, Lee, … Li Zhao) β€” jointly models future observations and actions, but unlike a plain World Action Model it makes the coupling bidirectional: an architecture for action-conditioned future-observation prediction and future-conditioned action generation, with the policy and dynamics models giving each other mutual corrective feedback during generation. Learning both directions together improves the sample-efficiency and perturbation-robustness of diffusion/flow-matching imitation policies. β†’ the bidirectional-coupling variant of the WAM+policy design (cf. its NeurIPS-2026 companion Why Latent Actions Fail on future-leakage in latent-action models).
  • βœ… PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics (2607.20653 β€” Tao, … Lu Gan, Chen) β€” keeps a physics model at the core (differentiable Material Point Method) while gaining learned generalization, instead of per-object material fitting (slow) or end-to-end dynamics (drifts unphysically). Material-from-Motion infers per-particle material + an unsupervised confidence (concentrates on what actually deforms β†’ where to probe next) in one forward pass; Residual-from-Dynamics injects a velocity residual inside the MPM cycle to close sim-to-real while preserving physical structure. Beats SOTA physics/learned baselines on real elastic & elastoplastic deformation β€” a deformable-object dynamics backend for planning. β†’ the physics-structured corner of World Models.
  • βœ… WHIRL: Intervention-Aware World Models with Real-World RL for Dexterous Manipulation (2609.06009 β€” Xu, Zhang, Feng, … Ajoudani, Renjing Xu) β€” turns one-off human takeovers into predictive risk signals: a latent world model with the usual dynamics/reward/termination heads plus a per-state intervention-probability head that predicts where a human would take over, then steers the policy away from intervention-prone states. On five real tasks with a 16-DoF LEAP Hand, +15–30 pts autonomous success and up to βˆ’84% operator-controlled training steps. β†’ another world-model-as-monitor instance (predict-to-avoid), cf. K-UBM above and World Models; safety-aware real-world RL.
  • βœ… StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement (2606.00267, CMU IntentLab) β€” don't ask a video WM only for the most-likely future; optimize the diffusion WM's initial noise to steer imagination toward high-impact yet plausible outcomes specified by text at inference (e.g. task failures), exposing where the policy breaks β€” then train on those failures. Works with SOTA driving + manipulation video WMs. β†’ world-model-as-stress-tester for evaluation/robustness; World Models, VLA Evaluation.

2.2 Vision-Language-Action models

  • βœ… FiberTune: Preserving Action-Fiber Visual Residuals in VLA Fine-Tuning (2606.08653) β€” action-supervised fine-tuning constrains only action-changing directions and lets action-equivalent visual structure collapse; FiberTune adds a training-time objective (online action probe β†’ filter action-predictive directions β†’ align residuals to a frozen visual teacher + rank regularization), no inference overhead. Gains on Ο€0.5 and OpenVLA-OFT (e.g. +10.7 pp SR(5) on CALVIN ABCβ†’D; SO-101 72.7%β†’78.1%). β†’ the VLA-robustness/fine-tuning thread.
  • ⚠️ StellaVLA: In-Context Structured Demonstration for Generalizable VLA (2608.11671) β€” conditions on one retrieved demo auto-converted offline into a structured demo (task plan + sub-goals + verbalized 3D motion), so the policy reasons rather than mimics pixels; transferable across real-robot / human-hand / XR demos. Reported #1 on VLA-Arena (0.63 vs 0.44/0.22). β†’ In-Context Imitation.
  • βœ… VLA-Feedback β€” "Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs" (2609.21022) β€” VLAs execute a pre-generated action chunk open-loop, so a target that moves mid-chunk is missed. A two-timescale design pairs low-frequency diffusion planning with high-frequency visual feedback: rather than fully denoising the chunk before execution, it keeps the final denoising step as a lightweight feedback interface, correcting each action against the latest observation before it fires ("plan ahead, react in the moment"). Matches GR00T on static LIBERO; lifts dynamic sim 27.5%β†’85.0% and real-robot 51%β†’73%. β†’ Real-Time Execution (reactive-during-chunk).
  • βœ… S2: See Less, Specify More β€” Visual Evidence Budgets for Generalizable VLAs (2606.02735, AIRoA β€” Wu, Matsushima, Ota) β€” a VLA must infer what to do and where to look from a coarse goal + full frame, on a fraction of a VLM's data. See Less imposes an explicit visual-evidence budget (act from task-sufficient evidence, learned straight from the control objective β€” no masks/region annotations); Specify More keeps the original instruction as a stable goal while relabeling trajectories into refined sub-task language. Showing the policy less of the image improves generalization. β†’ Independent Visual Representation / VLA generalization.
  • βœ… CounterAlign: Counterfactual Supervision for VLA (2608.21740, AIRoA / Tokyo Tech β€” Kondoh, Ota, Kanezaki, Wu) β€” relabels existing demos into counterfactual (instruction, action) mismatch pairs; a discriminator learns to catch the mismatch β†’ an instruction-grounded reward for offline RL (IQL + advantage-weighted flow matching). No new data, no hand labels; improves robustness to object-position/task perturbation on LIBERO-PRO and on the real TX-G2. β†’ Multi-Task VLA (instruction grounding), RL for VLA.
  • βœ… Real-Time Execution with Autoregressive Policies (2606.13355, KIST + SNU + Google Research β€” Lee, Park, You, Caciularu, Szpektor, Lim, Youngjae Yu) β€” real-time robotics has favored diffusion/flow policies because autoregressive ones have rollout-speed bottlenecks in synchronous inference. Shows AR policies can run real-time by adjusting the tokenization horizon + constrained decoding to guarantee strict latency bounds, which unlocks multi-trajectory decoding. Across sim + real, the AR policy beats equivalent flow-matching counterparts with faster completion β€” while keeping AR's native edge in convergence and instruction-following. β†’ Real-Time Execution (async, anytime decoding).
  • βœ… Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models (2607.10655) β€” standard fine-tuning supervises actions only, so a policy latches onto background/lighting/co-occurring objects β€” shortcut learning β€” and collapses under distractors (Ο€0.5 real-robot 60%β†’27% once distractors appear). AFP is a lightweight, policy-agnostic module (same RGB + instruction as the policy) predicting a continuous fovea-like mask over task-relevant objects + end-effector, used as an auxiliary grounding loss on image-token attention during fine-tuning (PCGrad projection removes components conflicting with the action gradient). Architecture unchanged; not in the inference loop. MimicGen (4 models Γ— 8 tasks): under-distractor success 39%β†’66%, never hurts clean; real Ο€0.5 27%β†’53%. Ships a 786-episode mask dataset. β†’ the "budget/where-to-look" thread with S2 above; Multi-Task VLA robustness.
  • ⚠️ Look Before You Move: Agent-Interface-Aligned Supervision for Open-World Mobile Manipulation (HKUST-GZ; author-announced, code/model "coming soon" β€” no preprint located yet) β€” agent-interface-aligned supervision for open-world mobile manipulation. β†’ mobile-manipulation VLA (listed on author announcement; verify against preprint when posted).
  • βœ… MessyMem: Learning-from-Doing Memory for Mobile Manipulation (2609.15976, Stanford IPRL β€” Muckelroy, S., Zhao, Bohg, Ho; project) β€” persistent memory for mobile manipulators that combines a spatially-grounded 3D scene graph + interaction-revealed properties (VLM records what actions expose, e.g. "drawer locked") + linked keyframes for fine-grained appearance recall; a new task retrieves the most relevant visual memories alongside the scene graph for a VLM planner. Over 25 consecutive tasks / 3+ hours, retrieves evidence from thousands of keyframes (>1 h old) β†’ 80.0% progress, +28.9 pts over the best external baseline; real TidyBot++ deployment with adaptation to environment change. β†’ the mobile-manipulation memory frontier, VLA Memory.
  • βœ… What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency (2609.13984, Li Auto + collaborators) β€” a T5-style empirical sweep rebuilding VLAs 60+ times under latency-paired conditions (fixed SigLIP2 + Qwen2.5 backbones, swept action-head design & module scale, each paired with measured on-device latency, real-robot validated). Headline: action-head performance is governed mainly by initialization, not decoder architecture / loss / inference budget β€” copying the last transformer layers of the language backbone into the head is the single largest lever, at no latency cost, and the only axis that helps at every scale. β†’ a design/efficiency reference for VLA Architectures, Real-Time Execution.
  • ⚠️ Disentangling Spurious Correlations in VLA Models by Predicting Domain-Invariant Latent Lookahead (SNU AAIG β€” J. Kim, E. Kim, Gi-Cheon Kang, Byoung-Tak Zhang; author-announced, preprint not yet located) β€” VLAs pick up undesired dependencies between task-irrelevant visual cues (lighting, camera perspective) and actions; this policy predicts a domain-invariant future latent that separates task-relevant signal from visual variation, disentangling spurious correlations to improve robustness under visual distribution shift. β†’ the anti-shortcut robustness thread with AFP/CounterAlign above (verify against preprint when posted).
  • βœ… MolmoAct2: Action Reasoning Models for Real-World Deployment (2605.02881, Allen AI) β€” a fully open action-reasoning VLA stack: MolmoER spatial/embodied-reasoning VLM backbone (3.3M-sample corpus, specialize-then-rehearse), OpenFAST open action tokenizer (millions of trajectories, 5 embodiments), and MolmoThink adaptive-depth reasoning (re-predicts depth tokens only for changed regions β†’ low latency). Across 7 sim+real benchmarks beats Ο€0.5; MolmoER surpasses GPT-5 / Gemini-Robotics-ER-1.5 on 13 embodied-reasoning benchmarks. β†’ open action-reasoning VLA, VLA Architectures.
  • βœ… MolmoBOT: Large-Scale Simulation Enables Zero-Shot Manipulation (2603.16861, Allen AI) β€” tests whether enough simulation diversity lets manipulation transfer to the real world with no real-world training data. Fully open procedural pipeline (MolmoBot-Engine / MolmoSpaces) releasing 1.8M expert trajectories; on tabletop pick-and-place 79.2% real-world vs Ο€0.5's 39.2%. β†’ the sim-only-transfer thesis, cf. World Models and sim-to-real (Β§2.6).
  • βœ… Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models (2603.19183, Stanford β€” Swann, McGranahan, Buurmeijer, Kennedy, Schwager) β€” trains SAEs on a VLA's hidden activations to find features corresponding to motion primitives + semantic concepts, some general across episodes and causally steerable; validated by steering on LIBERO and real DROID hardware. β†’ mechanistic-interpretability for VLAs (rare inside-the-model study).

2.3 Dexterous & visuo-tactile manipulation

  • βœ… Dex-X: Learning Visual-Tactile Dexterous Manipulation from Human Videos with Simulated Interaction (2609.07747) β€” reconstructs hand-object interaction in simulation where physically-grounded contact gives tactile supervision; teacher 65.9% (6 tasks, sim) distilled to a visual-tactile policy: 93% real cube-picking, 53% table-cleaning. β†’ egocentric-video pretraining Γ— tactile Γ— data pyramid.
  • βœ… Touch2Trace: Tactile-Driven Imitation Learning for Dexterous Cable Tracing (2609.15921) β€” tactile-driven IL for dexterous cable tracing on a real hand: 93% success, zero-shot to unseen cables/routing. β†’ tactile-first manipulation.
  • βœ… DexFLEX: Contact-Aware Foundation Controller for Command-Guided Dexterity (project, UCSD β€” Wan, Lai, Hou, Zhou, Fu, Christensen, Su) β€” turns upstream fingertip-motion drafts into contact-consistent joint commands: treats a draft as evidence about intent (not a trajectory to copy), proposes short-horizon motion chunks from the tactile-proprioceptive state, predicts their contact consequences, and selects the candidate that follows the command while preserving future contact stability. A Draft–Dream–Select inference loop mixes pure-prior proposals (recovery) with draft-seeded ones (responsiveness). β†’ the "servo on contact" dexterity line (cf. DexterityGen, T-Rex).
  • βœ… HUGS: Guiding Unified Dexterous Grasp Synthesis Across Modes and Scales via Learned Human Priors (2607.04554; project) β€” instead of retargeting human demos, learns an object-conditioned human prior that guides force-closure-aware optimization, giving one synthesizer across contact modes (two-finger β†’ bimanual) and object scales (2–30 cm). Synthesizes 3.2M grasps over 157K scenes; policies trained on it auto-select the contact mode in the real world (screws β†’ large boxes). β†’ dexterous grasp data generation, Dexterous-Hand Data Pyramid.
  • βœ… TeleDexter: Towards Human-Level Dexterous Teleoperation (2607.11481) β€” a hand-object co-tracking controller mapping operator intent into learned low-level contact execution, trained on co-tracking subgoals from human reference motion with a hybrid sparse-subgoal + dense-tracking reward (single-stage RL, zero-shot to real via action masking + domain randomization). 7 hard tasks (reorientation, long-horizon tool use) on two hands: 75% avg where baselines fail. β†’ high-quality dexterous data at the teleop layer of the data pyramid.
  • βœ… MiTaS: Multi-Resolution Tactile Imitation Learning for Contact-Rich Manipulation (2606.06281; project) β€” instead of choosing one tactile sensor, fuses heterogeneous sensors at different temporal resolutions: RGB + vision-based GelSight Mini + high-frequency event-based Evetac, via modality-specific conv stems + a transformer fusion module, conditioning a flow-matching policy. 5 real contact-rich tasks: 80% vs visual-tactile 54% / vision-only 31%; multi-tactile co-training helps even when Evetac is absent at eval. β†’ the sensor-fusion frontier of Tactile VLA.
  • βœ… FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors (2606.13102) β€” the first generalist foundation tactile policy: heterogeneous encoders project image-/array-/state-based tactile signals into unified morphology-aware latent tokens modeled by a shared tactile transformer expert. Pretrained on ~3,000 h across 26 sources / 21 sensors (human + robot), and transfers touch beyond sensors seen in pretraining β€” across sensors, hands, embodiments. β†’ the "foundation model for touch" thesis, Tactile VLA.

2.4 Humanoid whole-body loco-manipulation

  • ⚠️ Coordinated Humanoid Manipulation with Choice Policies (2512.25072, Berkeley β€” Qi, Wang, Lin, Yi, Ma, Sreenath, Malik; project) β€” Choice Policy = generate multiple candidate actions and learn to score them (fast inference + multimodal behavior), on a modular teleop + scalable learning stack; validated on dishwasher loading and whole-body whiteboard wiping. β†’ Humanoid VLA.
  • ⚠️ Weave: Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions (2609.16683) β€” contact-aware retargeting of human-object interactions β†’ jointly commands 29 body + 12 finger joints; 92.5% on trained interactions, 65.0% zero-shot on unseen sequences (9 objects). β†’ whole-body dexterity.
  • βœ… OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation (2606.26201) β€” long-horizon humanoid loco-manip needs robust meta-skills and seamless closed-loop chaining with recovery. A hierarchical framework centered on contact flow (CF) β€” key body trajectories + time-series binary contact signals. 98.7% Carry-Box, 76.5% Push-Stack-Boxes; +40.9% meta-skill and +66.5% skill-chaining over baselines. Its task suite is released CC-BY-4.0, and per the publisher it is trained on the commercial HiPHI-MOV corpus for which the public HiPHI dataset (below) is the free sample (vendor statement β€” modalitynet.com). β†’ contact-centric whole-body chaining; the peer-reviewed evidence for what the licensed corpus trains.
  • βœ… AdaPT: Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking (2608.20087, Noitom Robotics + collaborators; project) β€” learns pro tennis serve/rally styles from broadcast video (Nadal/Federer/Djokovic) and executes them on real humanoids via adaptive planning + tracking to bridge sim-to-real. Deployed on Unitree G1 and the full-size Dobot Atom (1.7 m), with in-the-wild serving without motion capture. β†’ dynamic, style-conditioned whole-body control.
  • βœ… HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction (2608.16222, Noitom Robotics; dataset) β€” a 617.5-hour optical-mocap dataset/benchmark for humanoid learning: 371.8 h whole-body movement + 245.7 h human-object interaction (object trajectories + meshes, 40 physical objects). Policies trained on it show matched-budget tracking that keeps improving with data scale and transfers to real humanoid hardware β€” the "one data foundation underneath" for the humanoid line. β†’ the humanoid data-scaling substrate.
  • βœ… Perceptive Behavior Foundation Model: Adapting Human Motion Priors to Robot-Centric Terrain (2606.08059, HKUST-GZ) β€” a human motion specifies intent but not the footholds/clearance/body-height/contact-timing the robot's terrain demands. Perceptive-BFM grounds human motion priors in onboard local perception to adapt them to terrain without changing the raw kinematic command interface. Deployed on a 29-DoF Unitree G1 over blocks, uneven ground, and stairs. β†’ perception-conditioned humanoid whole-body control.
  • βœ… MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction (2602.15733, X-Humanoid + HKUST-GZ + HKU) β€” learns coupled motion-terrain interaction directly from video: reconstructs human trajectories and terrain/object 3D geometry with SOTA 3D-vision models, extracts clean motion via a kinematic-consistency optimization, and retargets with a contact-invariant method. β†’ video-to-humanoid with explicit scene geometry.
  • βœ… CHIP: Adaptive Compliance for Humanoid Control through Hindsight Perturbation (2512.14689) β€” humanoids do agile locomotion but struggle at forceful manipulation (moving objects, wiping, pushing a cart). CHIP is a plug-and-play module giving controllable end-effector stiffness while preserving agile tracking of dynamic reference motions, via hindsight perturbation β€” no data augmentation, no extra reward tuning. β†’ the robot learns not just where to move but how compliant to be; Humanoid VLA.

2.5 RL / demonstration-efficient policy learning

  • βœ… FlashRFCL: Demonstration-Efficient Sparse-Reward RL via Reverse-Forward Curriculum Learning in High-Throughput GPU Simulation (project, UCSD β€” Choi, Stone Tao, Christensen) β€” combines scalable RL + high-throughput GPU sim + a demonstration auto-curriculum (a reverse curriculum then forward curriculum, leveraging even a single demonstration via per-demo state resets) to solve high-dimensional robotic tasks fast from sparse reward. Extends the ICLR-2024 RFCL recipe to high-throughput sim (ManiSkill-style). β†’ RL for VLA Β· data-efficiency.
  • βœ… SDPG: Efficient On-Policy Visual-RL via Stochastic Decoupled Policy Gradient (2609.20575, Yale (Apollo Lab) β€” co-first Haoxiang You & Yilang Liu, … Rakita, Ian Abraham; code) β€” end-to-end pixelsβ†’actions RL on a single RTX 4080, no teacher-student distillation, no GPU cluster. Two insights: (1) rewrite the likelihood-ratio policy gradient as regression onto locally-perturbed actions β€” mathematically identical, but it splits the update into a local-search step + a supervised policy-fit step that optimize independently; (2) not every env needs to render β€” mix ~64 rendered envs with many cheap physics-only ones to estimate the local reward landscape and cut gradient variance. Trains at ~distillation speed on visual MuJoCo, strongest on hard humanoid tasks; zero-shot sim-to-real to Unitree Go2 after ~2 h of sim training. Finding: end-to-end visual learning can beat distillation on harder tasks β€” pixels aren't something you must first distill away. β†’ RL for VLA, Independent Visual Representation.
  • βœ… UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms (2605.30313; code) β€” challenges the "physics must live on the GPU" default: decouples CPU-parallel simulation from GPU policy updates via a unified runtime for data movement/buffering/sync (MuJoCoUni + MotrixSim CPU physics backends; PPO/SAC/FlashSAC). A scalable RL infrastructure contribution. β†’ training-systems side of RL for VLA.
  • βœ… TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics (2602.19313) β€” instead of asking a VLM to output a reward score, read its internal token probabilities and turn them into dense zero-shot rewards β€” no reward-model training. On ManiRewardBench (130 tasks, 4 platforms) + Open-X, beats prior training-free VLM-reward methods on open models and rivals a trained reward model on progress estimation (tested on Qwen3-VL-8B, Molmo2, Gemini-2.5-Pro). β†’ training-free reward signals for RL for VLA.

2.6 Policy adaptation, deployment & sim-to-real

  • βœ… EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment (2606.12965) β€” training-free: keep policy learning in Cartesian space, then at inference lift diffusion sampling into the target robot's joint space via forward kinematics + Jacobian updates, adding whole-body collision-aware guidance after each denoising step (steer arm off collisions, preserve end-effector behavior). vs Cartesian-only: βˆ’46.1% collision, +28.5% success across 9 sim robots. β†’ Single-Checkpoint Multi-Robot, Cross-Embodiment.
  • βœ… FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space (2607.08877, Microsoft Research + ETH ZΓΌrich + UW) β€” latent-space DAgger for flow/diffusion policies: instead of fine-tuning the base model, a small steering network predicts the initial noise fed to the sampler; action inversion maps a human correction into the latent noise that would have generated it under the frozen policy. Adapts Ο€0.5 / GR00T N1.7 / Cosmos-Policy / diffusion policies within minutes, preserves prior capabilities, trains the adapter on a single 8 GB GPU. β†’ online correction without touching VLA/WAM weights; Real-Time Execution, RL for VLA.
  • βœ… TAM: Torque Adaptation Module for Robust Motion Transfer in Manipulation (2606.06218 β€” Son, Shkurti, Lee, Shah, Kim, Dieter Fox) β€” sits below the task policy and low-level controller and corrects torque commands from recent motion+torque history to match an "ideal robot," so the policy keeps its interface. Trained purely in sim, no real data; the same weights improve a real Franka Panda across RL box-pushing, IL cube-flipping, and MPC ball-on-plate β€” no task-policy retraining or per-task TAM fine-tuning. β†’ a sim-to-real dynamics-gap fix at the torque layer.
  • βœ… VLS: Steering Pretrained Robot Policies via Vision-Language Models (2602.03973) β€” training-free inference-time adaptation of a frozen diffusion/flow policy: a VLM synthesizes a trajectory-differentiable reward from the OOD observation+language, which steers the denoising toward test-time spatial/task requirements β€” policy weights untouched. +31% CALVIN, +13% LIBERO-PRO. β†’ VLM-guided action-steering at inference (cf. FlowDAgger above; RL for VLA).

2.7 Data generation, benchmarks & human-video pretraining

  • βœ… HuRo: Robotizing Human Videos for Scalable VLA Pretraining (2609.10706, RLWRLD + Yonsei; project) β€” converts heterogeneous egocentric human videos into robot-aligned observations + action trajectories (motion retargeting + visual robotization, inferring missing intermediate signals), yielding 630K robotized episodes / 142M frames from five sources. Scaling this in VLA pretraining lifts overall completion 51.5%β†’80.3% and OOD 34.9%β†’72.2%; ablations show visual robotization drives OOD robustness while retargeted action supervision adds gains beyond visual-only transfer. β†’ the Human Video β†’ Robot Transfer / egocentric-video pretraining scaling story.
  • ⚠️ MAGMA-GEN (author-announced, materials "coming soon" β€” no preprint located yet) β€” generating training data for embodied agents more efficiently than distillation or human annotation. β†’ synthetic data-generation for robot learning (verify against preprint when posted).
  • βœ… RoboReel β€” "Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation" (2609.08209; project) β€” a unified learning-from-observation benchmark: 10 tabletop tasks, 2,000 real human-demo videos, paired real+calibrated-sim environments, multi-view (incl. egocentric), and four test suites (visual distractors, domain shift, long-horizon). Evaluates 7+ LfO methods (incl. VLA variants) and compares language vs latent features vs human keypoints for humanβ†’robot transfer; finds long-horizon and low-tolerance tasks still hard. β†’ the evaluation counterpart to HuRo; Human Video β†’ Robot Transfer, VLA Evaluation.
  • βœ… SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale (2606.13497, KIT β€” Mattes, Li, Suliga, Roth, Vanjani, Reuss, Lioutikov) β€” auto-labels raw demos with structured spatial annotations (boxes, object trajectories, interaction-phase segments) and a reliability score per annotation, using the robot's own interactions as evidence of quality β€” so you can select high-quality training data without reviewing every example. Retains 3Γ— more usable data at high-precision operating points than detection-only baselines; improves downstream reasoning/policy learning. β†’ scalable annotation for embodied FMs.
  • βœ… WireCraft: A Simulation Benchmark for Industrial DLO Manipulation (2606.18097 β€” Zhu, ElMallah, Kim) β€” a deformable-linear-object (wire/cable) benchmark for industry: three task families (connector insertion, clip routing, channel seating), configurable assets/difficulty, two DLO physics models (articulated + deformable), and open trajectories from sim + real UR5. Benchmarks RL/IL/VLA under shared metrics (privileged RL >82%). β†’ deformable-object manipulation eval (cf. PhysCoRe Β§2.1).
  • βœ… LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition (2606.11628, LeCAR Lab) β€” two-stage: learn short-horizon intent ("what should happen next in the scene") from internet human video, then an embodiment-specific sensorimotor policy learned in massively-parallel sim converts intent β†’ actions ("learn what, let each robot figure out how"). Closed-loop 73% vs open-loop 28% on web-supervised tasks; one intent model drives two embodiments at comparable success. β†’ decoupled intent/execution for Human Video β†’ Robot Transfer Γ— dexterity.
  • βœ… World In Your Hands (WIYH): A Large-Scale, Open-Source Ecosystem for Human-Centric Manipulation in the Wild (2512.24310) β€” 1,000+ hours of in-the-wild human manipulation at mm-scale accuracy: the Oracle Suite wearable capture kit + auto-labeling, the WIYH Dataset (multimodal, hundreds of skills), and annotations/benchmarks perceptionβ†’action. Adding WIYH data lifts robot success 8%β†’60% in clutter; all data + hardware open. β†’ scaling beyond teleoperation, egocentric-video pretraining.

2.8 VLA safety, robustness & adversarial

  • βœ… Trajectory-Level Redirection Attacks on VLA (2606.12978, UIUC + Amazon β€” Puthumanaillam, Dongre, Thangeda, Nayyeri, Hakkani-TΓΌr, Ornik; more results) β€” a prompt-only threat model ("command-preserving trajectory redirection"): a single instruction edit that stays close to the benign command and never names the target outcome can still steer the whole action sequence to an attacker-chosen physical task. Found via an on-policy, rollout-based prompt search (model + environment fixed, only the instruction adapts). Demonstrated across several VLA families in LIBERO sim and on a real SO-100 arm β€” one edited prompt redirects the episode. β†’ the reliability flip-side of instruction grounding (cf. CounterAlign above, VLA Evaluation).
  • βœ… BarrierFormer: Transformer-Guided Predictive Barrier Enforcement for Safe Robot Control (2609.23896, ASU Safe Robotics Group β€” Chauhan et al.; project) β€” learns predictive safety during training so deployment is model-free and optimization-free: a transformer maps observation-action history straight to safe control, no online QP / no model knowledge. Combines transformer rollout prediction with control-barrier supervision; beats RL-, diffusion-, MPC-, and prior transformer-based safe controllers on safety rate + inference latency across 2D/3D linear & nonlinear systems, with zero-shot generalization. β†’ amortized control-barrier safety (a learned alternative to online CBF-QP).

2.9 Adjacent β€” driving, HRI & locomotion (out of this survey's manipulation scope)

  • βœ… SafeDriveVLA: Navigation-Conditioned World-Model Dreaming for Conflict-Aware End-to-End Autonomous Driving (Xie, Zhang, Wang, … Christensen, Chen) β€” a navigation-conditioned world model for conflict-aware end-to-end driving β€” a driving VLA+WM, listed for cross-domain reference.
  • βœ… BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving (Zhao, Wang, … Christensen, Ma, Zhou) β€” studies the open-loop vs closed-loop evaluation gap in end-to-end driving (an evaluation-methodology contribution). (Both via the Christensen-lab CoRL 2026 post.)
  • βœ… Observe Before You Alert: (Belief-State / Adaptive) Driver Alerting with Vision-Language Models (2609.08130) β€” VLAlert frames driver warning as a timing problem: a tri-action policy (SILENT / OBSERVE / ALERT) over a Qwen3-VL-4B safety-evidence generator, pooling hidden states from structured belief spans. On held-out ADAS takeover clips, R@5s 74.2%β†’88.7%, F1 0.585β†’0.686. A driver-monitoring VLM, listed for cross-domain reference.
  • βœ… Before Parc FermΓ©: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving (2605.31256 β€” Benfenati, Azimi, Risso, Carapellese, Jahier Pagliari, Burrello) β€” asks not how much but when to prune an embodied-LLM controller trained SFTβ†’RL. BPF folds pruning into RL so the model keeps adapting to closed-loop feedback after each pruning step until its "parc fermΓ©" (architecture frozen); better size-to-performance trade-offs than post-training pruning. A driving efficiency contribution, listed for cross-domain reference.
  • βœ… AI Coaching for Accelerating Human Skill Development with Reinforcement Learning (2606.25337, UPenn xLAB β€” Wang, Gu, Loquercio, Hu, Mangharam) β€” flips AI-assistant (which causes skill atrophy) into an AI coach that builds long-term human competence: Learning-to-Coach (L2C) infers the learner's latent skill in real time and scaffolds at the edge of capability (allowing "productive failures"), trained via PPO over a cognitive-science learner model. FPV drone-racing user study (33 people): βˆ’27.9% lap time after 40 min, beating AI-copilot baselines. A human-skill / HRI contribution, not manipulation.
  • βœ… Mind the Phase: Effective Rank and Representation Health in Legged Locomotion (Tommaselli et al.) β€” studies how locomotion policies respond to inputs across gait phases, using effective rank as a representation-health signal linked to smoother sim-to-real. A legged-locomotion analysis, out of manipulation scope.
  • βœ… Learning Contact Representation for Leg Odometry (2606.05501, Embry-Riddle β€” Girgin, Kilic) β€” self-supervised latent contact representation from joint kinematics only (denoising autoencoder + GMM over the latent space), improving leg odometry; finds contact-heuristic supervision can hurt and bigger context doesn't always help. A legged-odometry / state-estimation contribution, out of manipulation scope.

2.10 In-depth per-paper pages

Long-form analyses (Problem Β· Method Β· Results Β· Why-it-matters Β· Limitations, with arXiv links & figures) for the technically representative accepts:


3. Early insights (from the confirmed set)

  1. WAM converges on geometry-aware latent policy spaces, not pixels. SG-WAM predicts future dynamics in a policy-derived geometry-aware representation and wins OOD at 0.9B without embodied pretraining β€” reinforcing the wiki's latent/reactive WAM trend and the IROS-2026 "WAM-as-scaffold/representation" read (WAM vs VLA Robustness).
  2. Visual-tactile dexterity increasingly bootstraps from human video via sim. Dex-X's reconstruct-in-sim β†’ tactile supervision β†’ distill pipeline is the CoRL face of the data-pyramid L1β†’L5 bridge (video β†’ sim-tactile β†’ policy), and tactile-first IL (Touch2Trace) keeps maturing.
  3. VLA fine-tuning is being made robustness-preserving. FiberTune's "don't let action-loss collapse visual structure" is a training-time regularizer echoing the multi-task/instruction-collapse concerns in Multi-Task VLA.
  4. Whole-body humanoid manipulation is a first-class CoRL 2026 theme, now with a data substrate. Coordinated multi-limb (Choice Policies), human-interaction-retargeted whole-body dexterity (Weave), contact-flow skill chaining (OmniContact), and style-conditioned dynamic control (AdaPT tennis) all target body + hands together β€” and HiPHI's 617.5 h high-precision mocap is the shared "one data foundation underneath" that keeps improving with scale. The Humanoid VLA frontier is now data-bottleneck-first.
  5. Prediction is shifting from generating the action path to monitoring it β€” the sharpest new WAM-vs-VLA signal here. K-UBM rolls out a structured latent (Koopman) system as a pseudo planner and uses the predicted-vs-observed visual-flow discrepancy as a runtime monitor for event-triggered replanning; DAP couples future-observation prediction and action generation bidirectionally for mutual correction. Both reinforce the IROS-2026/ WAM vs VLA Robustness read that the world model is migrating out of the inference path into a supervisory/representational scaffold β€” here specifically as an execution monitor.
  6. VLAs are going reactive inside the chunk. Open-loop action-chunk execution misses moving targets; VLA-Feedback keeps the plan but corrects each action on the final denoising step against fresh vision ("plan ahead, react in the moment"), lifting dynamic-task success sharply — the Real-Time Execution frontier. Related pushes on the interface to vision: S2 (act from a learned visual-evidence budget) and SDPG (end-to-end pixels→actions RL that can beat state-distillation on hard tasks) both argue less-processed / less-distilled visual input can generalize better.
  7. Instruction grounding is being manufactured, not just collected. CounterAlign synthesizes counterfactual (instruction, action) mismatches from existing demos to learn an instruction-grounded offline-RL reward β€” a data-free complement to the Multi-Task VLA instruction-collapse fixes.

(Caveat: these insights rest on the confirmed-so-far set; they should be revisited against the full program.)


4. Caveats & how to extend

  • Acceptance verification: βœ… entries have an author/explicit CoRL-2026 statement; ⚠️ entries are reported/likely and should be re-checked against the official OpenReview list. ManiFlow (2509.01819) is CoRL 2025, not 2026 β€” deliberately excluded.
  • Coverage is partial by construction β€” 687 papers accepted; this lists a handful whose acceptance is publicly traceable now. Expect major gaps (grasping, sim-to-real, planning, multi-robot, mobile manipulation).
  • Next step: promote to a full session-taxonomy survey (like IROS 2026) once CoRL publishes the program; add per-paper in-depth pages for the flagship accepts.

5. Links

← Back to Reviews Β· Home