CoRL 2026 VLA Manipulation Survey - Heungwoo/research GitHub Wiki
CoRL 2026 β VLA & Manipulation Survey (preliminary, community-sourced)
Venue: 10th Conference on Robot Learning (CoRL) 2026 Β· Austin, TX (JW Marriott) Β· Nov 9β12, 2026 (workshops Nov 9, main Nov 10β12). Scale: 687 papers accepted Β· 32.8% acceptance Β· 2,094 active submissions (decisions announced; official per-paper program on OpenReview not yet public). β οΈ Sourcing caveat. There is no official accepted-papers list yet. This page is a preliminary, community-sourced selection of robot-manipulation papers whose CoRL 2026 acceptance is either explicitly stated by the authors (β ) or reported/announced but not yet independently verified (β οΈ). Every listed paper is a real arXiv preprint; the acceptance is what carries the caveat. To be replaced by a full survey once the official program publishes. Legend: β = author-confirmed CoRL 2026 Β· β οΈ = reported/likely, unverified.
1. At a glance
- 687 accepted / 2,094 submitted (32.8%) β CoRL stays selective and single-track. The manipulation/VLA/humanoid/WAM slice is large but not yet enumerable pre-program.
- ~47 author-confirmed in-scope papers (Β§2) across eight themes β World-Action / latent-dynamics (SG-WAM, Sensitivity-Shaping, K-UBM, DAP, PhysCoRe, WHIRL, StressDream), VLA (FiberTune, VLA-Feedback, S2, CounterAlign, Real-Time-AR, AFP, MessyMem, Efficient-VLA, MolmoAct2, MolmoBOT, SAE-VLA), dexterous & visuo-tactile (Dex-X, Touch2Trace, DexFLEX, HUGS, TeleDexter, MiTaS, FTP-1), whole-body humanoid (OmniContact, AdaPT, HiPHI, Perceptive-BFM, MeshMimic, CHIP), demonstration-efficient RL (FlashRFCL, SDPG, UniLab, TOPReward), policy adaptation / sim-to-real (EmbodiSteer, FlowDAgger, TAM, VLS), data-gen, benchmarks & human-video (HuRo, RoboReel, SPARC, WireCraft, LUCID, WIYH), and safety/adversarial (Trajectory-Redirection, BarrierFormer) β plus reported (β οΈ) StellaVLA/Choice-Policies/Weave/Look-Before-You-Move/Disentangling/MAGMA-GEN and adjacent driving/HRI/locomotion papers.
- What they signal (Β§3): all the 2026 threads the wiki tracks are present, and several sharpen: (a) prediction is being used to monitor execution, not just generate it (K-UBM's flow discrepancy β replan; WHIRL predicts human takeover β avoid; DAP's bidirectional obsβaction coupling); (b) VLAs are going reactive inside the action chunk (VLA-Feedback's final-step correction; Real-Time-AR); (c) the visual interface is being narrowed (S2 evidence budgets, AFP anti-shortcut foveation, SDPG end-to-end pixels); and (d) human video is being robotized at scale (HuRo 630K episodes) alongside a humanoid mocap substrate (HiPHI).
- Not in scope here: this is manipulation-centric; CoRL's pure locomotion/navigation/driving/theory tracks are excluded (SafeDriveVLA, BridgeSim, VLAlert, BPF, AI-Coaching, Mind-the-Phase are listed only as adjacent reference in Β§2.9).
2. Papers by theme (β confirmed Β· β οΈ reported)
2.1 World-Action Models / world-model & latent-dynamics policies
- β SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space (2608.01397) β aligns world modeling with action generation by predicting future dynamics in a geometry-aware, policy-derived representation space (not pixels). A 0.9B model with no large-scale embodied pretraining hits LIBERO 98.5% / LIBERO-Plus 73% and beats baselines in ID and OOD real-world eval. β the latent/representation-space WAM corner (World Models, VLA Hybrid Architectures).
- β Sensitivity Shaping for Latent Modeling (2606.14585, UCSD β Yu, Li, Zhang, Christensen, Gao) β generative dynamics models enable planning, but safe deployment needs to detect policy-induced OOD transitions; existing surrogates fail when dynamics are locally insensitive to critical action choices. Introduces support-conditioned control-sensitivity regularization so the learned dynamics respond sensitively to control changes in high-support regions β a world-model reliability contribution (planning-with-learned-dynamics safety), cf. World Models Β§6 hallucination/OOD.
- β K-UBM: Going with the Flow β Koopman Behavioral Models as Pseudo Planners for Visuo-Motor Dexterity (2602.07413, Georgia Tech β Yunhai Han, β¦ Harish Ravichandar) β treats manipulation as the coupled evolution of robot action + environmental visual flow in a shared latent space governed by a structured (Koopman) linear system. Instead of repeatedly predicting short action chunks, it rolls out the latent dynamical system as a "pseudo planner" for coherent long-horizon behavior + very fast inference. Crucially it predicts future visual flow and uses the predicted-vs-observed flow discrepancy as a runtime monitor β event-triggered replanning. 7 sim + 4 real dexterous tasks: competitive success with smoother execution, occlusion robustness, efficient inference. β prediction-as-execution-monitor (WAM used to supervise rather than generate the action path), cf. World Models, Dexterous Manipulation.
- β DAP: Dynamics-Aligned Flow Matching Policy (2510.27114 β Cho, Lee, β¦ Li Zhao) β jointly models future observations and actions, but unlike a plain World Action Model it makes the coupling bidirectional: an architecture for action-conditioned future-observation prediction and future-conditioned action generation, with the policy and dynamics models giving each other mutual corrective feedback during generation. Learning both directions together improves the sample-efficiency and perturbation-robustness of diffusion/flow-matching imitation policies. β the bidirectional-coupling variant of the WAM+policy design (cf. its NeurIPS-2026 companion Why Latent Actions Fail on future-leakage in latent-action models).
- β PhysCoRe: Physics-Corrected Residual World Models for Material-Aware Deformable Dynamics (2607.20653 β Tao, β¦ Lu Gan, Chen) β keeps a physics model at the core (differentiable Material Point Method) while gaining learned generalization, instead of per-object material fitting (slow) or end-to-end dynamics (drifts unphysically). Material-from-Motion infers per-particle material + an unsupervised confidence (concentrates on what actually deforms β where to probe next) in one forward pass; Residual-from-Dynamics injects a velocity residual inside the MPM cycle to close sim-to-real while preserving physical structure. Beats SOTA physics/learned baselines on real elastic & elastoplastic deformation β a deformable-object dynamics backend for planning. β the physics-structured corner of World Models.
- β WHIRL: Intervention-Aware World Models with Real-World RL for Dexterous Manipulation (2609.06009 β Xu, Zhang, Feng, β¦ Ajoudani, Renjing Xu) β turns one-off human takeovers into predictive risk signals: a latent world model with the usual dynamics/reward/termination heads plus a per-state intervention-probability head that predicts where a human would take over, then steers the policy away from intervention-prone states. On five real tasks with a 16-DoF LEAP Hand, +15β30 pts autonomous success and up to β84% operator-controlled training steps. β another world-model-as-monitor instance (predict-to-avoid), cf. K-UBM above and World Models; safety-aware real-world RL.
- β StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement (2606.00267, CMU IntentLab) β don't ask a video WM only for the most-likely future; optimize the diffusion WM's initial noise to steer imagination toward high-impact yet plausible outcomes specified by text at inference (e.g. task failures), exposing where the policy breaks β then train on those failures. Works with SOTA driving + manipulation video WMs. β world-model-as-stress-tester for evaluation/robustness; World Models, VLA Evaluation.
2.2 Vision-Language-Action models
- β FiberTune: Preserving Action-Fiber Visual Residuals in VLA Fine-Tuning (2606.08653) β action-supervised fine-tuning constrains only action-changing directions and lets action-equivalent visual structure collapse; FiberTune adds a training-time objective (online action probe β filter action-predictive directions β align residuals to a frozen visual teacher + rank regularization), no inference overhead. Gains on Ο0.5 and OpenVLA-OFT (e.g. +10.7 pp SR(5) on CALVIN ABCβD; SO-101 72.7%β78.1%). β the VLA-robustness/fine-tuning thread.
- β οΈ StellaVLA: In-Context Structured Demonstration for Generalizable VLA (2608.11671) β conditions on one retrieved demo auto-converted offline into a structured demo (task plan + sub-goals + verbalized 3D motion), so the policy reasons rather than mimics pixels; transferable across real-robot / human-hand / XR demos. Reported #1 on VLA-Arena (0.63 vs 0.44/0.22). β In-Context Imitation.
- β VLA-Feedback β "Catch Me If You Can: Real-Time Feedback Denoising for Responsive VLAs" (2609.21022) β VLAs execute a pre-generated action chunk open-loop, so a target that moves mid-chunk is missed. A two-timescale design pairs low-frequency diffusion planning with high-frequency visual feedback: rather than fully denoising the chunk before execution, it keeps the final denoising step as a lightweight feedback interface, correcting each action against the latest observation before it fires ("plan ahead, react in the moment"). Matches GR00T on static LIBERO; lifts dynamic sim 27.5%β85.0% and real-robot 51%β73%. β Real-Time Execution (reactive-during-chunk).
- β S2: See Less, Specify More β Visual Evidence Budgets for Generalizable VLAs (2606.02735, AIRoA β Wu, Matsushima, Ota) β a VLA must infer what to do and where to look from a coarse goal + full frame, on a fraction of a VLM's data. See Less imposes an explicit visual-evidence budget (act from task-sufficient evidence, learned straight from the control objective β no masks/region annotations); Specify More keeps the original instruction as a stable goal while relabeling trajectories into refined sub-task language. Showing the policy less of the image improves generalization. β Independent Visual Representation / VLA generalization.
- β CounterAlign: Counterfactual Supervision for VLA (2608.21740, AIRoA / Tokyo Tech β Kondoh, Ota, Kanezaki, Wu) β relabels existing demos into counterfactual (instruction, action) mismatch pairs; a discriminator learns to catch the mismatch β an instruction-grounded reward for offline RL (IQL + advantage-weighted flow matching). No new data, no hand labels; improves robustness to object-position/task perturbation on LIBERO-PRO and on the real TX-G2. β Multi-Task VLA (instruction grounding), RL for VLA.
- β Real-Time Execution with Autoregressive Policies (2606.13355, KIST + SNU + Google Research β Lee, Park, You, Caciularu, Szpektor, Lim, Youngjae Yu) β real-time robotics has favored diffusion/flow policies because autoregressive ones have rollout-speed bottlenecks in synchronous inference. Shows AR policies can run real-time by adjusting the tokenization horizon + constrained decoding to guarantee strict latency bounds, which unlocks multi-trajectory decoding. Across sim + real, the AR policy beats equivalent flow-matching counterparts with faster completion β while keeping AR's native edge in convergence and instruction-following. β Real-Time Execution (async, anytime decoding).
- β Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models (2607.10655) β standard fine-tuning supervises actions only, so a policy latches onto background/lighting/co-occurring objects β shortcut learning β and collapses under distractors (Ο0.5 real-robot 60%β27% once distractors appear). AFP is a lightweight, policy-agnostic module (same RGB + instruction as the policy) predicting a continuous fovea-like mask over task-relevant objects + end-effector, used as an auxiliary grounding loss on image-token attention during fine-tuning (PCGrad projection removes components conflicting with the action gradient). Architecture unchanged; not in the inference loop. MimicGen (4 models Γ 8 tasks): under-distractor success 39%β66%, never hurts clean; real Ο0.5 27%β53%. Ships a 786-episode mask dataset. β the "budget/where-to-look" thread with S2 above; Multi-Task VLA robustness.
- β οΈ Look Before You Move: Agent-Interface-Aligned Supervision for Open-World Mobile Manipulation (HKUST-GZ; author-announced, code/model "coming soon" β no preprint located yet) β agent-interface-aligned supervision for open-world mobile manipulation. β mobile-manipulation VLA (listed on author announcement; verify against preprint when posted).
- β MessyMem: Learning-from-Doing Memory for Mobile Manipulation (2609.15976, Stanford IPRL β Muckelroy, S., Zhao, Bohg, Ho; project) β persistent memory for mobile manipulators that combines a spatially-grounded 3D scene graph + interaction-revealed properties (VLM records what actions expose, e.g. "drawer locked") + linked keyframes for fine-grained appearance recall; a new task retrieves the most relevant visual memories alongside the scene graph for a VLM planner. Over 25 consecutive tasks / 3+ hours, retrieves evidence from thousands of keyframes (>1 h old) β 80.0% progress, +28.9 pts over the best external baseline; real TidyBot++ deployment with adaptation to environment change. β the mobile-manipulation memory frontier, VLA Memory.
- β What Makes an Efficient VLA? Navigating Action-Head Design, Scaling, and Latency (2609.13984, Li Auto + collaborators) β a T5-style empirical sweep rebuilding VLAs 60+ times under latency-paired conditions (fixed SigLIP2 + Qwen2.5 backbones, swept action-head design & module scale, each paired with measured on-device latency, real-robot validated). Headline: action-head performance is governed mainly by initialization, not decoder architecture / loss / inference budget β copying the last transformer layers of the language backbone into the head is the single largest lever, at no latency cost, and the only axis that helps at every scale. β a design/efficiency reference for VLA Architectures, Real-Time Execution.
- β οΈ Disentangling Spurious Correlations in VLA Models by Predicting Domain-Invariant Latent Lookahead (SNU AAIG β J. Kim, E. Kim, Gi-Cheon Kang, Byoung-Tak Zhang; author-announced, preprint not yet located) β VLAs pick up undesired dependencies between task-irrelevant visual cues (lighting, camera perspective) and actions; this policy predicts a domain-invariant future latent that separates task-relevant signal from visual variation, disentangling spurious correlations to improve robustness under visual distribution shift. β the anti-shortcut robustness thread with AFP/CounterAlign above (verify against preprint when posted).
- β MolmoAct2: Action Reasoning Models for Real-World Deployment (2605.02881, Allen AI) β a fully open action-reasoning VLA stack: MolmoER spatial/embodied-reasoning VLM backbone (3.3M-sample corpus, specialize-then-rehearse), OpenFAST open action tokenizer (millions of trajectories, 5 embodiments), and MolmoThink adaptive-depth reasoning (re-predicts depth tokens only for changed regions β low latency). Across 7 sim+real benchmarks beats Ο0.5; MolmoER surpasses GPT-5 / Gemini-Robotics-ER-1.5 on 13 embodied-reasoning benchmarks. β open action-reasoning VLA, VLA Architectures.
- β MolmoBOT: Large-Scale Simulation Enables Zero-Shot Manipulation (2603.16861, Allen AI) β tests whether enough simulation diversity lets manipulation transfer to the real world with no real-world training data. Fully open procedural pipeline (MolmoBot-Engine / MolmoSpaces) releasing 1.8M expert trajectories; on tabletop pick-and-place 79.2% real-world vs Ο0.5's 39.2%. β the sim-only-transfer thesis, cf. World Models and sim-to-real (Β§2.6).
- β Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models (2603.19183, Stanford β Swann, McGranahan, Buurmeijer, Kennedy, Schwager) β trains SAEs on a VLA's hidden activations to find features corresponding to motion primitives + semantic concepts, some general across episodes and causally steerable; validated by steering on LIBERO and real DROID hardware. β mechanistic-interpretability for VLAs (rare inside-the-model study).
2.3 Dexterous & visuo-tactile manipulation
- β Dex-X: Learning Visual-Tactile Dexterous Manipulation from Human Videos with Simulated Interaction (2609.07747) β reconstructs hand-object interaction in simulation where physically-grounded contact gives tactile supervision; teacher 65.9% (6 tasks, sim) distilled to a visual-tactile policy: 93% real cube-picking, 53% table-cleaning. β egocentric-video pretraining Γ tactile Γ data pyramid.
- β Touch2Trace: Tactile-Driven Imitation Learning for Dexterous Cable Tracing (2609.15921) β tactile-driven IL for dexterous cable tracing on a real hand: 93% success, zero-shot to unseen cables/routing. β tactile-first manipulation.
- β DexFLEX: Contact-Aware Foundation Controller for Command-Guided Dexterity (project, UCSD β Wan, Lai, Hou, Zhou, Fu, Christensen, Su) β turns upstream fingertip-motion drafts into contact-consistent joint commands: treats a draft as evidence about intent (not a trajectory to copy), proposes short-horizon motion chunks from the tactile-proprioceptive state, predicts their contact consequences, and selects the candidate that follows the command while preserving future contact stability. A DraftβDreamβSelect inference loop mixes pure-prior proposals (recovery) with draft-seeded ones (responsiveness). β the "servo on contact" dexterity line (cf. DexterityGen, T-Rex).
- β HUGS: Guiding Unified Dexterous Grasp Synthesis Across Modes and Scales via Learned Human Priors (2607.04554; project) β instead of retargeting human demos, learns an object-conditioned human prior that guides force-closure-aware optimization, giving one synthesizer across contact modes (two-finger β bimanual) and object scales (2β30 cm). Synthesizes 3.2M grasps over 157K scenes; policies trained on it auto-select the contact mode in the real world (screws β large boxes). β dexterous grasp data generation, Dexterous-Hand Data Pyramid.
- β TeleDexter: Towards Human-Level Dexterous Teleoperation (2607.11481) β a hand-object co-tracking controller mapping operator intent into learned low-level contact execution, trained on co-tracking subgoals from human reference motion with a hybrid sparse-subgoal + dense-tracking reward (single-stage RL, zero-shot to real via action masking + domain randomization). 7 hard tasks (reorientation, long-horizon tool use) on two hands: 75% avg where baselines fail. β high-quality dexterous data at the teleop layer of the data pyramid.
- β MiTaS: Multi-Resolution Tactile Imitation Learning for Contact-Rich Manipulation (2606.06281; project) β instead of choosing one tactile sensor, fuses heterogeneous sensors at different temporal resolutions: RGB + vision-based GelSight Mini + high-frequency event-based Evetac, via modality-specific conv stems + a transformer fusion module, conditioning a flow-matching policy. 5 real contact-rich tasks: 80% vs visual-tactile 54% / vision-only 31%; multi-tactile co-training helps even when Evetac is absent at eval. β the sensor-fusion frontier of Tactile VLA.
- β FTP-1: A Generalist Foundation Tactile Policy Across Tactile Sensors (2606.13102) β the first generalist foundation tactile policy: heterogeneous encoders project image-/array-/state-based tactile signals into unified morphology-aware latent tokens modeled by a shared tactile transformer expert. Pretrained on ~3,000 h across 26 sources / 21 sensors (human + robot), and transfers touch beyond sensors seen in pretraining β across sensors, hands, embodiments. β the "foundation model for touch" thesis, Tactile VLA.
2.4 Humanoid whole-body loco-manipulation
- β οΈ Coordinated Humanoid Manipulation with Choice Policies (2512.25072, Berkeley β Qi, Wang, Lin, Yi, Ma, Sreenath, Malik; project) β Choice Policy = generate multiple candidate actions and learn to score them (fast inference + multimodal behavior), on a modular teleop + scalable learning stack; validated on dishwasher loading and whole-body whiteboard wiping. β Humanoid VLA.
- β οΈ Weave: Whole-Body Dexterous Loco-Manipulation from Human-Object Interactions (2609.16683) β contact-aware retargeting of human-object interactions β jointly commands 29 body + 12 finger joints; 92.5% on trained interactions, 65.0% zero-shot on unseen sequences (9 objects). β whole-body dexterity.
- β OmniContact: Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation (2606.26201) β long-horizon humanoid loco-manip needs robust meta-skills and seamless closed-loop chaining with recovery. A hierarchical framework centered on contact flow (CF) β key body trajectories + time-series binary contact signals. 98.7% Carry-Box, 76.5% Push-Stack-Boxes; +40.9% meta-skill and +66.5% skill-chaining over baselines. Its task suite is released CC-BY-4.0, and per the publisher it is trained on the commercial HiPHI-MOV corpus for which the public HiPHI dataset (below) is the free sample (vendor statement β modalitynet.com). β contact-centric whole-body chaining; the peer-reviewed evidence for what the licensed corpus trains.
- β AdaPT: Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking (2608.20087, Noitom Robotics + collaborators; project) β learns pro tennis serve/rally styles from broadcast video (Nadal/Federer/Djokovic) and executes them on real humanoids via adaptive planning + tracking to bridge sim-to-real. Deployed on Unitree G1 and the full-size Dobot Atom (1.7 m), with in-the-wild serving without motion capture. β dynamic, style-conditioned whole-body control.
- β HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction (2608.16222, Noitom Robotics; dataset) β a 617.5-hour optical-mocap dataset/benchmark for humanoid learning: 371.8 h whole-body movement + 245.7 h human-object interaction (object trajectories + meshes, 40 physical objects). Policies trained on it show matched-budget tracking that keeps improving with data scale and transfers to real humanoid hardware β the "one data foundation underneath" for the humanoid line. β the humanoid data-scaling substrate.
- β Perceptive Behavior Foundation Model: Adapting Human Motion Priors to Robot-Centric Terrain (2606.08059, HKUST-GZ) β a human motion specifies intent but not the footholds/clearance/body-height/contact-timing the robot's terrain demands. Perceptive-BFM grounds human motion priors in onboard local perception to adapt them to terrain without changing the raw kinematic command interface. Deployed on a 29-DoF Unitree G1 over blocks, uneven ground, and stairs. β perception-conditioned humanoid whole-body control.
- β MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction (2602.15733, X-Humanoid + HKUST-GZ + HKU) β learns coupled motion-terrain interaction directly from video: reconstructs human trajectories and terrain/object 3D geometry with SOTA 3D-vision models, extracts clean motion via a kinematic-consistency optimization, and retargets with a contact-invariant method. β video-to-humanoid with explicit scene geometry.
- β CHIP: Adaptive Compliance for Humanoid Control through Hindsight Perturbation (2512.14689) β humanoids do agile locomotion but struggle at forceful manipulation (moving objects, wiping, pushing a cart). CHIP is a plug-and-play module giving controllable end-effector stiffness while preserving agile tracking of dynamic reference motions, via hindsight perturbation β no data augmentation, no extra reward tuning. β the robot learns not just where to move but how compliant to be; Humanoid VLA.
2.5 RL / demonstration-efficient policy learning
- β FlashRFCL: Demonstration-Efficient Sparse-Reward RL via Reverse-Forward Curriculum Learning in High-Throughput GPU Simulation (project, UCSD β Choi, Stone Tao, Christensen) β combines scalable RL + high-throughput GPU sim + a demonstration auto-curriculum (a reverse curriculum then forward curriculum, leveraging even a single demonstration via per-demo state resets) to solve high-dimensional robotic tasks fast from sparse reward. Extends the ICLR-2024 RFCL recipe to high-throughput sim (ManiSkill-style). β RL for VLA Β· data-efficiency.
- β SDPG: Efficient On-Policy Visual-RL via Stochastic Decoupled Policy Gradient (2609.20575, Yale (Apollo Lab) β co-first Haoxiang You & Yilang Liu, β¦ Rakita, Ian Abraham; code) β end-to-end pixelsβactions RL on a single RTX 4080, no teacher-student distillation, no GPU cluster. Two insights: (1) rewrite the likelihood-ratio policy gradient as regression onto locally-perturbed actions β mathematically identical, but it splits the update into a local-search step + a supervised policy-fit step that optimize independently; (2) not every env needs to render β mix ~64 rendered envs with many cheap physics-only ones to estimate the local reward landscape and cut gradient variance. Trains at ~distillation speed on visual MuJoCo, strongest on hard humanoid tasks; zero-shot sim-to-real to Unitree Go2 after ~2 h of sim training. Finding: end-to-end visual learning can beat distillation on harder tasks β pixels aren't something you must first distill away. β RL for VLA, Independent Visual Representation.
- β UniLab: A Heterogeneous Architecture for Robot RL Beyond GPU-Dominant Paradigms (2605.30313; code) β challenges the "physics must live on the GPU" default: decouples CPU-parallel simulation from GPU policy updates via a unified runtime for data movement/buffering/sync (MuJoCoUni + MotrixSim CPU physics backends; PPO/SAC/FlashSAC). A scalable RL infrastructure contribution. β training-systems side of RL for VLA.
- β TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics (2602.19313) β instead of asking a VLM to output a reward score, read its internal token probabilities and turn them into dense zero-shot rewards β no reward-model training. On ManiRewardBench (130 tasks, 4 platforms) + Open-X, beats prior training-free VLM-reward methods on open models and rivals a trained reward model on progress estimation (tested on Qwen3-VL-8B, Molmo2, Gemini-2.5-Pro). β training-free reward signals for RL for VLA.
2.6 Policy adaptation, deployment & sim-to-real
- β EmbodiSteer: Steering Embodiment-Agnostic Visuomotor Policies with Joint-Space Guidance for Zero-Shot Cross-Embodiment Deployment (2606.12965) β training-free: keep policy learning in Cartesian space, then at inference lift diffusion sampling into the target robot's joint space via forward kinematics + Jacobian updates, adding whole-body collision-aware guidance after each denoising step (steer arm off collisions, preserve end-effector behavior). vs Cartesian-only: β46.1% collision, +28.5% success across 9 sim robots. β Single-Checkpoint Multi-Robot, Cross-Embodiment.
- β FlowDAgger: Human-in-the-Loop Adaptation of Generative Robot Policies in Latent Space (2607.08877, Microsoft Research + ETH ZΓΌrich + UW) β latent-space DAgger for flow/diffusion policies: instead of fine-tuning the base model, a small steering network predicts the initial noise fed to the sampler; action inversion maps a human correction into the latent noise that would have generated it under the frozen policy. Adapts Ο0.5 / GR00T N1.7 / Cosmos-Policy / diffusion policies within minutes, preserves prior capabilities, trains the adapter on a single 8 GB GPU. β online correction without touching VLA/WAM weights; Real-Time Execution, RL for VLA.
- β TAM: Torque Adaptation Module for Robust Motion Transfer in Manipulation (2606.06218 β Son, Shkurti, Lee, Shah, Kim, Dieter Fox) β sits below the task policy and low-level controller and corrects torque commands from recent motion+torque history to match an "ideal robot," so the policy keeps its interface. Trained purely in sim, no real data; the same weights improve a real Franka Panda across RL box-pushing, IL cube-flipping, and MPC ball-on-plate β no task-policy retraining or per-task TAM fine-tuning. β a sim-to-real dynamics-gap fix at the torque layer.
- β VLS: Steering Pretrained Robot Policies via Vision-Language Models (2602.03973) β training-free inference-time adaptation of a frozen diffusion/flow policy: a VLM synthesizes a trajectory-differentiable reward from the OOD observation+language, which steers the denoising toward test-time spatial/task requirements β policy weights untouched. +31% CALVIN, +13% LIBERO-PRO. β VLM-guided action-steering at inference (cf. FlowDAgger above; RL for VLA).
2.7 Data generation, benchmarks & human-video pretraining
- β HuRo: Robotizing Human Videos for Scalable VLA Pretraining (2609.10706, RLWRLD + Yonsei; project) β converts heterogeneous egocentric human videos into robot-aligned observations + action trajectories (motion retargeting + visual robotization, inferring missing intermediate signals), yielding 630K robotized episodes / 142M frames from five sources. Scaling this in VLA pretraining lifts overall completion 51.5%β80.3% and OOD 34.9%β72.2%; ablations show visual robotization drives OOD robustness while retargeted action supervision adds gains beyond visual-only transfer. β the Human Video β Robot Transfer / egocentric-video pretraining scaling story.
- β οΈ MAGMA-GEN (author-announced, materials "coming soon" β no preprint located yet) β generating training data for embodied agents more efficiently than distillation or human annotation. β synthetic data-generation for robot learning (verify against preprint when posted).
- β RoboReel β "Monkey See, Can Monkey Do? A Benchmark for Evaluating Robot Skill Learning by Observation" (2609.08209; project) β a unified learning-from-observation benchmark: 10 tabletop tasks, 2,000 real human-demo videos, paired real+calibrated-sim environments, multi-view (incl. egocentric), and four test suites (visual distractors, domain shift, long-horizon). Evaluates 7+ LfO methods (incl. VLA variants) and compares language vs latent features vs human keypoints for humanβrobot transfer; finds long-horizon and low-tolerance tasks still hard. β the evaluation counterpart to HuRo; Human Video β Robot Transfer, VLA Evaluation.
- β SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale (2606.13497, KIT β Mattes, Li, Suliga, Roth, Vanjani, Reuss, Lioutikov) β auto-labels raw demos with structured spatial annotations (boxes, object trajectories, interaction-phase segments) and a reliability score per annotation, using the robot's own interactions as evidence of quality β so you can select high-quality training data without reviewing every example. Retains 3Γ more usable data at high-precision operating points than detection-only baselines; improves downstream reasoning/policy learning. β scalable annotation for embodied FMs.
- β WireCraft: A Simulation Benchmark for Industrial DLO Manipulation (2606.18097 β Zhu, ElMallah, Kim) β a deformable-linear-object (wire/cable) benchmark for industry: three task families (connector insertion, clip routing, channel seating), configurable assets/difficulty, two DLO physics models (articulated + deformable), and open trajectories from sim + real UR5. Benchmarks RL/IL/VLA under shared metrics (privileged RL >82%). β deformable-object manipulation eval (cf. PhysCoRe Β§2.1).
- β LUCID: Learning Embodiment-Agnostic Intent Models from Unstructured Human Videos for Scalable Dexterous Robot Skill Acquisition (2606.11628, LeCAR Lab) β two-stage: learn short-horizon intent ("what should happen next in the scene") from internet human video, then an embodiment-specific sensorimotor policy learned in massively-parallel sim converts intent β actions ("learn what, let each robot figure out how"). Closed-loop 73% vs open-loop 28% on web-supervised tasks; one intent model drives two embodiments at comparable success. β decoupled intent/execution for Human Video β Robot Transfer Γ dexterity.
- β World In Your Hands (WIYH): A Large-Scale, Open-Source Ecosystem for Human-Centric Manipulation in the Wild (2512.24310) β 1,000+ hours of in-the-wild human manipulation at mm-scale accuracy: the Oracle Suite wearable capture kit + auto-labeling, the WIYH Dataset (multimodal, hundreds of skills), and annotations/benchmarks perceptionβaction. Adding WIYH data lifts robot success 8%β60% in clutter; all data + hardware open. β scaling beyond teleoperation, egocentric-video pretraining.
2.8 VLA safety, robustness & adversarial
- β Trajectory-Level Redirection Attacks on VLA (2606.12978, UIUC + Amazon β Puthumanaillam, Dongre, Thangeda, Nayyeri, Hakkani-TΓΌr, Ornik; more results) β a prompt-only threat model ("command-preserving trajectory redirection"): a single instruction edit that stays close to the benign command and never names the target outcome can still steer the whole action sequence to an attacker-chosen physical task. Found via an on-policy, rollout-based prompt search (model + environment fixed, only the instruction adapts). Demonstrated across several VLA families in LIBERO sim and on a real SO-100 arm β one edited prompt redirects the episode. β the reliability flip-side of instruction grounding (cf. CounterAlign above, VLA Evaluation).
- β BarrierFormer: Transformer-Guided Predictive Barrier Enforcement for Safe Robot Control (2609.23896, ASU Safe Robotics Group β Chauhan et al.; project) β learns predictive safety during training so deployment is model-free and optimization-free: a transformer maps observation-action history straight to safe control, no online QP / no model knowledge. Combines transformer rollout prediction with control-barrier supervision; beats RL-, diffusion-, MPC-, and prior transformer-based safe controllers on safety rate + inference latency across 2D/3D linear & nonlinear systems, with zero-shot generalization. β amortized control-barrier safety (a learned alternative to online CBF-QP).
2.9 Adjacent β driving, HRI & locomotion (out of this survey's manipulation scope)
- β SafeDriveVLA: Navigation-Conditioned World-Model Dreaming for Conflict-Aware End-to-End Autonomous Driving (Xie, Zhang, Wang, β¦ Christensen, Chen) β a navigation-conditioned world model for conflict-aware end-to-end driving β a driving VLA+WM, listed for cross-domain reference.
- β BridgeSim: Unveiling the OL-CL Gap in End-to-End Autonomous Driving (Zhao, Wang, β¦ Christensen, Ma, Zhou) β studies the open-loop vs closed-loop evaluation gap in end-to-end driving (an evaluation-methodology contribution). (Both via the Christensen-lab CoRL 2026 post.)
- β Observe Before You Alert: (Belief-State / Adaptive) Driver Alerting with Vision-Language Models (2609.08130) β VLAlert frames driver warning as a timing problem: a tri-action policy (SILENT / OBSERVE / ALERT) over a Qwen3-VL-4B safety-evidence generator, pooling hidden states from structured belief spans. On held-out ADAS takeover clips, R@5s 74.2%β88.7%, F1 0.585β0.686. A driver-monitoring VLM, listed for cross-domain reference.
- β Before Parc FermΓ©: RL-Time Pruning for Efficient Embodied LLMs in Autonomous Driving (2605.31256 β Benfenati, Azimi, Risso, Carapellese, Jahier Pagliari, Burrello) β asks not how much but when to prune an embodied-LLM controller trained SFTβRL. BPF folds pruning into RL so the model keeps adapting to closed-loop feedback after each pruning step until its "parc fermΓ©" (architecture frozen); better size-to-performance trade-offs than post-training pruning. A driving efficiency contribution, listed for cross-domain reference.
- β AI Coaching for Accelerating Human Skill Development with Reinforcement Learning (2606.25337, UPenn xLAB β Wang, Gu, Loquercio, Hu, Mangharam) β flips AI-assistant (which causes skill atrophy) into an AI coach that builds long-term human competence: Learning-to-Coach (L2C) infers the learner's latent skill in real time and scaffolds at the edge of capability (allowing "productive failures"), trained via PPO over a cognitive-science learner model. FPV drone-racing user study (33 people): β27.9% lap time after 40 min, beating AI-copilot baselines. A human-skill / HRI contribution, not manipulation.
- β Mind the Phase: Effective Rank and Representation Health in Legged Locomotion (Tommaselli et al.) β studies how locomotion policies respond to inputs across gait phases, using effective rank as a representation-health signal linked to smoother sim-to-real. A legged-locomotion analysis, out of manipulation scope.
- β Learning Contact Representation for Leg Odometry (2606.05501, Embry-Riddle β Girgin, Kilic) β self-supervised latent contact representation from joint kinematics only (denoising autoencoder + GMM over the latent space), improving leg odometry; finds contact-heuristic supervision can hurt and bigger context doesn't always help. A legged-odometry / state-estimation contribution, out of manipulation scope.
2.10 In-depth per-paper pages
Long-form analyses (Problem Β· Method Β· Results Β· Why-it-matters Β· Limitations, with arXiv links & figures) for the technically representative accepts:
- WAM / latent-dynamics: SG-WAM Β· K-UBM Β· StressDream
- VLA: MolmoAct2 Β· MolmoBOT Β· VLA-Feedback Β· SAE-VLA
- Dexterous & tactile: FTP-1
- Humanoid: HiPHI Β· CHIP
- RL / steering: TOPReward Β· VLS
- Data & human-video: HuRo Β· LUCID Β· World In Your Hands
3. Early insights (from the confirmed set)
- WAM converges on geometry-aware latent policy spaces, not pixels. SG-WAM predicts future dynamics in a policy-derived geometry-aware representation and wins OOD at 0.9B without embodied pretraining β reinforcing the wiki's latent/reactive WAM trend and the IROS-2026 "WAM-as-scaffold/representation" read (WAM vs VLA Robustness).
- Visual-tactile dexterity increasingly bootstraps from human video via sim. Dex-X's reconstruct-in-sim β tactile supervision β distill pipeline is the CoRL face of the data-pyramid L1βL5 bridge (video β sim-tactile β policy), and tactile-first IL (Touch2Trace) keeps maturing.
- VLA fine-tuning is being made robustness-preserving. FiberTune's "don't let action-loss collapse visual structure" is a training-time regularizer echoing the multi-task/instruction-collapse concerns in Multi-Task VLA.
- Whole-body humanoid manipulation is a first-class CoRL 2026 theme, now with a data substrate. Coordinated multi-limb (Choice Policies), human-interaction-retargeted whole-body dexterity (Weave), contact-flow skill chaining (OmniContact), and style-conditioned dynamic control (AdaPT tennis) all target body + hands together β and HiPHI's 617.5 h high-precision mocap is the shared "one data foundation underneath" that keeps improving with scale. The Humanoid VLA frontier is now data-bottleneck-first.
- Prediction is shifting from generating the action path to monitoring it β the sharpest new WAM-vs-VLA signal here. K-UBM rolls out a structured latent (Koopman) system as a pseudo planner and uses the predicted-vs-observed visual-flow discrepancy as a runtime monitor for event-triggered replanning; DAP couples future-observation prediction and action generation bidirectionally for mutual correction. Both reinforce the IROS-2026/ WAM vs VLA Robustness read that the world model is migrating out of the inference path into a supervisory/representational scaffold β here specifically as an execution monitor.
- VLAs are going reactive inside the chunk. Open-loop action-chunk execution misses moving targets; VLA-Feedback keeps the plan but corrects each action on the final denoising step against fresh vision ("plan ahead, react in the moment"), lifting dynamic-task success sharply β the Real-Time Execution frontier. Related pushes on the interface to vision: S2 (act from a learned visual-evidence budget) and SDPG (end-to-end pixelsβactions RL that can beat state-distillation on hard tasks) both argue less-processed / less-distilled visual input can generalize better.
- Instruction grounding is being manufactured, not just collected. CounterAlign synthesizes counterfactual (instruction, action) mismatches from existing demos to learn an instruction-grounded offline-RL reward β a data-free complement to the Multi-Task VLA instruction-collapse fixes.
(Caveat: these insights rest on the confirmed-so-far set; they should be revisited against the full program.)
4. Caveats & how to extend
- Acceptance verification: β entries have an author/explicit CoRL-2026 statement; β οΈ entries are reported/likely and should be re-checked against the official OpenReview list. ManiFlow (2509.01819) is CoRL 2025, not 2026 β deliberately excluded.
- Coverage is partial by construction β 687 papers accepted; this lists a handful whose acceptance is publicly traceable now. Expect major gaps (grasping, sim-to-real, planning, multi-robot, mobile manipulation).
- Next step: promote to a full session-taxonomy survey (like IROS 2026) once CoRL publishes the program; add per-paper in-depth pages for the flagship accepts.
5. Links
- Official: corl.org Β· Call for Papers Β· Workshops
- Confirmed papers: see Β§2 (author-confirmed accepts with arXiv/project links, grouped by theme). Adjacent driving papers in Β§2.7.
- Related topic reviews: World Models Β· VLA Hybrid Architectures Β· Multi-Task VLA Β· In-Context Imitation Β· Dexterous-Hand Data Pyramid Β· Humanoid VLA Β· Egocentric Video Pre-Training
- Other venue surveys: IROS 2026 Β· RSS 2026 Β· ICML 2026 Β· ICLR 2026 Β· CVPR 2026 Β· CoRL 2025