ICRA 2026 Topic RL Data - Heungwoo/research GitHub Wiki
ICRA 2026 — RL, Datasets, Sim & Representation for Manipulation (Topic Analysis)
Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program. Paper IDs (e.g.
ThI1I.302) are the program session codes. ← Back to ICRA-2026-VLA-Manipulation-Survey
This is the methodological backbone of the ICRA 2026 manipulation track: 39 papers whose primary contribution is how a manipulation policy is learned, rather than a specific manipulation skill. They cluster around four pillars — reinforcement learning (sample-efficiency, sim-to-real, residual/offline RL), large manipulation datasets & data engines, simulation and benchmarks, and learned representations / world models. Where the VLA papers (see the survey) ask what architecture maps pixels to actions, these papers ask where the training signal comes from and how cheaply it scales. The dominant story this year is data scaling without real-robot cost — real-to-sim reconstruction, video-diffusion data synthesis, automatic affordance annotation — paired with RL that is finally safe and sample-efficient enough to run on real hardware.
Sub-trends
1. RL for manipulation: sample-efficiency, safety, and offline-to-online
The largest cluster makes RL practical on real robots. Failure-Aware RL (TuI2I.331, arXiv 2601.07821) is the headline result: a world-model safety critic plus an offline-trained recovery policy that cuts intervention-requiring failures by 73.1% while raising performance 11.3% during real-world offline-to-online post-training, shipped with a FailureBench of spill/breakage scenarios. SHaRe-RL (TuI1I.222, arXiv 2509.13949) tackles the same sample-efficiency/safety problem for contact-rich industrial assembly by structuring skills into manipulation primitives, folding in human demonstrations and online corrections, and bounding per-axis interaction forces — demonstrated on Harting connector inserts with 0.2–0.4 mm clearance. RM-RL (WeI2I.137) targets precise manipulation (chemistry/biology where spillage invalidates a task) without pre-collected expert demos. Push-Grasp Synergy (TuI1I.196) uses self-supervised deep RL to learn pushing that exposes occluded targets for grasping in clutter, and Collision-Free Object Goal Pushing (ThI2I.269) adds safe-corridor obstacle avoidance to quadruped non-prehensile pushing. The connective tissue with the VLA track is the RL for VLA line and FPO — RL is now the standard post-training lever, not just a from-scratch trainer.
2. Sim-to-real and physical-parameter grounding
Closing the sim-to-real gap is treated as a grounding problem rather than brute-force domain randomization. Phys2Real (ThI1I.302, arXiv 2510.11689) fuses VLM-inferred priors over physical parameters (mass, center-of-mass) with online interactive estimation and uncertainty-aware fusion, lifting a top-weighted T-block push from 23% (domain-randomization baseline) to 57%, and the bottom-weighted case from 79% to 100%. The broader theme — let a foundation model guess the physics, then refine it from a few interactions — recurs across the RL papers, and is the residual/online-adaptation flavor of sim-to-real.
3. Real-to-sim & generative data engines
The most active 2026 trend is generating training data without real-robot teleoperation. Re³Sim (ThI2I.86, arXiv 2502.08645) reconstructs real scenes with 3D-photorealistic real-to-sim (Gaussian-splat geometry + neural rendering inside a physics simulator) to close the geometric and visual gap. AnchorDream (TuI2I.274, arXiv 2512.11797) repurposes pretrained video diffusion as an embodiment-aware world model: it conditions on robot-motion renderings to anchor the embodiment and prevent hallucination, scaling a handful of teleop demos into large datasets for +36.4% in sim and nearly 2× in real. SAGrid (ThI1I.340) attacks the asset bottleneck with automatic affordance annotation on in-the-wild 3D assets, scaling simulation-ready content for household tasks. These three are complementary points on the real→sim→data spectrum: reconstruct (Re³Sim), hallucinate-but-anchor (AnchorDream), and auto-annotate at scale (SAGrid).
4. Simulation environments & benchmarks
Beyond data engines, several papers ship simulators and benchmarks as the contribution. LeHome (ThI2I.306, arXiv 2604.22363) is a high-fidelity deformable-object household simulator spanning six physical categories (liquid, gaseous fluid, granular, linear, thin-shell, volumetric) with PBD/FEM dynamics, a causal Action Graph, and a focus on low-cost embodiments. Judo (TuI2I.199) is an open-source package for sampling-based MPC built for reproducibility and benchmarking. GraspClutter6D (below) doubles as both a dataset and a perception/grasping benchmark for cluttered scenes.
5. Large manipulation datasets & data curation
Datasets remain a primary artifact. GraspClutter6D (WeI2I.39, arXiv 2504.06866) is a large real-world cluttered-grasping dataset — 1,000 scenes, 14.1 objects/scene, 62.6% occlusion, 736K 6D poses, 9.3B feasible grasps over 52K RGB-D images — that benchmarks SOTA segmentation, pose estimation, and grasp detection. On the curation side, RoboSQ (WeI2I.209) introduces semantic queries to extract task-aligned subsets (e.g. only well-captioned or cooking-related demos) from large, noisy datasets — the data-engineering counterpart to the data-generation engines above. MLLM-Fabric (ThI1I.363, arXiv 2507.04351) releases a 220-fabric multimodal dataset (RGB + visuotactile + pressure) for fabric sorting/selection.
6. Representation, world-model & flow-policy learning
A cluster improves the representation a policy learns over. FlowDreamer (WeI1I.403, arXiv 2505.10075) is an RGB-D world model that makes motion explicit via 3D scene flow (U-Net predicts flow, diffusion predicts the next frame), beating baseline RGB-D world models by ~7–11% on prediction/planning. Dense-Jump Flow Matching (ThBT2.2, arXiv 2509.13574) diagnoses why adding integration steps degrades flow-matching policies (late-time oversampling + non-Lipschitz velocity near t=1) and fixes it with U-shaped time scheduling + a dense-jump inference schedule, for up to +23.7% — directly relevant to the FPO flow-policy line. ManiMorph (TuI1LB.17) models held objects as dynamic extensions of the kinematic chain (morphology-aware multi-task representation). Denoising Particle Filters (WeI2I.38) reframes learned state estimation with single-step denoising objectives.
7. Self-supervised, transfer & VLM-driven control
VLMs and self-supervision show up as the high-level reasoner or feature source over RL/learned skills. Hierarchical DLO Routing (TuAT1.2, arXiv 2510.19268) pairs in-context VLM planning with RL low-level skills plus a failure-recovery reorientation step for deformable-linear-object routing, reaching 92.5% success (~+50% over the next baseline). I-FailSense (ThI2I.75, arXiv 2509.16072) post-trains a VLM with lightweight per-layer classification heads for general robotic failure detection, generalizing zero-shot from semantic-misalignment training to broader failure categories. MLLM-Fabric (above) uses an MLLM with explanation-guided distillation. Neuromorphic Event-Camera Grasping (WeI1I.394) uses transfer learning in a multi-task event-vision model for grasp-position detection.
8. Co-design, failure-active control & adjacent learning-for-control
A tail of papers applies learned/reward models to design and resilience. Learning to Design Soft Hands Using Reward Models (TuI1I.46, arXiv 2510.17086) optimizes tendon-driven soft-hand morphology with a Cross-Entropy Method + learned reward (CEM-RM) from teleop data, halving design evaluations. Fail-Active Trajectory Generation (DEFT) (WeI1I.294) trains diffusion policies conditioned on embodiment and task so a damaged robot still completes its task. OmniRetarget (WeAT1.4, arXiv 2509.26633) is an interaction-preserving data-generation engine for humanoid loco-manipulation, augmenting one demo across embodiments/terrains for proprioceptive RL on a Unitree G1.
Standout deep-dives
Failure-Aware RL — safe real-world offline-to-online RL (TuI2I.331)
arXiv 2601.07821. A world-model safety critic + offline recovery policy that reduces intervention-requiring failures 73.1% and improves performance 11.3% on real-world RL post-training, with a FailureBench of common intervention scenarios. The clearest answer this year to "RL on real robots is too dangerous." Connects to the RL for VLA post-training theme.
Phys2Real — VLM physics priors + interactive adaptation (ThI1I.302)
arXiv 2510.11689. Real-to-sim-to-real with three parts: Gaussian-splat geometry, VLM-inferred priors over physical parameters, and online interaction-based estimation fused by uncertainty. Top-weighted T-block push 23% → 57%, bottom-weighted 79% → 100% over domain randomization. The "let the VLM guess the physics" sim-to-real recipe.
AnchorDream — video diffusion as an embodiment-anchored data engine (TuI2I.274)
arXiv 2512.11797. Conditions a pretrained video-diffusion model on robot-motion renderings to anchor the embodiment (no hallucinated arms) while synthesizing novel objects/scenes, scaling a few teleop demos into large datasets: +36.4% sim, ~2× real. The strongest generative-data result of the group.
GraspClutter6D — the cluttered-grasping dataset/benchmark (WeI2I.39)
arXiv 2504.06866. 1,000 scenes · 14.1 objects/scene · 62.6% occlusion · 736K 6D poses · 9.3B grasps · 52K RGB-D images across bins/shelves/tables and four cameras. A reference dataset and benchmark for perception+grasping under heavy occlusion (RA-L).
Dense-Jump Flow Matching — fixing multi-step flow-policy degradation (ThBT2.2)
arXiv 2509.13574. Diagnoses the counterintuitive "more Euler steps → worse policy" effect (late-time oversampling + non-Lipschitz velocity near t=1) and fixes it with non-uniform (U-shaped) training time + dense-jump inference: up to +23.7%. Direct methodological input to the flow-matching VLA / FPO line.
Hierarchical DLO Routing — VLM planner + RL skills (TuAT1.2)
arXiv 2510.19268. In-context VLM high-level planning over RL-trained low-level skills, with a failure-recovery reorientation step, for long-horizon cable/rope routing: 92.5% success, ~+50% over the next baseline. A clean template for VLM-on-top-of-RL hierarchies.
Complete paper list (39)
| code | Title | arXiv |
|---|---|---|
| ThBT2.2 | Dense-Jump Flow Matching with Non-Uniform Time Scheduling for Robotic Policies: Mitigating Multi-Step Inference Degradation | 2509.13574 |
| ThI1I.109 | Reference-Free Sampling-Based Model Predictive Control | 2511.19204 |
| ThI1I.302 | Phys2Real: Fusing VLM Priors with Interactive Online Adaptation for Uncertainty-Aware Sim-To-Real Manipulation | 2510.11689 |
| ThI1I.340 | SAGrid: Scaling Robot Simulation through Automatic Affordance Annotation on In-The-Wild 3D Assets | — |
| ThI1I.363 | MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection | 2507.04351 |
| ThI2I.189 | Impact-Robust Posture Optimization for Aerial Manipulation | 2602.13762 |
| ThI2I.269 | Learning Collision-Free Object Goal Pushing for Quadruped Robots with Safe Corridors | — |
| ThI2I.306 | LeHome: A Simulation Environment for Deformable Object Manipulation in Household Scenarios | 2604.22363 |
| ThI2I.356 | The Challenges of Using Robots to Automate the Recycling of Electronic Devices | — |
| ThI2I.75 | I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models | 2509.16072 |
| ThI2I.86 | Re³Sim: Generating High-Fidelity Simulation Data Via 3D-Photorealistic Real-To-Sim for Robotic Manipulation | 2502.08645 |
| TuAT1.2 | Hierarchical DLO Routing with Reinforcement Learning and In-Context Vision-Language Models | 2510.19268 |
| TuI1I.196 | Learning Push-Grasp Synergy for Occluded Objects in Cluttered Environments | — |
| TuI1I.222 | SHaRe-RL: Structured, Interactive Reinforcement Learning for Contact-Rich Industrial Assembly Tasks | 2509.13949 |
| TuI1I.46 | Learning to Design Soft Hands Using Reward Models | 2510.17086 |
| TuI1I.87 | Shell-Type Soft Jig for Holding Objects During Disassembly | 2509.13802 |
| TuI1LB.17 | ManiMorph: Object Representations in Robot Manipulators Morphology for Improving Multi-Task Manipulation Performance | — |
| TuI2I.125 | Robotic Cell Manipulation at the Solid-Liquid Interface for Cryopreservation | — |
| TuI2I.153 | Integrated Hydrogel Patterning and Dynamic Microparticle Manipulation Using Optoelectronic Tweezers | — |
| TuI2I.199 | Judo: A User-Friendly Open-Source Package for Sampling-Based Model Predictive Control | 2506.17184 |
| TuI2I.22 | Safety-Critical and Distributed Nonlinear Predictive Controllers for Teams of Quadrupedal Robots | 2503.14656 |
| TuI2I.268 | Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning | 2509.13534 |
| TuI2I.274 | AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis | 2512.11797 |
| TuI2I.331 | Failure-Aware RL: Reliable Offline-To-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation | 2601.07821 |
| TuI2I.374 | Controlling Deformable Objects with Non-Negligible Dynamics: A Shape-Regulation Approach to End-Point Positioning | 2402.16114 |
| WeAT1.4 | OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction | 2509.26633 |
| WeI1I.294 | Moving On, Even When You're Broken: Fail-Active Trajectory Generation Via Diffusion Policies Conditioned on Embodiment and Task (DEFT) | 2602.02895 |
| WeI1I.329 | The Translational/Rotational Piezoelectric Impact Drive Mechanism for Cell/Tissue Extraction from Mouse Cranial Window | — |
| WeI1I.394 | Neuromorphic Event Camera-Based Object Recognition and Grasping Position Detection Using a Transfer Learning-Enhanced Multi-Task Model | — |
| WeI1I.403 | FlowDreamer: A RGB-D World Model with Flow-Based Motion Representations for Robot Manipulation | 2505.10075 |
| WeI2I.130 | ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models | 2603.01490 |
| WeI2I.137 | RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation | 2510.15189 |
| WeI2I.178 | IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human–Robot Interaction | 2510.07778 |
| WeI2I.209 | RoboSQ: Semantic Queries for Task-Aligned Robot Training Data | — |
| WeI2I.219 | CASSR: Continuous A-Star Search through Reachability for Real Time Footstep Planning | 2603.02989 |
| WeI2I.257 | Magnetic-Acoustic Microbubble Microrobot for Targeted Mechanical Stimulation of Cancer Cells | — |
| WeI2I.38 | Denoising Particle Filters: Learning State Estimation with Single-Step Objectives | 2602.19651 |
| WeI2I.387 | Primal-Dual iLQR for GPU-Accelerated Learning and Control in Legged Robots | 2506.07823 |
| WeI2I.39 | GraspClutter6D: A Large-Scale Real-World Dataset for Robust Perception and Grasping in Cluttered Scenes | 2504.06866 |
Related
← Back to ICRA-2026-VLA-Manipulation-Survey