ICRA 2026 Topic RL Data - Heungwoo/research GitHub Wiki

ICRA 2026 — RL, Datasets, Sim & Representation for Manipulation (Topic Analysis)

Venue: IEEE ICRA 2026 · Vienna, Austria · June 1–5, 2026 Compiled against the official PaperCept program. Paper IDs (e.g. ThI1I.302) are the program session codes. ← Back to ICRA-2026-VLA-Manipulation-Survey

This is the methodological backbone of the ICRA 2026 manipulation track: 39 papers whose primary contribution is how a manipulation policy is learned, rather than a specific manipulation skill. They cluster around four pillars — reinforcement learning (sample-efficiency, sim-to-real, residual/offline RL), large manipulation datasets & data engines, simulation and benchmarks, and learned representations / world models. Where the VLA papers (see the survey) ask what architecture maps pixels to actions, these papers ask where the training signal comes from and how cheaply it scales. The dominant story this year is data scaling without real-robot cost — real-to-sim reconstruction, video-diffusion data synthesis, automatic affordance annotation — paired with RL that is finally safe and sample-efficient enough to run on real hardware.

Sub-trends

1. RL for manipulation: sample-efficiency, safety, and offline-to-online

The largest cluster makes RL practical on real robots. Failure-Aware RL (TuI2I.331, arXiv 2601.07821) is the headline result: a world-model safety critic plus an offline-trained recovery policy that cuts intervention-requiring failures by 73.1% while raising performance 11.3% during real-world offline-to-online post-training, shipped with a FailureBench of spill/breakage scenarios. SHaRe-RL (TuI1I.222, arXiv 2509.13949) tackles the same sample-efficiency/safety problem for contact-rich industrial assembly by structuring skills into manipulation primitives, folding in human demonstrations and online corrections, and bounding per-axis interaction forces — demonstrated on Harting connector inserts with 0.2–0.4 mm clearance. RM-RL (WeI2I.137) targets precise manipulation (chemistry/biology where spillage invalidates a task) without pre-collected expert demos. Push-Grasp Synergy (TuI1I.196) uses self-supervised deep RL to learn pushing that exposes occluded targets for grasping in clutter, and Collision-Free Object Goal Pushing (ThI2I.269) adds safe-corridor obstacle avoidance to quadruped non-prehensile pushing. The connective tissue with the VLA track is the RL for VLA line and FPO — RL is now the standard post-training lever, not just a from-scratch trainer.

2. Sim-to-real and physical-parameter grounding

Closing the sim-to-real gap is treated as a grounding problem rather than brute-force domain randomization. Phys2Real (ThI1I.302, arXiv 2510.11689) fuses VLM-inferred priors over physical parameters (mass, center-of-mass) with online interactive estimation and uncertainty-aware fusion, lifting a top-weighted T-block push from 23% (domain-randomization baseline) to 57%, and the bottom-weighted case from 79% to 100%. The broader theme — let a foundation model guess the physics, then refine it from a few interactions — recurs across the RL papers, and is the residual/online-adaptation flavor of sim-to-real.

3. Real-to-sim & generative data engines

The most active 2026 trend is generating training data without real-robot teleoperation. Re³Sim (ThI2I.86, arXiv 2502.08645) reconstructs real scenes with 3D-photorealistic real-to-sim (Gaussian-splat geometry + neural rendering inside a physics simulator) to close the geometric and visual gap. AnchorDream (TuI2I.274, arXiv 2512.11797) repurposes pretrained video diffusion as an embodiment-aware world model: it conditions on robot-motion renderings to anchor the embodiment and prevent hallucination, scaling a handful of teleop demos into large datasets for +36.4% in sim and nearly 2× in real. SAGrid (ThI1I.340) attacks the asset bottleneck with automatic affordance annotation on in-the-wild 3D assets, scaling simulation-ready content for household tasks. These three are complementary points on the real→sim→data spectrum: reconstruct (Re³Sim), hallucinate-but-anchor (AnchorDream), and auto-annotate at scale (SAGrid).

4. Simulation environments & benchmarks

Beyond data engines, several papers ship simulators and benchmarks as the contribution. LeHome (ThI2I.306, arXiv 2604.22363) is a high-fidelity deformable-object household simulator spanning six physical categories (liquid, gaseous fluid, granular, linear, thin-shell, volumetric) with PBD/FEM dynamics, a causal Action Graph, and a focus on low-cost embodiments. Judo (TuI2I.199) is an open-source package for sampling-based MPC built for reproducibility and benchmarking. GraspClutter6D (below) doubles as both a dataset and a perception/grasping benchmark for cluttered scenes.

5. Large manipulation datasets & data curation

Datasets remain a primary artifact. GraspClutter6D (WeI2I.39, arXiv 2504.06866) is a large real-world cluttered-grasping dataset — 1,000 scenes, 14.1 objects/scene, 62.6% occlusion, 736K 6D poses, 9.3B feasible grasps over 52K RGB-D images — that benchmarks SOTA segmentation, pose estimation, and grasp detection. On the curation side, RoboSQ (WeI2I.209) introduces semantic queries to extract task-aligned subsets (e.g. only well-captioned or cooking-related demos) from large, noisy datasets — the data-engineering counterpart to the data-generation engines above. MLLM-Fabric (ThI1I.363, arXiv 2507.04351) releases a 220-fabric multimodal dataset (RGB + visuotactile + pressure) for fabric sorting/selection.

6. Representation, world-model & flow-policy learning

A cluster improves the representation a policy learns over. FlowDreamer (WeI1I.403, arXiv 2505.10075) is an RGB-D world model that makes motion explicit via 3D scene flow (U-Net predicts flow, diffusion predicts the next frame), beating baseline RGB-D world models by ~7–11% on prediction/planning. Dense-Jump Flow Matching (ThBT2.2, arXiv 2509.13574) diagnoses why adding integration steps degrades flow-matching policies (late-time oversampling + non-Lipschitz velocity near t=1) and fixes it with U-shaped time scheduling + a dense-jump inference schedule, for up to +23.7% — directly relevant to the FPO flow-policy line. ManiMorph (TuI1LB.17) models held objects as dynamic extensions of the kinematic chain (morphology-aware multi-task representation). Denoising Particle Filters (WeI2I.38) reframes learned state estimation with single-step denoising objectives.

7. Self-supervised, transfer & VLM-driven control

VLMs and self-supervision show up as the high-level reasoner or feature source over RL/learned skills. Hierarchical DLO Routing (TuAT1.2, arXiv 2510.19268) pairs in-context VLM planning with RL low-level skills plus a failure-recovery reorientation step for deformable-linear-object routing, reaching 92.5% success (~+50% over the next baseline). I-FailSense (ThI2I.75, arXiv 2509.16072) post-trains a VLM with lightweight per-layer classification heads for general robotic failure detection, generalizing zero-shot from semantic-misalignment training to broader failure categories. MLLM-Fabric (above) uses an MLLM with explanation-guided distillation. Neuromorphic Event-Camera Grasping (WeI1I.394) uses transfer learning in a multi-task event-vision model for grasp-position detection.

8. Co-design, failure-active control & adjacent learning-for-control

A tail of papers applies learned/reward models to design and resilience. Learning to Design Soft Hands Using Reward Models (TuI1I.46, arXiv 2510.17086) optimizes tendon-driven soft-hand morphology with a Cross-Entropy Method + learned reward (CEM-RM) from teleop data, halving design evaluations. Fail-Active Trajectory Generation (DEFT) (WeI1I.294) trains diffusion policies conditioned on embodiment and task so a damaged robot still completes its task. OmniRetarget (WeAT1.4, arXiv 2509.26633) is an interaction-preserving data-generation engine for humanoid loco-manipulation, augmenting one demo across embodiments/terrains for proprioceptive RL on a Unitree G1.

Standout deep-dives

Failure-Aware RL — safe real-world offline-to-online RL (TuI2I.331)

arXiv 2601.07821. A world-model safety critic + offline recovery policy that reduces intervention-requiring failures 73.1% and improves performance 11.3% on real-world RL post-training, with a FailureBench of common intervention scenarios. The clearest answer this year to "RL on real robots is too dangerous." Connects to the RL for VLA post-training theme.

Phys2Real — VLM physics priors + interactive adaptation (ThI1I.302)

arXiv 2510.11689. Real-to-sim-to-real with three parts: Gaussian-splat geometry, VLM-inferred priors over physical parameters, and online interaction-based estimation fused by uncertainty. Top-weighted T-block push 23% → 57%, bottom-weighted 79% → 100% over domain randomization. The "let the VLM guess the physics" sim-to-real recipe.

AnchorDream — video diffusion as an embodiment-anchored data engine (TuI2I.274)

arXiv 2512.11797. Conditions a pretrained video-diffusion model on robot-motion renderings to anchor the embodiment (no hallucinated arms) while synthesizing novel objects/scenes, scaling a few teleop demos into large datasets: +36.4% sim, ~2× real. The strongest generative-data result of the group.

GraspClutter6D — the cluttered-grasping dataset/benchmark (WeI2I.39)

arXiv 2504.06866. 1,000 scenes · 14.1 objects/scene · 62.6% occlusion · 736K 6D poses · 9.3B grasps · 52K RGB-D images across bins/shelves/tables and four cameras. A reference dataset and benchmark for perception+grasping under heavy occlusion (RA-L).

Dense-Jump Flow Matching — fixing multi-step flow-policy degradation (ThBT2.2)

arXiv 2509.13574. Diagnoses the counterintuitive "more Euler steps → worse policy" effect (late-time oversampling + non-Lipschitz velocity near t=1) and fixes it with non-uniform (U-shaped) training time + dense-jump inference: up to +23.7%. Direct methodological input to the flow-matching VLA / FPO line.

Hierarchical DLO Routing — VLM planner + RL skills (TuAT1.2)

arXiv 2510.19268. In-context VLM high-level planning over RL-trained low-level skills, with a failure-recovery reorientation step, for long-horizon cable/rope routing: 92.5% success, ~+50% over the next baseline. A clean template for VLM-on-top-of-RL hierarchies.

Complete paper list (39)

code Title arXiv
ThBT2.2 Dense-Jump Flow Matching with Non-Uniform Time Scheduling for Robotic Policies: Mitigating Multi-Step Inference Degradation 2509.13574
ThI1I.109 Reference-Free Sampling-Based Model Predictive Control 2511.19204
ThI1I.302 Phys2Real: Fusing VLM Priors with Interactive Online Adaptation for Uncertainty-Aware Sim-To-Real Manipulation 2510.11689
ThI1I.340 SAGrid: Scaling Robot Simulation through Automatic Affordance Annotation on In-The-Wild 3D Assets —
ThI1I.363 MLLM-Fabric: Multimodal Large Language Model-Driven Robotic Framework for Fabric Sorting and Selection 2507.04351
ThI2I.189 Impact-Robust Posture Optimization for Aerial Manipulation 2602.13762
ThI2I.269 Learning Collision-Free Object Goal Pushing for Quadruped Robots with Safe Corridors —
ThI2I.306 LeHome: A Simulation Environment for Deformable Object Manipulation in Household Scenarios 2604.22363
ThI2I.356 The Challenges of Using Robots to Automate the Recycling of Electronic Devices —
ThI2I.75 I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models 2509.16072
ThI2I.86 Re³Sim: Generating High-Fidelity Simulation Data Via 3D-Photorealistic Real-To-Sim for Robotic Manipulation 2502.08645
TuAT1.2 Hierarchical DLO Routing with Reinforcement Learning and In-Context Vision-Language Models 2510.19268
TuI1I.196 Learning Push-Grasp Synergy for Occluded Objects in Cluttered Environments —
TuI1I.222 SHaRe-RL: Structured, Interactive Reinforcement Learning for Contact-Rich Industrial Assembly Tasks 2509.13949
TuI1I.46 Learning to Design Soft Hands Using Reward Models 2510.17086
TuI1I.87 Shell-Type Soft Jig for Holding Objects During Disassembly 2509.13802
TuI1LB.17 ManiMorph: Object Representations in Robot Manipulators Morphology for Improving Multi-Task Manipulation Performance —
TuI2I.125 Robotic Cell Manipulation at the Solid-Liquid Interface for Cryopreservation —
TuI2I.153 Integrated Hydrogel Patterning and Dynamic Microparticle Manipulation Using Optoelectronic Tweezers —
TuI2I.199 Judo: A User-Friendly Open-Source Package for Sampling-Based Model Predictive Control 2506.17184
TuI2I.22 Safety-Critical and Distributed Nonlinear Predictive Controllers for Teams of Quadrupedal Robots 2503.14656
TuI2I.268 Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning 2509.13534
TuI2I.274 AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis 2512.11797
TuI2I.331 Failure-Aware RL: Reliable Offline-To-Online Reinforcement Learning with Self-Recovery for Real-World Manipulation 2601.07821
TuI2I.374 Controlling Deformable Objects with Non-Negligible Dynamics: A Shape-Regulation Approach to End-Point Positioning 2402.16114
WeAT1.4 OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction 2509.26633
WeI1I.294 Moving On, Even When You're Broken: Fail-Active Trajectory Generation Via Diffusion Policies Conditioned on Embodiment and Task (DEFT) 2602.02895
WeI1I.329 The Translational/Rotational Piezoelectric Impact Drive Mechanism for Cell/Tissue Extraction from Mouse Cranial Window —
WeI1I.394 Neuromorphic Event Camera-Based Object Recognition and Grasping Position Detection Using a Transfer Learning-Enhanced Multi-Task Model —
WeI1I.403 FlowDreamer: A RGB-D World Model with Flow-Based Motion Representations for Robot Manipulation 2505.10075
WeI2I.130 ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models 2603.01490
WeI2I.137 RM-RL: Role-Model Reinforcement Learning for Precise Robot Manipulation 2510.15189
WeI2I.178 IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human–Robot Interaction 2510.07778
WeI2I.209 RoboSQ: Semantic Queries for Task-Aligned Robot Training Data —
WeI2I.219 CASSR: Continuous A-Star Search through Reachability for Real Time Footstep Planning 2603.02989
WeI2I.257 Magnetic-Acoustic Microbubble Microrobot for Targeted Mechanical Stimulation of Cancer Cells —
WeI2I.38 Denoising Particle Filters: Learning State Estimation with Single-Step Objectives 2602.19651
WeI2I.387 Primal-Dual iLQR for GPU-Accelerated Learning and Control in Legged Robots 2506.07823
WeI2I.39 GraspClutter6D: A Large-Scale Real-World Dataset for Robust Perception and Grasping in Cluttered Scenes 2504.06866

Related

← Back to ICRA-2026-VLA-Manipulation-Survey