ICRA 2026 Topic Perception - Heungwoo/research GitHub Wiki
ICRA 2026 โ Perception for Manipulation (Topic Analysis)
Venue: IEEE ICRA 2026 ยท Vienna, Austria ยท June 1โ5, 2026 Compiled against the official ICRA 2026 PaperCept program. Numbers below are taken from the official program and from authors' arXiv/project pages; where a number could not be confirmed against an authoritative source it is omitted rather than guessed.
This page covers the 82 papers in the ICRA 2026 "Perception for Manipulation" cluster โ work whose primary contribution is seeing well enough to act, rather than the action policy itself. Across modern robot learning, the policy architecture has commoditized faster than the perception that feeds it: diffusion/flow policies and VLAs are increasingly off-the-shelf, but they remain only as good as the object pose, 3D geometry, affordance, or correspondence handed to them. Perception is the manipulation bottleneck, and this cluster is where ICRA 2026 attacks it. Three throughlines dominate: (1) 3D and 4D scene representation (point clouds, Gaussian splatting, neural fields, meshes) as the substrate manipulation policies want; (2) pose, shape, and affordance estimation for known, category-level, and fully novel objects, increasingly with explicit uncertainty; and (3) foundation-model perception (SAM, DINO, video-diffusion, MLLMs) repurposed as zero-shot front-ends for grasping and flow. A recurring sub-current is hard-case perception โ transparent/specular glassware, deformable linear objects (cables/cloth), articulated mechanisms, and heavily occluded clutter โ exactly the cases where commodity RGB-D and discriminative depth break.
Sub-trends
1. 6-DoF object pose estimation & tracking
Classical 6-DoF pose remains a live research target, now pushed toward distributions and dynamics rather than single point estimates. SE(3)-PoseFlow (WeI1I.194; arXiv 2511.01501) does flow-matching on the SE(3) manifold to produce a full sample-based pose distribution, explicitly modeling multi-modality from symmetry and occlusion and reporting SOTA on Real275/YCB-V/LM-O. MGS-Track (TuI2I.316) tracks monocular 6-DoF pose via a masked 3D prior plus online Gaussian splatting, targeting depth-sensor-free deployment. Learning 6D Object Pose Estimation with Event Cameras (TuI1I.397) attacks high-speed and adverse-lighting regimes where RGB/RGB-D fail, training on synthetic data with domain randomization. PartPose (ThI2I.337) reframes 6D pose for multi-part deformable objects (cable-attached appliances, pouch drinks) by attending to graspable parts. DynOPETs (WeI2I.33; arXiv 2503.19625) is the supporting benchmark: 175 object instances under simultaneous camera and object motion with synchronized 6-DoF annotations, evaluating 18 methods โ squarely targeting the moving-camera/moving-object regime that static pose datasets ignore.
2. Category-level & novel-object pose / detection
Where instance-level pose assumes a CAD model, this thread drops that assumption. Category-Level Object Shape and Pose Estimation in Less Than a Millisecond (TuI2I.250; arXiv 2509.18979, MIT-SPARK) is the standout: a self-consistent-field solver over a linear active-shape model that runs one iteration in ~100 microseconds with a certificate of global optimality โ fast enough to be used as an outlier-rejection inner loop. GFreeDet2 (WeAT3.1; building on GFreeDet, arXiv 2412.01552, BOP-Challenge-2024 winner) goes fully model-free: reconstruct 3D Gaussian object models from multi-view RGB references, then do 2D + 6D detection of unseen objects via SAM/DINOv2 mask matching. PIRATR (WeI1I.214) does parametric 3D object inference with transformers directly in point clouds, jointly estimating multi-class 6-DoF poses and class-specific parameters. Plug-And-Play Shape Matching (TuI1I.353) refines grasps on unknown objects with a mesh-free, training-free geometric module.
3. 3D / point-cloud / 4D representations for manipulation
The largest cluster: what representation should feed a policy. FP3 (TuAT1.3; arXiv 2503.08950) is a 3D foundation policy โ a DiT pre-trained on 60k point-cloud trajectories with a Uni3D encoder, learning new tasks at >90% success from only 80 demos. GP3 (ThI1I.52) builds a geometry-aware policy from multi-view images; VO-DP (WeI1I.252) argues for vision-only semantic-geometric features as an alternative to point clouds; 3D Dynamics-Aware Manipulation (TuI2I.323) adds explicit depth-wise 3D foresight to world-model policies. Gaussian splatting recurs as the 3D substrate: GAF (ThI1I.185; arXiv 2506.14135) makes a Gaussian Action Field โ a 4D representation that jointly reconstructs the current scene, predicts future frames, and estimates action via Gaussian motion (a "Vision-to-4D-to-Action" paradigm); Informative Object-Centric Next Best View (ThI2I.51) and Real-To-Sim with Gaussian Splatting of Soft-Body Interactions (ThI2I.288) use 3DGS for active sensing and policy evaluation respectively. Meshes and latent maps also appear: Subsecond 3D Mesh Generation (WeI2I.230; arXiv 2512.24428) produces a manipulation-ready mesh from a single RGB-D image in <1s (92% pick-and-place success), and Seeing the Bigger Picture (WeI1I.89) shows a 3D latent map beats image-only policies for mobile manipulation.
4. Affordance detection & grounding
Affordances bridge perception and action โ where and how to interact. Coupled Particle Filters for Robust Affordance Estimation (TuI1I.225) disambiguates graspable vs. movable regions with two coupled recursive estimators. Visual Category-Guided One-Shot Open Affordance Grounding (TuI2I.345) leverages visual foundation models for one-shot open-vocabulary grounding. RoboPCA (WeI2I.198) learns pose-centered affordances (contact regions + contact poses) from human demos. NaturalVLM (WeI2I.5) and T-FunS3D (ThI2I.137) push toward language-conditioned/3D functionality grounding โ the latter doing task-driven hierarchical open-vocabulary 3D functionality segmentation. AdapGrasp (TuI1I.86) pairs a stiffness-and-affordance dataset with a transformer grasp model, while RoboHitch (TuI1I.97) learns visual affordance from disordered keypoints for knot tying.
5. Articulated-object perception
Understanding kinematic structure is a perception problem in its own right. PokeNet (TuI1I.242) learns joint parameters of articulated objects from human observations without CAD priors. Kinematify (ThI1I.333; arXiv 2511.01294) synthesizes high-DoF articulated objects (exported as URDF) from a single RGB image or text via MCTS structural inference plus geometry-driven joint optimization โ no motion data required. UniDoorManip (ThI1I.330; arXiv 2403.02604) builds a large-scale door environment (6 categories, thousands of instances) and learns a universal door policy from partial/occluded point clouds, a canonical articulated-perception-for-control setting.
6. Transparent / specular & deformable perception (the hard cases)
RGB-D's failure modes get dedicated treatment. Diffusion Knows Transparency (TuI2I.90) repurposes a video-diffusion prior for transparent-object depth and normals where stereo/ToF break. TORM (WeI1I.362) reconstructs and manipulates transparent objects via multi-view segmentation; SilRef (ThI2I.406) jointly optimizes visual silhouette and tactile pose for transparent manipulation; SPILL (ThI1I.383) estimates size, pose, and internal liquid level of transparent glassware for bartending. Liquid recurs in Toward Multimodal Liquid-Level Estimation (TuI2LB.16). Deformables form a parallel hard-case cluster: CloSE (TuI2I.219; arXiv 2504.05033) is a compact shape-/orientation-agnostic cloth-state representation built on a topological dGLI disk; CVF-DLO (WeI1I.16) and Interactive Robotic Moving Cable Segmentation (ThI2I.364) tackle tangled/branched cable routing and motion-correlation segmentation; Uncertainty-Aware Stereo Grasp Point Selection (TuI1LB.8) adds prediction-reliability to cable grasping.
7. Visual servoing & active/interactive perception
Closed-loop perception-for-control. Perception-Control Coupled Visual Servoing for Textureless Objects (TuI1I.68) uses a keypoint-based EKF to servo on feature-poor surfaces; An Autonomous Hardware-Agnostic Vision-Servoed System (TuI1I.325) servos nanoliter microdevice injection without custom calibration. Active/interactive perception threads through RUMI (ThI2I.374; arXiv 2408.10450), which plans contact-rich rummaging via mutual information between object-pose belief and trajectory in occluded bins; COMPASS (TuI1I.294), active sensing in confined spaces; Active-Perceptive Language-Oriented Grasp (TuI1I.431) for heavily cluttered scenes; and SceneComplete (TuAT3.5; arXiv 2410.23643), which composes pretrained perception modules into a full segmented 3D scene from a single view for robust grasping.
8. Keypoint / dense-correspondence & flow
Correspondence as the perceptual interface to action. Sparse Meets Dense (WeI1I.228) fuses sparse and dense correspondence for rigid-deformable interactions (hanging clothes, dressing). NovaFlow (ThI1I.134; arXiv 2510.08568) distills generated videos into 3D actionable object flow, transferring zero-shot across a Franka and a Spot; Dream2Flow (TuI2I.123; arXiv 2512.24766) likewise reconstructs 3D object flow from generated video as an embodiment-agnostic interface for open-world manipulation; Actron3D (ThI2I.114) learns actionable neural functions from uncalibrated RGB-only human video for transferable 6-DoF skills.
9. Foundation-model perception (SAM / DINO / MLLM / video-diffusion) for manipulation
A cross-cutting enabler rather than a single application. SAM/DINO underpin GFreeDet2 (WeAT3.1) and Subsecond 3D Mesh Generation (Florence-2 + SAM2 + Depth-Anything-v2). MLLMs are distilled into 3D for grasping in Point2Act (WeI1I.92; arXiv 2508.03099), which builds 3D relevancy fields and produces a grasp in <20s. Video-diffusion priors drive Diffusion Knows Transparency (TuI2I.90), NovaFlow, and Dream2Flow. VERM (ThI1I.407) leverages foundation models to construct a "virtual eye" that prunes multi-camera redundancy for 3D manipulation, and Clutt3R-Seg (WeI2I.181) does sparse-view 3D instance segmentation for language-grounded grasping.
Standout deep-dives
These papers were cross-checked against an authoritative source (arXiv + project/GitHub) for both identity and a concrete number.
- Category-Level Object Shape & Pose in <1 ms (TuI2I.250) โ arXiv 2509.18979 (MIT-SPARK). Self-consistent-field iteration over a linear active-shape model; ~100 ยตs per iteration, with a global-optimality certificate. Reframes category-level pose as a tiny eigenproblem fast enough to use for outlier rejection. Code.
- FP3: A 3D Foundation Policy (TuAT1.3) โ arXiv 2503.08950. DiT pre-trained on 60k point-cloud trajectories (Uni3D encoder); fine-tunes to new tasks at >90% success from 80 demos in novel environments with unseen objects. A 3D answer to image-only foundation policies.
- SE(3)-PoseFlow (WeI1I.194) โ arXiv 2511.01501. Flow-matching on SE(3) yields a full pose distribution (not a point estimate); SOTA on Real275, YCB-V, LM-O, with downstream active-perception and uncertainty-aware grasp synthesis.
- GAF: Gaussian Action Field (ThI1I.185) โ arXiv 2506.14135. 4D "Vision-to-4D-to-Action" representation extending 3DGS with learnable motion; +15.7% success over Diffusion Policy and +7.3% over Act3D, real-time on a single GPU.
- NovaFlow (ThI1I.134) โ arXiv 2510.08568 (RAI Institute / Brown). Synthesizes a video from language, distills 3D actionable object flow, and executes zero-shot across rigid/articulated/deformable objects on a Franka and a Spot โ no demonstrations, no embodiment-matched data.
- GFreeDet2 (WeAT3.1) โ built on GFreeDet, arXiv 2412.01552, best-overall + best-fast in the model-free 2D-detection track of BOP Challenge 2024. Gaussian-splatting object models from RGB references + SAM/DINOv2 enable model-free 2D+6D detection of unseen objects.
- Subsecond 3D Mesh Generation (WeI2I.230) โ arXiv 2512.24428 (Yale). Single RGB-D image โ manipulation-ready mesh in <1s (Florence-2 + SAM2 + Depth-Anything-v2 + FlashVDM-distilled Hunyuan3D 2.0); 92% pick-and-place success.
Complete paper list (82)
| Code | Title | arXiv |
|---|---|---|
| ThI1I.114 | Latent Representations for Visual Proprioception in Inexpensive Robots | 2504.14634 |
| ThI1I.134 | NovaFlow: Zero-Shot Manipulation Via Actionable Flow from Generated Videos | 2510.08568 |
| ThI1I.185 | GAF: Gaussian Action Field As a 4D Representation for Dynamic World Modeling in Robotic Manipulation | 2506.14135 |
| ThI1I.187 | Beyond Domain Randomization: Event-Inspired Perception for Visually Robust Adversarial Imitation from Videos | 2505.18899 |
| ThI1I.195 | Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery | 2603.03181 |
| ThI1I.198 | OHMM-PA: A Learning from Demonstration Approach Using Online Hidden Markov Models with Path Planning | โ |
| ThI1I.212 | CoVAR: Co-Generation of Video and Action for Robotic Manipulation Via Multi-Modal Diffusion | 2512.16023 |
| ThI1I.25 | Haptic Stiffness Perception Using Hand Exoskeletons in Tactile Robotic Telemanipulation | 2412.02613 |
| ThI1I.330 | UniDoorManip: Learning Universal Door Manipulation Policy Over Large-Scale and Diverse Door Manipulation Environments | 2403.02604 |
| ThI1I.333 | Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects | 2511.01294 |
| ThI1I.383 | SPILL: Size, Pose, and Internal Liquid Level Estimation of Transparent Glassware for Robotic Bartending | โ |
| ThI1I.407 | VERM: Leveraging Foundation Models to Create a Virtual Eye for Efficient 3D Robotic Manipulation | 2512.16724 |
| ThI1I.52 | GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation | 2509.15733 |
| ThI1I.99 | Mash, Spread, Slice! Learning to Manipulate Object States Via Visual Spatial Progress | 2509.24129 |
| ThI2I.114 | Actron3D: Learning Actionable Neural Functions from Videos for Transferable Robotic Manipulation | 2510.12971 |
| ThI2I.121 | Improving Robotic Manipulation Robustness Via NICE Scene Surgery | 2511.22777 |
| ThI2I.137 | T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation | โ |
| ThI2I.182 | Instrumentation for Imitation Learning: Enhancing Training Datasets for Clothes Hanger Insertion | 2605.23847 |
| ThI2I.197 | OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-To-Robot Action Transfer | 2603.14401 |
| ThI2I.258 | VistaBot: View-Robust Robot Manipulation Via Spatiotemporal-Aware View Synthesis | 2604.21914 |
| ThI2I.288 | Real-To-Sim Robot Policy Evaluation with Gaussian Splatting Simulation of Soft-Body Interactions | 2511.04665 |
| ThI2I.337 | PartPose: Attentive 6D Pose Estimation by Focusing on Graspable Parts of Multi-Part Deformable Objects | โ |
| ThI2I.364 | Interactive Robotic Moving Cable Segmentation by Motion Correlation | โ |
| ThI2I.374 | RUMI: Rummaging Using Mutual Information | 2408.10450 |
| ThI2I.406 | SilRef: Joint Visual Silhouette and Tactile Pose Optimization for Transparent Object Manipulation | โ |
| ThI2I.51 | Informative Object-Centric Next Best View for Object-Aware 3D Gaussian Splatting in Cluttered Scenes | โ |
| TuAT1.3 | FP3: A 3D Foundation Policy for Robotic Manipulation | 2503.08950 |
| TuAT3.5 | SceneComplete: Open-World 3D Scene Completion in Cluttered Real World Environments for Robot Manipulation | 2410.23643 |
| TuAT3.6 | Robust Bayesian Scene Reconstruction with Retrieval-Augmented Priors for Precise Grasping and Planning | 2411.19461 |
| TuI1I.12 | Distributional Treatment of Real2Sim2Real for Object-Centric Agent Adaptation in Vision-Driven DLO Manipulation | 2502.18615 |
| TuI1I.225 | Coupled Particle Filters for Robust Affordance Estimation | 2603.15223 |
| TuI1I.242 | PokeNet: Learning Kinematic Models of Articulated Objects from Human Observations | 2602.02741 |
| TuI1I.294 | COMPASS: Confined-Space Manipulation Planning with Active Sensing Strategy | 2509.14787 |
| TuI1I.325 | An Autonomous and Hardware-Agnostic Vision-Servoed System for Microdevice Injection | โ |
| TuI1I.353 | Plug-And-Play Shape Matching Module for Zero-Shot Mesh-Free Grasp Refinement on Unknown Objects | โ |
| TuI1I.364 | Fine-Grained Classification for Depth Estimation from Monocular Microscopy for Robotic Micromanipulation of Motile Cells | โ |
| TuI1I.397 | Learning 6D Object Pose Estimation with Event Cameras Using Synthetic Data and Domain Randomization | โ |
| TuI1I.431 | Active-Perceptive Language-Oriented Grasp Policy for Heavily Cluttered Scenes | โ |
| TuI1I.68 | Perception-Control Coupled Visual Servoing for Textureless Objects Using Keypoint-Based EKF | 2602.06834 |
| TuI1I.86 | AdapGrasp: A Stiffness and Grasp Affordance Dataset with a Transformer-Based Adaptive Grasp Model | โ |
| TuI1I.97 | RoboHitch: Learning Visual Affordance from Disordered Keypoints for Hitch Knots Tying | 2605.24394 |
| TuI1LB.8 | Uncertainty-Aware Stereo Grasp Point Selection for Deformable Linear Objects | โ |
| TuI2I.123 | Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow | 2512.24766 |
| TuI2I.198 | Tactile Memory for Continuous Policy Blending in Unified Force-Impedance Control | โ |
| TuI2I.219 | CloSE: A Geometric Shape-Agnostic Cloth State Representation | 2504.05033 |
| TuI2I.250 | Category-Level Object Shape and Pose Estimation in Less Than a Millisecond | 2509.18979 |
| TuI2I.316 | MGS-Track: Monocular 6DoF Pose Tracking Via Masked 3D Prior and Online Gaussian Splatting | โ |
| TuI2I.323 | 3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight | 2502.10028 |
| TuI2I.336 | The Price Is Not Right: Neuro-Symbolic Methods Outperform VLAs on Structured Long-Horizon Manipulation Tasks with Significantly Lower Energy Consumption | 2602.19260 |
| TuI2I.345 | Visual Category-Guided One-Shot Open Affordance Grounding | โ |
| TuI2I.364 | ILeSiA: Interactive Learning of Robot Situational Awareness from Camera Input | 2409.20173 |
| TuI2I.436 | Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies towards Visual Robustness | 2505.08627 |
| TuI2I.90 | Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation | 2512.23705 |
| TuI2LB.16 | Toward Multimodal Liquid-Level Estimation for Closed-Loop Robotic Pouring | โ |
| WeAT3.1 | GFreeDet2: Exploiting Gaussian Splatting and Foundation Models for RGB-Based Model-Free 2D and 6D Detection of Unseen Objects | 2412.01552 |
| WeI1I.131 | Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation | 2509.17125 |
| WeI1I.159 | Bi-Manual Joint Camera Calibration and Scene Representation | 2505.24819 |
| WeI1I.16 | CVF-DLO: Cross-Visual-Field Branched Deformable Linear Objects Route Estimation | โ |
| WeI1I.194 | SE(3)-PoseFlow: Estimating 6D Pose Distributions for Uncertainty-Aware Robotic Manipulation | 2511.01501 |
| WeI1I.214 | PIRATR: Parametric Object Inference for Robotic Applications with Transformers in 3D Point Clouds | 2602.05557 |
| WeI1I.228 | Sparse Meets Dense: Correspondence Guided Robotic Manipulation with Rigid-Deformable Interactions | โ |
| WeI1I.252 | VO-DP: Semantic-Geometric Adaptive Diffusion Policy for Vision-Only Robotic Manipulation | 2510.15530 |
| WeI1I.268 | Ego-Vision World Model for Humanoid Contact Planning | 2510.11682 |
| WeI1I.271 | GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning | 2602.04231 |
| WeI1I.362 | TORM: Transparent Objects Reconstruction and Manipulation with Multi-View Segmentation | โ |
| WeI1I.382 | Fixture-Free Automated Sewing System Using Dual-Arm Manipulator and High-Speed Fabric Edge Detection | โ |
| WeI1I.88 | EdgeGrasp: Enhancing Edge Perception for 7-DoF Grasping Pose Estimation in Cluttered Scenes | โ |
| WeI1I.89 | Seeing the Bigger Picture: 3D Latent Mapping for Mobile Manipulation Policy Learning | 2510.03885 |
| WeI1I.92 | Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping | 2508.03099 |
| WeI2I.131 | Learning to Grasp by Integrating Human Preferences and Success Feedback | โ |
| WeI2I.181 | Clutt3R-Seg: Sparse-View 3D Instance Segmentation for Language-Grounded Grasping in Cluttered Scenes | 2602.11660 |
| WeI2I.196 | Visual-Auditory Extrinsic Contact Estimation | 2409.14608 |
| WeI2I.198 | RoboPCA: Pose-Centered Affordance Learning from Human Demonstrations for Robot Manipulation | 2603.07691 |
| WeI2I.223 | From Swept Contact to Pose: Probe-Aware Registration Via Complementary-Shape Docking | โ |
| WeI2I.230 | Subsecond 3D Mesh Generation for Robot Manipulation | 2512.24428 |
| WeI2I.254 | Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets | 2601.09605 |
| WeI2I.289 | Beyond the Patch: Exploring Vulnerabilities of Visuomotor Policies Via Viewpoint-Consistent 3D Adversarial Object | 2603.04913 |
| WeI2I.312 | CAVER: Curious AudioVisual Exploring Robot | 2511.07619 |
| WeI2I.327 | GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-Trained Robot Policy Enhancement | 2511.03400 |
| WeI2I.33 | DynOPETs: A Versatile Benchmark for Dynamic Object Pose Estimation and Tracking in Moving Camera Scenarios | 2503.19625 |
| WeI2I.331 | OmniMap: A General Mapping Framework Integrating Optics, Geometry, and Semantics | 2509.07500 |
| WeI2I.5 | NaturalVLM: Leveraging Fine-Grained Natural Language for Affordance-Guided Visual Manipulation | 2403.08355 |
Related
โ Back to ICRA-2026-VLA-Manipulation-Survey