RSS 2026 SID - Heungwoo/research GitHub Wiki
SID: Sliding into Distribution for Robust Few-Demonstration Manipulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #8 Authors: Yicheng Ma, Wei Yu, Zhian Su, Xidan Zhang, Huixu Dong arXiv: 2605.13428 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

The teaser shows SID's two-regime view of a manipulation episode: gripper poses starting out-of-distribution follow a learned object-centric motion field (black arrows) that "slides" them into the demonstrated in-distribution region (dashed funnel), where a lightweight egocentric execution policy (blue arrows) takes over for the contact-rich interaction — illustrated by the energy-landscape inset whose gradient vanishes near the demonstration manifold.
Problem
Few-demonstration visuomotor policies fail mostly through distribution shift: with low-coverage data, test-time object poses, viewpoints, and disturbances put the robot in states the demonstrations never covered. The authors observe that an episode has two qualitatively different regimes — an underdetermined approach phase and a locally sensitive, interaction-rich execution phase — and that a single end-to-end policy must resolve both at once, making it brittle under pose shifts and perturbations.
Method
SID (Grasp Lab, Zhejiang University + Torch Kernel Co.) factorizes control into four components: (i) an object-centric motion field f_θ learned from canonicalized approach-phase demonstrations, implemented as a gradient-descent-style dynamical system over a pose-aligned SE(3) potential that produces large corrective motions far from the demonstration manifold and vanishes near convergence; (ii) a kinematically consistent point-cloud reprojection augmentation that perturbs the end-effector pose, reprojects segmented wrist-camera point clouds via fixed hand–eye calibration, and updates relative actions to preserve action–observation consistency (also generating ID/OOD labels); (iii) an egocentric execution policy trained with conditioned flow matching on wrist-centric point clouds and gripper width; and (iv) two inference pipelines — open-loop handoff (SID-O) and a closed-loop variant (SID-C) whose auxiliary ID-confidence head triggers field-based re-alignment when observations drift OOD. Canonicalization uses 6D pose estimation and SAM2/SAM3 object segmentation. Each task uses only two raw demonstrations, expanded to 100 training samples by augmentation.
Results
On six real-world tasks (Open Drawer, Pour Water, Hang Tape, Hang Cup, PnP-Box, Multi-PnP-Box; 50 trials each), SID-C reaches 86–92% success under OOD initializations with 2 demos, versus π0.5 at 0–14% OOD (trained with 100 demos) and retrieval baselines MT3/Ret-BC at 36–78% (10 demos); ACT and DP3 largely collapse OOD. In the dynamic-disturbance setting SID-C scores 82–88% across four tasks, beating the strongest baseline π0.5 (68–80%), and both SID variants stay reliable in cluttered scenes and on recomposed long-horizon tasks built by reusing learned sub-skills. The abstract's headline: ~90% OOD success with two demonstrations and under a 10% drop with distractors and external disturbances.
Significance
A clean articulation of "online distribution recovery" as an alternative to scaling data: instead of covering the state space, steer the system back to where the policy is competent. Complements the equivariance and coarse-to-fine lines discussed in Review-Dexterous-Manipulation, and its OOD-confidence-gated re-alignment echoes the runtime-monitoring themes in Review-VLA-Evaluation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home