RSS 2026 GHOST - Heungwoo/research GitHub Wiki

GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 2 · paper #55 Authors: Sriram Krishna, Ben Eisner, Haotian Zhan, Ying Yuan, Haoyu Zhen, Chuang Gan, Shubham Tulsiani, David Held arXiv: 2606.10025 · program page

Summary compiled from the arXiv paper (v1, CMU Robotics Institute + UMass Amherst); all numbers quoted from the paper. Trend context: RSS 2026 survey.

GHOST hierarchical training and rollout (Figure 1 of arXiv 2606.10025, © the authors)

Figure 1: robot teleop demos on multiple tasks (top) train both the high-level 3D goal predictor π_hi and the low-level goal-conditioned policy π_lo, while optional human demos on a novel task (left) train only π_hi; at rollout on a novel task (bottom), the system alternates Plan (π_hi predicts a 3D end-effector sub-goal, visualized on the point cloud) and Act (π_lo executes toward it).

Problem

Imitation-learned visuomotor policies struggle to generalize beyond the training distribution, and incorporating human video normally requires noisy action retargeting into the robot's embodiment. The paper asks whether factorizing control into an embodiment-agnostic sub-goal predictor and an embodiment-specific controller improves both in-distribution performance and out-of-distribution transfer.

Method

GHOST splits the policy into (i) a high-level π_hi — a decoder-only transformer over frozen DINOv3 patch tokens from multi-view RGB-D (patch tokens augmented with 3D coordinates, plus gripper-pose, Flan-T5 language, and a learnable human/robot embodiment token) that predicts a dense per-patch GMM over the next sub-goal, represented as 4 end-effector 3D keypoints (gripper base, two fingertips, grasp center; MANO palm/thumb/index/grasp-center for human hands) — and (ii) a low-level π_lo, a Diffusion Policy conditioned on a simple spatial interface: predicted 3D keypoints projected into each camera and converted to dense distance-field end-effector heatmaps. Sub-goals are extracted automatically at gripper open/close transitions in robot teleop data (manually annotated for human demos); human demonstrations (hand poses from off-the-shelf trackers, scale-resolved with Grounded-SAM + depth) train only π_hi, keeping π_lo purely robot-trained. Tasks use 17–50 demos each; evaluation is 30 trials per method with bootstrap CIs.

Results

Hierarchy alone helps in-distribution: on fold-onesie, final success jumps from 10% (flat Diffusion Policy) to 80% (GHOST robot-only), and on hammer-pin from 16.7% to 50%. Human demos then unlock OOD transfer: 63.3% on mug-on-table (novel object combination, vs 13.3% DP / 28.3% MimicPlay), 56.7% final on fold-onesie-ood (novel instance, vs 0% MimicPlay), 36.7% on fold-towel (novel category + skill composition, vs 16.7% MimicPlay), and 70% on hammer-pin with a novel target pin. An oracle ablation shows π_lo generalizes zero-shot to towels (40%→90% when π_hi gets robot-quality sub-goals), locating the OOD bottleneck in the high-level human-robot visual domain gap.

Significance

A clean demonstration that 3D sub-goals are a practical embodiment-agnostic interface for folding human video into robot skill learning without action retargeting — directly relevant to the human-data threads in Review-Human-Video-Transfer and the hierarchical-policy discussions in Review-Dexterous-Manipulation and Review-VLA-Evaluation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home