RSS 2026 Supervised Mixture of Experts for Surgical Grasping - Heungwoo/research GitHub Wiki

Supervised Mixture-of-Experts for Surgical Grasping and Retraction

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #2 Authors: Lorenzo Mazza, Ariel Rodriguez arXiv: 2601.21971 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Collaborative bowel grasping and retraction roll-outs (Figure 1 of arXiv 2601.21971, © the authors)

Three columns of policy roll-outs: the human-robot collaborative task on the OpenHELP phantom (left), zero-shot generalization to ex vivo porcine bowel (middle), and a preliminary in vivo porcine roll-out (right). In each, the surgeon-controlled gripper (bottom instrument) points to the target bowel segment (red rhomboid), then the robot assistant gripper (top instrument) grasps, waits, and retracts while maintaining tension.

Problem

Imitation learning for surgical robotics is limited by data scarcity, single-endoscope visual constraints, deformable-tissue dynamics, and strict safety/latency requirements that preclude large VLA models. Prior surgical policies (e.g., SRT-H) needed multi-camera setups and roughly 16,000 demonstrations; the authors target collaborative bowel grasping and retraction — where a robot assistant reads visual cues from the surgeon's instrument — from under 150 demonstrations using stereo endoscopic images only.

Method

A supervised Mixture-of-Experts (MoE) block is added on top of the Action Chunking Transformer (ACT): the task is segmented into H = 5 phases, with H parallel action phase-experts, H gripper phase-experts (Bernoulli heads), and a gating network trained as a phase classifier with explicit cross-entropy supervision from automatically derived phase labels; final actions are phase-weighted mixtures. The policy is vision-only (stereo image pair, no proprioception), predicting chunks of delta Cartesian tip motions plus binary gripper actions. Setup uses two UR5e arms (one holding a Karl Storz stereo endoscope) on the OpenHELP phantom; 120 fixed-viewpoint episodes plus 50 random-viewpoint episodes; training took 3 hours for ACT/ACT+MoE vs 8 h for π0.5 and 14 h for SmolVLA.

Results

In-distribution (20 phantom roll-outs): ACT+MoE reaches 17/20 end-to-end (85%) vs 10/20 for ACT (50%) — a 70% relative gain — while π0.5 and SmolVLA both score 0/20 end-to-end; ACT+MoE runs at 27 Hz vs 10 Hz (π0.5) and 3.3 Hz (SmolVLA). Out-of-distribution (novel grasp locations, reduced illumination, partial occlusions), ACT+MoE achieves 13/20 end-to-end vs 6/20 for ACT. Zero-shot on ex vivo porcine bowel it succeeds 12/15 (80%), and after retraining with the random-viewpoint data it hits 18/22 (82%) on unseen camera viewpoints; qualitative in vivo porcine roll-outs are also shown.

Significance

Evidence that in extreme low-data, safety-critical regimes, a lightweight specialist with explicit phase supervision beats generalist VLAs outright — a useful data point for Review-Dexterous-Manipulation and for debates on when VLA scale helps (Review-VLA-Evaluation).

← Back to RSS 2026 survey · RSS-2026-Papers · Home