RSS 2026 AxisGuide - Heungwoo/research GitHub Wiki

AxisGuide: Grounding Robot Action Coordinate System in RGB Observations for Robust Visuomotor Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 3 · paper #125 Authors: Jiyun Jang, Yujin Sung, Woosung Joung, Daewon Chae, Sangwon Lee, Sohwi Kim, Jinkyu Kim, Jungbeom Lee arXiv: 2606.06761 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

AxisGuide motivation and effect (Figure 1 of arXiv 2606.06761, © the authors)

"Pick up the black bowl": conventional policies (left) receive only RGB, and success (green dots) stays confined near training placements (blue squares) while unseen locations fail (red crosses). AxisGuide (right) renders the robot base-frame x/y/z axes at the end-effector into the observation, and the success region expands across a much wider range of unseen object positions.

Problem

Behavior-cloned visuomotor policies can "understand" a scene semantically yet fail to execute correct low-level actions under distribution shift — even a simple pick-up degrades sharply at unseen object locations. The authors attribute this to insufficient action understanding: the policy cannot interpret how the robot's base-frame action coordinate system (which pixel direction is +x, +y, +z) manifests in the camera view, since the robot base is rarely visible.

Method

AxisGuide is a lightweight guidance signal: using camera intrinsics/extrinsics and the end-effector pose (no depth), it projects unit translations along the base-frame axes into image space and renders three arrows anchored at the projected gripper pixel, encoded as an R/G/B action-coordinate cue image A_t (3×H×W). The cue is channel-wise concatenated with RGB (6-channel input; only the first conv layer of the vision backbone widens), avoiding overlay occlusion. It plugs into Diffusion Policy and SmolVLA, in single- and multi-view configurations.

Results

On the spatial-generalization probe, AxisGuide lifts LIBERO Pick Up (Bowl) success at unseen locations from 52.38% to 65.71% (+13.33 pp) and real-world Pick Up (Pear) from 30.12% to 50.00% (+19.88 pp). Single-view LIBERO: DP+AxisGuide reaches 82.67/93.33/100.0% on Pick&Place/Drawer/Stove versus 69.33/86.67/92.00% for DP (also beating DP+KYC). Real-world UR5e: Flip Pot rises 73.33→93.33% (multi-view) and 36.65→50.00% (single-view). Multi-task LIBERO with SmolVLA improves across all four suites (e.g., Spatial 66.5→72.0).

Significance

A near-free input-space fix for the understanding-vs-execution gap: making action-space semantics visible in pixels improves spatial generalization across policy families, complementing camera-conditioning approaches like KYC. Related wiki threads: Review-Dexterous-Manipulation · Review-Cross-Embodiment.

← Back to RSS 2026 survey · RSS-2026-Papers · Home