RSS 2026 Universal Pose Pretraining for Generalizable - Heungwoo/research GitHub Wiki

Universal Pose Pretraining for Generalizable Vision-Language-Action Policies

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #145 Authors: Haitao Lin, Hanyang Yu, Jingshun Huang, He Zhang, Yonggen Ling, Ping Tan, Xiangyang Xue, Yanwei Fu (Tencent Robotics X · Futian Laboratory · HKUST · Fudan University · Shanghai Innovation Institute) arXiv: 2602.19710 · program page

Summary compiled from the arXiv paper (v3); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Pose-VLA decoupled pre-training vs. existing VLAs (Figure 1 of arXiv 2602.19710, © the authors)

Figure 1: Overview of Pose-VLA (paper title "Pose-VLA"). (a) Existing VLAs couple a VQA-dominated VLM directly to sparse, embodiment-specific action supervision. (b) Pose-VLA instead uses unified pose prediction as a proxy task: a pre-training phase extracts universal 3D spatial priors from diverse 3D non-robotic data via discrete pose tokens, then an alignment phase adapts these priors to a specific embodiment with only a few (100) demonstrations.

Problem

VLA models often suffer feature collapse and low training efficiency because they entangle high-level perception with sparse, embodiment-specific action supervision. Backbones optimized for Visual Question Answering excel at semantic identification but overlook the subtle 3D state variations that dictate distinct action patterns.

Method

Pose-VLA decouples training into a pre-training phase that extracts universal 3D spatial priors in a unified camera-centric space, and a post-training phase for efficient embodiment alignment in robot-specific action space. It introduces discrete pose tokens as a universal representation, letting the model integrate spatial grounding from diverse 3D datasets with geometry-level trajectories from robot demonstrations. Pre-training follows a two-stage pipeline: first establishing spatial grounding via pose prediction, then motion alignment through trajectory supervision. For fair comparison with RGB-only baselines, Pose-VLA is evaluated using only RGB input (masking depth and raymap modalities).

Results

On RoboTwin 2.0, Pose-VLA sets a new state of the art with a 79.5% average success rate, reaching 79.1% in the challenging Hard setting — a 14.0-point margin over the π0 baseline and over 45 points above the vanilla PaliGemma baseline. On LIBERO it attains 96.0% overall (surpassing π0, second only to π0.5), including 92.4% on the multi-stage LIBERO-Long suite (tying π0.5). On 3D spatial grounding (Omni3D / Objectron), it reports a 16.1% absolute improvement over the strongest baseline. Real-world experiments show robust generalization across diverse objects with only 100 demonstrations per task.

Significance

Argues that injecting 3D spatial priors via pose pre-training prevents the feature collapse of VQA-oriented backbones and improves data efficiency for generalizable manipulation — relevant to the co-training thread Review-LBM-Cotraining and to RL.

← Back to RSS 2026 survey · RSS-2026-Papers · Home