RSS 2026 LAP - Heungwoo/research GitHub Wiki
LAP: Language-Action Pre-training Enables Zero-Shot Cross-Embodiment Transfer
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #203 Authors: Lihan Zha, Asher James Hancock, Mingtong Zhang, Tenny Yin, Yixuan Huang, Dhruv Shah, Allen Z. Ren, Anirudha Majumdar (Princeton University · Physical Intelligence) arXiv: 2602.10556 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: Overview of LAP. Low-level end-effector actions (e.g. "Move forward 3cm", "tilt forward 30 degrees") are expressed directly as natural-language "language-actions" and used to supervise a PaliGemma-3B/Gemma3 backbone, unifying robot action prediction with motion-prediction VQA in a "Semantically Unified Action Space." The right panels show zero-shot generalization to unseen embodiments (SOTA VLA 0% vs. LAP-3B 52%) and transferable-representation scaling curves.
Problem
Despite large-scale multi-embodiment pre-training, existing VLAs stay tightly coupled to their training embodiments and rarely work zero-shot on new robots — even minor differences (an altered gripper or wrist-camera location) usually demand costly per-embodiment fine-tuning.
Method
Language-Action Pre-training (LAP) represents low-level robot actions directly in natural language ("language-actions"), aligning action supervision with the pre-trained VLM's input–output distribution. It requires no learned tokenizer, no costly annotation, and no embodiment-specific architecture; language-actions are parsed from raw actions under a fixed coordinate convention. The instantiation, LAP-3B, initializes its VLM backbone from PaliGemma-3B and uses a Mixture-of-Transformers architecture combining the LAP-trained VLM with a lightweight diffusion-based action expert for real-time control (following π0.5, from which it differs only in action representation). It is trained on open-sourced robot datasets including Open X-Embodiment, and can additionally co-train with VQA data.
Results
Across multiple novel robot embodiments and manipulation tasks, LAP-3B attains over 50% average zero-shot success (52% reported) — roughly a 2× improvement (~30 points absolute) over the strongest prior VLAs — while all open-sourced VLA baselines collapse to near-zero zero-shot success. It consistently outperforms replicated π0.5 and π0 baselines by about 15 percentage points, and the paper reports efficient adaptation, favorable scaling, and additional gains from unifying action prediction with VQA through co-training.
Significance
Presented as the first VLA to achieve substantial zero-shot transfer to unseen real embodiments without embodiment-specific fine-tuning, advancing cross-embodiment generalist policies discussed in Review-Human-Video-Transfer and Review-Realtime-Execution.
← Back to RSS 2026 survey · RSS-2026-Papers · Home