RSS 2026 StereoVLA - Heungwoo/research GitHub Wiki
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #88 Authors: Shengliang Deng, Mi Yan, Yixin Zheng, Jiayi Su, Wenhao Zhang, Xiaoguang Zhao, Heming Cui, Zhizheng Zhang, He Wang arXiv: 2512.21970 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1. Left: prior VLAs supplement single-view RGB with sensor depth, wrist-view RGB, or side-view RGB, each with drawbacks — depth ambiguity, noisy depth on transparent objects, limited/occluded wrist views, and multi-camera deployment overhead. Middle: StereoVLA takes a single stereo RGB pair and, alongside the action head, co-trains Interaction-Region Depth Estimation and Camera Parameter Estimation. Right: the model stays robust under near-hemispheric camera viewpoint variation.
Problem
VLA models excel at generalist manipulation but lack fine-grained spatial awareness and viewpoint robustness, largely because pretrained RGB encoders prioritize semantic alignment over explicit geometry. Existing remedies — wrist-mounted cameras, depth sensors, or extra third-person cameras — add hardware, suffer sensor noise on transparent/specular surfaces, or hurt viewpoint robustness. Stereo vision, which supplies geometric cues via binocular disparity with a single camera, had not been leveraged for VLAs.
Method
StereoVLA is presented as the first VLA to incorporate rich geometric cues from large-scale synthetic stereo data. A Geometric-and-Semantic (GeoSem) vision encoder fuses geometry-centric features from the FoundationStereo depth foundation model (both stereo views) with semantically rich tokens from PrismaticVLM (left view) into hybrid visual tokens. Two synergistic co-training objectives are added: Interaction-Region Depth Estimation (predicting depth in task-relevant regions rather than background) for spatial reasoning, and Camera Parameter Estimation to implicitly align camera and robot frames for viewpoint robustness. To overcome scarce real stereo data, the authors generate 5 million synthetic stereo trajectories with domain randomization; the final model is trained for 300k steps on 32 NVIDIA H800 GPUs at batch size 384.
Results
Against baselines using various input modalities, StereoVLA reports a 33.4% absolute gain in real-world success rate. Across four increasingly hard evaluation groups it scores 77.2 / 62.2 / 53.9 / 52.2%, versus π0.5 (65.6 / 51.7 / 42.2 / 39.4), GraspVLA (71.7 / 23.3 / 11.1 / 7.2), and SpatialVLA-ED (35.6 / 6.7 / 3.9 / 2.8). On a small-object (1–2 cm) grasping task where all baselines completely fail (0%), StereoVLA reaches 33%. A synthetic-pretraining ablation shows 26.6% / 13.3% without pretraining rising to 73.3% / 53.3% with it, and the model is robust to near-hemispheric camera pose variation.
Significance
Positions stereo RGB as a lightweight, single-camera route to geometry-aware manipulation that sidesteps depth-sensor failure modes and multi-camera overhead — a recurring concern in Review-VLA-Architecture and geometry-conditioned policy work.
← Back to RSS 2026 survey · RSS-2026-Papers · Home