ICML 2026 Any3D VLA - Heungwoo/research GitHub Wiki
Any3D-VLA: Enhancing VLA Robustness via Diverse Point Clouds — Domain-agnostic 3D from simulator, sensor, and model-estimated point clouds
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Xianzhe Fan, Shengliang Deng, Xiaoyang Wu, Yuxiang Lu, Zhuoling Li, Mi Yan, Yujia Zhang, Zhizheng Zhang, He Wang, Hengshuang Zhao Traction (2026-06): 6 citations (arXiv)

Problem
Most Vision-Language-Action (VLA) models take only 2D images as visual input, so their spatial understanding is inherited from 2D backbones and they are brittle in scenarios with small objects, viewpoint changes, and occlusions. The community has tried injecting 3D in several ways (depth-pretrained or spatial encoders such as VGGT, depth-map inputs, or point-cloud inputs), but these are limited by the lack of a pretrained point encoder, reliance on non-parametric 3D tokenizers, point clouds treated independently from 2D, or by requiring high-precision depth hardware.
The core question: how can 3D information be incorporated to enhance VLA capabilities while overcoming (1) scarce 3D data and (2) the cross-environment domain gap induced by depth-scale biases? A pilot study across observation spaces and visual representations finds that explicitly lifting visual input into point clouds yields representations that best complement their corresponding 2D representations.
Method
Any3D-VLA is a plug-in pipeline for existing VLA backbones. Given RGB images with optional depth, it:
- Lifts the visual input into point clouds and compresses them (3D compression).
- Runs a pre-trained point cloud encoder to produce point-wise embeddings.
- Aligns and fuses the 3D embeddings with the corresponding 2D patch features, feeding the fused 2D–3D representation into the downstream backbone (joint training of the VLM and a flow-matching action expert).
To address data scarcity and domain gap, the authors introduce a hybrid point-cloud training strategy that unifies three point-cloud sources within one training pipeline — simulator, sensor, and model-estimated point clouds — and construct a large-scale synthetic RGBD pre-training dataset covering diverse depth sources. This teaches domain-agnostic 3D representations, so at deployment the model does not need expensive depth hardware or stringent data-collection conditions and stays robust when depth is noisy or scale-biased. Diverse point-cloud inputs also serve as a form of data augmentation.

Results
Any3D-VLA is tested with two input settings — sensor-based and model-estimated point clouds — under real-world perturbations (intra-class scale/shape variation, viewpoint changes) and fine-tuned on new tasks (e.g., flower arrangement) with small real-world demonstration sets:
- Zero-shot real world: maximum overall accuracy 62.5%, outperforming the best baseline by 29.2%.
- Real-world post-training: maximum overall accuracy reaches 93.3%.
- The fused point-cloud–2D representation and hybrid training yield robust sim-to-real generalization even as depth quality and scale vary.
- Additional results are reported on the LIBERO and CALVIN standardized benchmarks to validate generalization, with analysis of the performance gap to π0.5.
Significance
Any3D-VLA provides a general, modular way to inject 3D geometry into 2D VLAs without depending on a specific depth sensor, addressing both the data-scarcity and domain-gap bottlenecks that have limited prior 3D-VLA work. By unifying simulator, sensor, and model-estimated point clouds, it learns depth-source-agnostic representations that transfer robustly from simulation to noisy real-world deployment.
Links
- arXiv: 2602.00807
- ICML 2026: https://icml.cc/virtual/2026/poster/60472
← Back to ICML-2026