ICML 2026 GeoMoLa - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Yunchao Zhang, Yijia Weng, Ruizhe Liu, Ming Hu, Leonidas Guibas, Yanchao Yang
Learning motion latents for robotic manipulation typically relies on extracting motion patterns from visual sequences. But effective action abstractions require understanding three-dimensional geometric transformations, not just appearance. When latents are learned by reconstructing pixels, they tend to encode visual appearance rather than the physical motion that actually drives manipulation — and existing strong methods further depend on multi-view reconstruction.
GeoMoLa (Geometry-Aware Motion Latents) learns discrete motion latent codes by predicting how point clouds evolve during manipulation, rather than reconstructing visual observations. This is a four-dimensional objective — spatial geometry changing through time — which forces the latent representations to encode actual physical motion instead of appearance patterns. The method operates from single-view RGB-D input only, whereas prior approaches require multi-view reconstruction. The learned codes act as transferable motion abstractions: applying them to novel scenes produces physically consistent transformations regardless of visual context.
flowchart LR
A[Single-view RGB-D] --> B[Encode discrete motion latents]
B --> C[Predict point-cloud evolution<br/>4D: geometry x time]
C --> D[Latents encode physical motion,<br/>not appearance]
D --> E[Transfer codes to novel scenes -><br/>physically consistent transforms]
E --> F[Manipulation policy]
Per the abstract, GeoMoLa achieves state-of-the-art performance using only single-view RGB-D input, succeeding across diverse manipulation benchmarks while existing methods require multi-view reconstruction. Ablations reveal that geometric prediction is the key driver of performance, quantitatively validating that manipulation depends on spatial understanding. The learned codes exhibit effective motion abstraction, transferring to novel scenes as physically consistent transformations. Real-world experiments confirm robustness, achieving robust manipulation with minimal demonstrations in cluttered environments where geometric reasoning determines success.
(No arXiv preprint was available at the time of writing; quantitative tables are not yet public. Results above are drawn from the ICML abstract.)
GeoMoLa argues that effective motion latents for robot control emerge better from understanding motion through its three-dimensional effects than from pixel-level patterns. By swapping a visual-reconstruction objective for a 4D point-cloud-evolution objective, it obtains transferable, geometry-grounded action abstractions from a single RGB-D view — lowering the sensing burden (no multi-view rig) while improving robustness in cluttered, low-demonstration settings.
- ICML 2026: https://icml.cc/virtual/2026/poster/62176
← Back to ICML-2026