ICML 2026 The Lie We Tell - Heungwoo/research GitHub Wiki
The Lie We Tell — Correcting the Euclidean Fallacy in diffusion VLA policies via SE(3) tangent-space score matching
Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Affiliations: Bing-Cheng Chuang, I-Hsuan Chu, Bor-Jiun Lin, YuanFu Yang, Min Sun, Chun-Yi Lee Traction (2026-06): 0 citations (arXiv)

Problem
Diffusion-based Vision-Language-Action (VLA) policies achieve strong manipulation results, but the authors argue they commit a fundamental geometric error they call the Euclidean Fallacy: representing SE(3) gripper poses as flat R^12 vectors and corrupting them with additive Gaussian noise. The Special Euclidean group SE(3) is a curved Riemannian manifold, not a flat vector space, so operations valid in R^n do not preserve rigid-transformation structure. Three failure modes follow: (1) manifold drift, where generated poses violate SO(3) constraints (non-orthogonal, non-unit-determinant rotations); (2) broken equivariance under coordinate-frame changes; and (3) non-geodesic trajectories that incur excessive kinematic cost. The paper frames this as a structural incompatibility rather than a criticism of prior engineering.
Method
The authors introduce the Lie Diffuser Actor (LDA), a diffusion framework that operates intrinsically on SE(3). Noise is injected through left-invariant SDEs; the network predicts scores in the tangent space (the Lie algebra se(3), as twists ξ); and samples are retracted onto the manifold via the exponential map. This construction is shown (with proofs in the appendix) to eliminate manifold drift by construction while guaranteeing coordinate-frame equivariance and geodesic (kinematic) optimality.
The neural architecture combines geometric context encoding with an iterative denoising Transformer backbone that refines a sequence of H SE(3) waypoints over diffusion steps, each pose tokenized with learned positional embeddings plus a sinusoidal time embedding, and a dedicated tangent-space prediction head. To separate the contribution of intrinsic geometry from architecture, the authors also add a Graph Attention Network (GAT) encoder over the 3D point cloud and run a factorial ablation.

Results
On CALVIN ABC→D (zero-shot transfer to held-out environment D), LDA improves average task-chain length from 3.27 to 3.51 (+7.3%) over the 3D Diffuser Actor baseline. The full Table 1 ablation (ABC→D) shows: baseline 3D Diffuser Actor (600K) = 3.27; +GAT encoder only = 3.474; +Lie diffusion only (w/o GAT) = 3.368; full LDA (300K) = 3.512, confirming both components are individually effective and complementary. On the more diverse ABCD→D setting LDA reaches 3.584 vs. the baseline's 3.288, with the Euclidean baseline showing training instability. A manifold-constraint analysis shows Euclidean diffusion violates SO(3) constraints by 7+ orders of magnitude while LDA stays on the manifold throughout reverse diffusion. On a real robot (20 trials/task), LDA matches or beats the baseline on most tasks: Move Doll 100 vs 90, Sort Blocks 75 vs 55, Stack Cups 60 vs 55 (Block-in-Box 75 vs 80).
Significance
LDA reframes a widely-overlooked representational choice in diffusion VLA policies as a correctable geometric defect, and shows that respecting the SE(3) Lie-group structure yields measurable long-horizon and real-robot gains without scaling data or model size. Because the modification is orthogonal to vision-language backbone advances, the intrinsic-geometry recipe could transfer to other trajectory-level diffusion policies.
Links
- arXiv: 2606.01847
- ICML 2026: https://icml.cc/virtual/2026/poster/64672
← Back to ICML-2026