ICLR 2026 EquAct - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 / Category: VLA Architecture — Equivariance / Trend tag: Spatial / 3D for VLA
flowchart LR
PC[Point cloud] --> UNet["SE(3)-equivariant point-cloud U-Net<br/>spherical Fourier features"]
Lang[Language] --> iFiLM["SE(3)-invariant FiLM layers"]
iFiLM --> UNet
UNet --> Act[Action]
Multi-task manipulation policies tend to break under novel 3D object poses because their networks have no built-in geometric consistency. The paper aims for theoretical generalization to unseen scene transformations rather than relying on data augmentation. EquAct is an open-loop (keyframe / next-best-pose) policy in the RVT/PerAct lineage — it predicts discrete end-effector poses from language + point cloud rather than continuous closed-loop control.
EquAct is an SE(3)-equivariant transformer over point clouds. Two ingredients: (1) an efficient SE(3)-equivariant point-cloud U-Net using spherical Fourier features for policy reasoning, and (2) SE(3)-invariant Feature-wise Linear Modulation (iFiLM) layers for language conditioning, so language influences the policy without breaking equivariance.
State-of-the-art across 18 RLBench tasks under SE(3) and SE(2) scene perturbations and across varying training-data sizes, plus four physical robot tasks. Specific per-task numbers (not stated in abstract).
Stakes out the "principled equivariance" position in 2026's spatial cluster — orthogonal to the VLM-centric papers (Spatial Forcing, FALCON, SP-VLA) that bolt spatial signals onto a 2D backbone. The iFiLM trick (language-conditioning that respects equivariance) is the technically novel piece.
- OpenReview: https://openreview.net/forum?id=d1wuA8oIH0
- arXiv: https://arxiv.org/abs/2505.21351 — "EquAct: An SE(3)-Equivariant Multi-Task Transformer for Open-Loop Robotic Manipulation" (Xupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters, Robert Platt — Northeastern University)
← Back to ICLR-2026