RSS 2026 PointACT - Heungwoo/research GitHub Wiki

PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 1 · paper #73 Authors: Shizhe Chen, Paul Pacaud, Cordelia Schmid arXiv: 2605.21414 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

PointACT architecture (Figure 2 of arXiv 2605.21414, © the authors)

Left: the dual-system architecture — a frozen VLM (vision encoder + text tokenizer) supplies semantic embeddings, while a trainable Point Action Expert fuses robot state, action tokens, and multi-scale point-cloud tokens via bottleneck window self-attention, FFN, downsample pooling, and cross-attention to VLM outputs. Right: detail of the bottleneck window self-attention, where action tokens are broadcast into each spatial window of point tokens and pooled back into a global action representation.

Problem

Most VLAs perceive only 2D images, which weakens fine-grained geometric reasoning and spatial grounding needed for precise 3D manipulation. Prior 3D-aware VLAs either inject depth cues as auxiliary signals on 2D features or feed only coarse last-layer point features to action experts, and pretrained 3D representations had shown limited transfer to policy learning.

Method

PointACT is a dual-system 3D-aware VLA: a frozen Qwen2.5-VL backbone plus a ~300M-parameter point-action expert. Point clouds (workspace-cropped, 1 cm voxels, up to 4,096 points) are encoded with a Point Transformer v3 (PTv3-Large: 5 stages, (3,3,3,12,3) layers, 64→768 dims) initialized from self-supervised 3D pretraining. At every hierarchical scale, evolving action tokens attend to point tokens via an efficient bottleneck window self-attention (action tokens broadcast into K disjoint spatial windows, then average-pooled), interleaved with cross-attention to the VLM embeddings. Training is behavior cloning with a regression head (L2 on 16-step action chunks) or a point-anchored classification head for keypoint poses; 2×H100 GPUs, 20K-50K steps.

Results

On LIBERO (single front view, no wrist camera, no robot-data pretraining), PointACT averages 96.0% across the four suites (Spatial 97.4, Object 99.6, Goal 96.2, Long 90.6), beating the reproduced EO1 (93.1%) under matched training. On RLBench-10Tasks it reaches 82.3% mean success versus 74% for HybridVLA, 73.2% reproduced EO1, and 64.5% ACT3D — the abstract's ~10% gain over state-of-the-art pretrained VLAs. Ablations show injecting point tokens into the VLM backbone hurts (EO1+Point drops to 18.6% on RLBench), while multi-scale point-action interaction beats last-layer point features (82.3 vs 69.7).

Significance

Evidence that where 3D enters a VLA matters: tightly coupling hierarchical point-cloud geometry into the action expert (not the VLM) yields stable gains, and pretrained 3D encoders finally pay off for policies — with a frozen VLM and modest trainable parameters. Related wiki threads: Review-LBM-Cotraining · Review-Human-Video-Transfer.

← Back to RSS 2026 survey · RSS-2026-Papers · Home