ICLR 2026 Spatial Forcing - Heungwoo/research GitHub Wiki

Spatial Forcing โ€” Implicit Spatial Representation Alignment for VLAs

Venue: ICLR 2026 / Category: VLA Architecture โ€” Spatial grounding / Trend tag: Spatial / 3D for VLA Affiliation: HKUST (Guangzhou) ยท Tsinghua University ยท Westlake University ยท Zhejiang University ยท South China University of Technology

Approach diagram

flowchart LR
  RGB[RGB obs I] --> Tok[Vision tokenizer<br/>SigLIP+DINOv2 / PaliGemma]
  Tok --> L1[VLA layer 1]
  L1 --> Lk[... layer 24<br/>middle-deep layer]
  Lk -->|aligned tokens x_V_i| Action[Action head<br/>flow matching / parallel decode]
  Lk -.->|cosine-similarity loss| MLP[BatchNorm + 2-layer MLP]
  MLP -.-> Comp[Cosine alignment]
  VGGT[Frozen VGGT 3D foundation model<br/>frame-wise + global Alternating Attention] --> Spat[Per-pixel spatial repr f_3D I]
  PE[+ positional embedding E] --> Spat
  Spat -.-> Comp
  Action --> Out[Action]
Loading

Problem

Current VLAs inherit 2D-pretrained VLM backbones (PaliGemma, Prismatic, SigLIP+DINOv2) and lack spatial awareness. Two existing fixes have flaws:

  1. Explicit 3D inputs (depth maps, point clouds, lidar โ€” SpatialVLA, GeoVLA, 3D-CAVLA, EVO-0): suffer from sensor noise, hardware heterogeneity, and the fact that large robot datasets like Open-X-Embodiment / BC-Z don't contain depth at all.
  2. 2D-image depth estimation (SpatialVLA, EVO-0 use a depth estimator): bottlenecked by estimator quality.

A lightweight depth probing experiment in the paper makes the gap concrete: freeze the visual embeddings from OpenVLA-OFT and train only a DPT head to regress depth โ€” the resulting depth maps are blurry / wrong. Visual embeddings learned from 2D-only RGB action-imitation do not encode usable spatial structure.

Detailed Method

Spatial Forcing alignment loss

Multi-view images I are fed to a frozen VGGT (Visual Geometry Grounded Transformer, Wang et al. 2025) which uses Alternating Attention (frame-wise + global) to produce per-pixel spatial representations f^{3D}(I). VGGT is chosen because it (a) is trained on 2Dโ€“3D paired data, (b) inherently handles multi-view consistency, (c) produces dense spatial features compatible with patch tokens.

To each VLA visual token x^V_i (the per-pixel patch features after VLA layer โ„“), apply BatchNorm ฮ“ then a 2-layer MLP for dim compatibility, then maximize cosine similarity to the corresponding VGGT feature plus a positional embedding E:

L_align = โˆ’ (1/N) ฮฃ_i S[ MLPยทฮ“(x^V_i), f^{3D}_i(I) + E ]

The positional embedding E is added to the target spatial features so that the supervised tokens preserve relative-position information critical to the auto-regressive VLA. The total loss is L_SF = L_action + ฮฑ ยท L_align.

Which layer to align?

The VLM backbone (Prismatic) has 32 causal-attention layers. Empirical sweep (Table 2) shows layer 24 is best โ€” relatively deep but not the deepest. Reasoning: aligning at deep layers implicitly forces shallow layers to inherit spatial structure (because gradients flow back), while constraining shallow layers directly causes spatial info to be lost in subsequent layers. Layers 28โ€“32 begin to converge into modality-agnostic space (Huang et al. 2024) so they're less amenable to vision-target supervision.

Inference is unchanged

The model trained with SF runs identically to a vanilla VLA at deployment โ€” no VGGT, no depth, no extra modules. SF is purely a training-time alignment signal. This is a major selling point versus 3D-input methods.

Training setup

  • Two base models:
    • OpenVLA-OFT on LIBERO โ€” Prismatic VLM (SigLIP + DINOv2 fused vision backbone, Open-X-Embodiment-pretrained), trained on 8ร— H100 for 150K iterations.
    • ฯ€0 on RoboTwin โ€” PaliGemma backbone, trained with LoRA on 1ร— H100 for 30K iterations.
  • All other settings follow the official OpenVLA-OFT / ฯ€0 protocols.

Comprehensive Results

LIBERO benchmark (500 trials per task; Table 1)

Per-suite success rates (%):

Method Spatial Object Goal Long Average
Diffusion Policy [RSS'23] 78.3 92.5 68.3 50.5 72.4
TraceVLA [ICLR'25] 84.6 85.2 75.1 54.1 74.8
Octo [RSS'24] 78.9 85.7 84.6 51.1 75.1
OpenVLA [CoRL'24] 84.7 88.4 79.2 53.7 76.5
Dita [ICCV'25] 84.2 96.3 85.4 63.8 82.4
CoT-VLA [CVPR'25] 87.5 91.6 87.6 69.0 83.9
ฯ€0-FAST [RSS'25] 96.4 96.8 88.6 60.2 85.5
ฯ€0 [RSS'25] 96.8 98.8 95.8 85.2 94.2
UniVLA [RSS'25] 96.5 96.8 95.6 92.0 95.2
OpenVLA-OFT [RSS'25] 97.6 98.4 97.9 94.5 97.1
(grey: 3D sensor input)
SpatialVLA 88.2 89.9 78.6 55.5 78.1
GeoVLA 98.4 99.0 96.6 96.6 97.7
3D-CAVLA 98.2 99.8 98.2 96.1 98.1
Spatial Forcing (Ours) 99.4 99.6 98.8 96.0 98.5

Beats every 2D method including OpenVLA-OFT (+1.4 pp average), and matches or beats explicit-3D methods (GeoVLA, 3D-CAVLA) without using any depth or point-cloud input.

RoboTwin 2.0 (bimanual; easy 100 trials, hard 300 trials/task; Figure 4)

Numbers reported in figure form. Key claim: SF gets the highest average SR over ฯ€0 base across both easy (in-domain) and hard (domain-randomized: clutter, textures, lighting, table heights) settings. The hard-setting gap is larger โ€” interpreted as SF teaching the model to use spatial relationships rather than spurious background/lighting cues.

Real-world bimanual AgileX (Figure 6)

6-DoF Piper ร— 2 + 1-DoF gripper, primary + dual wrist cameras. Trained on 40 demonstrations per single-arm task / 20 for bimanual. 10 trials per variation ร— 4 variations = 40 trials per single-arm task; 20 trials for dual-arm.

Task w/o SF โ†’ w/ SF (SR%)
Stack glass cups (light variation, deceptive reflections) 15 โ†’ 62.5 (+47.5)
Grasp right-side vegetable (target object variation) 10 โ†’ 47.5
Place green block (height variation) 67.5 โ†’ 85
Lift pot (dual-arm, balance) 30 โ†’ 42.5

The transparent-cups task is the most striking: SF captures underlying spatial relationships rather than overfitting to lighting reflections.

Ablation Studies (Table 2, 1ร— H100)

Target representation choice

Target Spatial Object Goal Long Avg
(no SF baseline) 96.8 94.8 92.8 86.2 92.7
SigLIP 95.2 94.8 94.0 91.8 94.0
DINOv2 93.4 95.2 93.8 93.8 94.1
VGGT w/o positional embedding 97.8 100.0 96.6 84.4 94.7
VGGT w/ PE 97.2 99.2 96.8 94.2 96.9

VGGT > SigLIP / DINOv2 confirms that 3D-aware target features matter, and the PE component specifically rescues long-horizon performance (84.4 โ†’ 94.2 on LIBERO-Long).

Aligned layer (out of 32)

Layer Spatial Object Goal Long Avg
1 96.8 99.4 99.0 83.0 94.6
8 96.2 98.4 95.6 92.4 95.7
16 97.4 98.8 95.8 83.2 93.8
24 97.2 99.2 96.8 94.2 96.9
32 (deepest) 98.8 99.4 96.2 84.8 94.8

Best at 24 โ€” middle-deep, not deepest.

Training-iteration efficiency (Figure 5(a))

SF reaches the same SR as baseline OpenVLA-OFT 3.8ร— faster (e.g. ~20K iters with SF matches ~150K iters without).

Iterations SR (no SF) SR (with SF)
2K โ€” 72.7
5K โ€” 87.5
20K โ€” 93.7
50K โ€” 96.5
150K 92.7 96.9

Data-efficiency (Figure 5(b))

Training data SF Spatial SF Object SF Goal SF Long SF Avg
1% 32.8 67.8 44.8 23.6 42.3
5% 73.2 83.4 80.6 66.0 75.8
100% 97.2 99.2 96.8 94.2 96.9

SF reaches 75.8% with only 5% of LIBERO data and is 5.9ร— more data-efficient at matched success rate.

t-SNE (Figure 5(c))

After alignment, the VLA's visual embeddings show same distribution shape as VGGT while keeping their cluster center distinct โ€” i.e. it absorbs spatial structure without collapsing onto the target.

Limitations (as discussed by authors)

The paper itself does not enumerate a formal limitations section. Practical caveats inferable from the text:

  1. Compute caveat: the training-efficiency / data-efficiency / layer / target ablations are run on 1ร— H100 rather than 8ร— H100, so reported ablation numbers are slightly lower than the headline 8-GPU 150K numbers. (Stated openly in Sec. 3.3: "because of limitations of computational resources.")
  2. VGGT dependence. The whole pipeline rests on VGGT being a good 3D foundation model. If VGGT mis-estimates 3D for unusual viewpoints / domains, the alignment signal degrades.
  3. Layer choice is empirical. No theoretical justification of why layer 24 of 32 is optimal โ€” the authors propose an information-theoretic intuition (deep features are modality-agnostic, alignment forces shallow layers to absorb spatial structure) but it's not formalized.
  4. No formal limitations of the VGGT-vs-PE-vs-no-PE long-horizon collapse โ€” the without-PE config gets 84.4 on Long but 100 on Object, suggesting the alignment may "over-correct" for some task types. Authors note PE specifically restores Long performance.

Significance & Positioning

  • vs SpatialVLA, GeoVLA, 3D-CAVLA, EVO-0 (explicit 3D input): SF matches or beats all of them on LIBERO without sensor depth โ€” and crucially, can be applied to large datasets that don't have depth (Open-X-Embodiment, DROID, BC-Z).
  • vs SpatialVLA, EVO-0 (depth-from-2D estimation): these depend on a depth estimator's quality; SF sidesteps that bottleneck because the alignment target VGGT is itself a strong 3D foundation model with multi-view consistency.
  • vs ฯ€0 / OpenVLA-OFT (2D baselines): SF is a drop-in training-time auxiliary loss. No inference cost. +1.4 pp on LIBERO over OpenVLA-OFT, and dramatic data/iteration efficiency wins.
  • vs Spatially Guided / FALCON: SF takes the lightest stance โ€” keep the 2D backbone, just add a representation-alignment loss. FALCON adds explicit spatial-to-action geometry priors; Spatially Guided modifies inputs. SF stays purely in latent space.
  • vs PA3FF: PA3FF aligns to a different geometric target; SF specifically argues for VGGT due to its multi-view consistency and joint frame-wise + global attention.
  • Aligns with the broader REPA / 3DRS / Genhancer trend (Yu et al., Huang et al.) of using representation supervision to bake structure into generative or VLA stacks at training time.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ