ICLR 2026 Spatial to Actions - Heungwoo/research GitHub Wiki

FALCON — Grounding VLAs in Spatial Foundation Priors

Venue: ICLR 2026 Category: VLA Architecture — Spatial grounding Trend tag: Spatial / 3D for VLA Affiliations: ByteDance Seed, NUS, NTU, Tsinghua, SMU

Approach diagram

flowchart LR
  RGB[Multi-view RGB] --> VLM[Kosmos-2 VLM<br/>~1.6B params]
  RGB --> ESM[Embodied Spatial Model<br/>~1.0B, VGGT-style]
  Depth[Optional depth + valid mask] -.-> ESM
  Pose[Optional camera pose 7D] -.-> ESM
  VLM --> AT[Semantic action token]
  ESM --> ST[Spatial tokens M×Ds]
  ST --> Pool[Max-pool] --> MLP[Adapter MLP]
  AT --> Add[Element-wise add]
  MLP --> Add
  Add --> Head[Spatial-Enhanced Action Head<br/>MLP or LSTM]
  Head --> Act["Action chunk a_t..a_{t+C-1}"]
Loading

Problem

Most VLAs (OpenVLA, π0, GR00T, CogACT) act in 3D but read the world through 2D encoders, leaving a spatial reasoning gap that limits generalization to new heights, scales, clutter, and viewpoints. Two existing remedies both have failure modes:

  1. Explicit 3D inputs (point clouds, depth) — PointVLA, GeoVLA, 3D-VLA — require special sensors and break when those inputs are missing or differ from training (poor modality transferability).
  2. Weak 3D cues injected into the VLM — pseudo-depth, learnable spatial embeddings (SpatialVLA, Evo-0) — disrupt the VLM's pretrained vision-language alignment and require costly re-tuning to recover language reasoning.

FALCON's bet: inject rich 3D spatial tokens, but route them to the action head rather than the VLM, preserving language alignment while still exposing geometry to the controller.

Detailed Method

Architecture (~2.9B parameters total)

  • VLM backbone: Kosmos-2 (~1.6B params).
  • Embodied Spatial Model (ESM): ~1.0B params, built on VGGT (Wang et al., 2025a). Tokenizes images via DINO into visual tokens T_vis, concatenates a learnable camera token t_cam, then runs N alternating cross- and self-attention blocks (T_spl, t̂_cam) = E_spl(T_vis, t_cam). Outputs M×Ds spatial tokens plus refined camera tokens used by depth/pose heads (these are auxiliary supervision targets, not action targets).
  • Spatial-Enhanced Action Head: max-pools spatial tokens → adapter MLP D → element-wise added to the VLM's semantic action token t̂_act. The fused vector goes to either an MLP or LSTM action predictor that emits a chunk of C 7-DoF actions (Δxyz + Euler + gripper).

Optional 3D conditions

  • Camera pose P ∈ R^7 (intrinsics + normalized extrinsics) → MLP encoder produces ground-truth camera token t_gt-cam.
  • Depth map + validity mask → normalized, concatenated with mask, passed through a 14×14 conv stack to produce depth tokens T_dpt aligned with image tokens.
  • Stochastic injection during training: (T_spl, t̂_cam) = E_spl(T_vis + b_d·T_dpt, b_p·t_gt-cam + (1−b_p)·t_cam) where b_d, b_p ∼ Bernoulli(p=0.66). This gives FALCON modality transferability — same model handles RGB-only, RGB-D, and RGB-D+pose without retraining.

Training objective (per-chunk)

L = Σ_{i=t}^{t+C-1} MSE(â_i,pose, a_i,pose) + λ · BCE(â_i,gripper, a_i,gripper) ESM auxiliary losses follow VGGT (depth, point map, pose).

Two-stage post-training

  • Stage 1 (alignment): freeze VLM, action predictor, ESM; train only the lightweight adapter D, with zero-initialized final linear so spatial tokens contribute negligibly at the start. LR 1e-4, batch 128, no warmup.
  • Stage 2 (joint refinement): unfreeze VLM and adapter; ESM and action predictor stay frozen. The VLM implicitly learns to incorporate spatial cues into its semantic features without disrupting its pretraining.

ESM training

Performed before VLA post-training, on 16× A100, ~2 days. Random 1–12 frames per scene, 24 images per batch, AdamW with split LRs (1e-6 backbone, 1e-5 heads), Bernoulli p=0.66.

Hardware

Main FALCON training on 32× A100. Pre-training of the 2D Kosmos-VLA-2D uses LR 2e-5, batch 128, 0.25-epoch warmup on CALVIN / 2.5k-step warmup on OXE. Real-world Stage 1 uses batch 512.

Comprehensive Results

CALVIN (long-horizon, 5 tasks chained, Avg. Len. = expected number of consecutive successes)

Method Setting 1 2 3 4 5 Avg. Len.
GR-1 ABCD→D 94.9 89.6 84.4 78.9 73.1 4.21
UP-VLA ABCD→D 96.2 92.1 87.9 84.2 81.2 4.42
RoboVLM ABCD→D 96.7 93.0 89.9 86.5 82.6 4.49
FALCON ABCD→D 97.2 93.3 90.3 88.0 84.0 4.53
3D Diffuser Actor ABC→D (zero-shot) 93.8 80.3 66.2 53.3 41.2 3.35
GR-1 ABC→D 85.4 71.2 59.6 49.7 40.1 3.06
UP-VLA ABC→D 92.8 86.5 81.5 76.9 69.9 4.08
RoboVLM ABC→D 98.0 93.6 85.4 77.8 70.4 4.25
Seer-Large ABC→D 96.3 91.6 86.1 80.3 74.0 4.28
FALCON ABC→D 98.4 94.5 88.6 82.5 75.5 4.40

Notable: FALCON beats 3D Diffuser Actor (which uses ground-truth point clouds) by +1.05 Avg. Len. in zero-shot ABC→D, and beats 3DDP by +4.13 — strong evidence that implicit spatial-token injection beats explicit point-cloud input.

SimplerEnv WidowX (Bridge)

Method Spoon→Towel Carrot→Plate Stack Block Eggplant→Basket Avg
RT-1-X 0.0 4.2 0.0 0.0 1.1
OpenVLA 0.0 0.0 0.0 4.1 1.0
Octo-Base 12.5 8.3 0.0 43.1 16.0
RoboVLM 45.8 20.8 4.2 79.2 37.5
SpatialVLA 16.7 25.0 29.2 100.0 42.7
FALCON 62.5 41.7 20.8 100.0 56.3

SimplerEnv Google Robot

Method Pick Coke Move Near Open/Close Drawer-Apple Avg
OpenVLA 16.3 46.2 35.6 0.0 24.5
TraceVLA 28.0 53.7 57.0 0.0 34.7
RT-2-X (55B!) 78.7 77.9 25.0 3.7 46.3
RoboVLM 77.3 61.7 43.5 24.1 51.7
SpatialVLA 86.0 77.9 57.4 0.0 55.3
FALCON 90.7 79.2 39.8 41.7 62.9

The Drawer-Apple column is striking: most baselines collapse to 0%; FALCON reaches 41.7%, with even the 55B closed-source RT-2-X at only 3.7%. This is the spatial-reasoning task hardest for 2D-only methods.

Real-world (11 tasks: 9 base + 4 spatial-understanding + few-shot)

  • Base Tasks (9 task suites with cluttered scenes): FALCON 70.0% mean S.R., +25.6 over SpatialVLA's 44.4%.
  • Few-shot: FALCON beats second-best by 27.5% in Simple setting and 27% in Unseen-Average. On the "open drawer + place bread" Unseen-Object split, FALCON 80% vs near-zero for the rest.
  • Spatial-understanding tasks: RoboVLM struggles with object-scale variations (collisions on big blocks, premature releases on small ones); FALCON robust across scales.

Ablation Studies

Spatial-token injection point (CALVIN ABC→D Avg. Len.)

Variant Where spatial tokens go Avg. Len. (ABCD→D) Avg. Len. (ABC→D)
FALCON-VLM-tokens Concatenated into VLM input 4.00 3.79
Cross-Attention VLM acts as query, spatial tokens key/value 3.98 3.68
FiLM-Gated Spatial tokens generate γ, β for action token 4.04 3.76
FALCON (Element-wise add) Action head only 4.08 3.91

Two takeaways: (1) injecting spatial tokens into the VLM input disrupts pretrained alignment — Avg. Len. drops 3.91 → 3.79 in zero-shot. (2) Among action-head fusion methods, parameter-free element-wise addition wins over cross-attention and FiLM.

Modality transferability (CALVIN)

Method Train modality Test modality Avg. Len. (ABCD→D) Avg. Len. (ABC→D)
Kosmos-VLA (no ESM) RGB RGB 4.01 3.48
Kosmos-VLA (no ESM) RGB-D RGB-D 4.05 3.98
FALCON RGB RGB 4.08 3.91
FALCON RGB RGB+D at test 4.09 3.95
FALCON RGB-D RGB-D 4.09 3.97
FALCON RGB-D RGB at test 4.07 3.95

FALCON trained RGB-only matches Kosmos-VLA trained RGB-D, and adding test-time depth gives small further gains. The ESM gracefully degrades when 3D inputs disappear at test time — the central modality-transferability claim. Real-world: depth + pose at test bumps the cup-height-change task 60% → 80%.

ESM zero-shot depth estimation (CALVIN, validating the ESM is doing real geometry)

Method Depth in Camera in δ<1.25 ↑ Abs. Rel ↓
VGGT – – 91.33 8.53
ESM ✗ ✗ 90.91 8.61
ESM ✓ ✗ 99.79 0.91
ESM ✓ ✓ 99.47 0.87

ESM matches VGGT with image-only input and improves dramatically when 3D modalities are available — confirming the spatial tokens contain real geometric information rather than learned correlations.

Limitations (as stated)

The paper does not include a dedicated "Limitations" section. Implicit limitations from the design:

  • ESM training requires 16 A100 × 2 days plus VGGT-style multi-task supervision (depth, point map, pose) — not trivially reproducible.
  • Spatial-token injection is post-training by design (Sec. A.3): the authors deliberately avoid pre-training with spatial tokens to keep optimization tractable, but this also means spatial competence is bolted onto an existing VLA rather than learned end-to-end.
  • All real-world tasks are tabletop manipulation; the framework's generalization to mobile or whole-body settings is not evaluated.

Significance & Positioning

FALCON is one of three ICLR 2026 papers triangulating "where does the 3D signal enter a VLA?":

  • Inside the VLM via pretraining: SP-VLA (point/box/trajectory grounding pretraining of the VLM).
  • As an alignment loss against a frozen 3D encoder: Spatial Forcing.
  • Into the action head as tokens, with the VLM untouched: FALCON.

Empirically FALCON's design choice is well supported — the FALCON-VLM-tokens ablation (CALVIN ABC→D 3.79 Avg. Len.) shows VLM-side injection actively hurts zero-shot generalization, while head-side injection gives 3.91. Versus point-cloud-input baselines (PointVLA, 3D Diffuser Actor) FALCON wins on CALVIN ABC→D despite using only RGB at training time, with optional depth/pose only as a test-time bonus.

The Embodied Spatial Model is also a notable contribution beyond the VLA paradigm — it generalizes VGGT to ingest optional depth/pose without architectural change, opening a path for any 2D foundation model to acquire 3D priors. The Bernoulli-stochastic conditioning trick is the practical machinery enabling that.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️