ICLR 2026 Spatial to Actions - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Spatial grounding Trend tag: Spatial / 3D for VLA Affiliations: ByteDance Seed, NUS, NTU, Tsinghua, SMU
flowchart LR
RGB[Multi-view RGB] --> VLM[Kosmos-2 VLM<br/>~1.6B params]
RGB --> ESM[Embodied Spatial Model<br/>~1.0B, VGGT-style]
Depth[Optional depth + valid mask] -.-> ESM
Pose[Optional camera pose 7D] -.-> ESM
VLM --> AT[Semantic action token]
ESM --> ST[Spatial tokens M×Ds]
ST --> Pool[Max-pool] --> MLP[Adapter MLP]
AT --> Add[Element-wise add]
MLP --> Add
Add --> Head[Spatial-Enhanced Action Head<br/>MLP or LSTM]
Head --> Act["Action chunk a_t..a_{t+C-1}"]
Most VLAs (OpenVLA, π0, GR00T, CogACT) act in 3D but read the world through 2D encoders, leaving a spatial reasoning gap that limits generalization to new heights, scales, clutter, and viewpoints. Two existing remedies both have failure modes:
- Explicit 3D inputs (point clouds, depth) — PointVLA, GeoVLA, 3D-VLA — require special sensors and break when those inputs are missing or differ from training (poor modality transferability).
- Weak 3D cues injected into the VLM — pseudo-depth, learnable spatial embeddings (SpatialVLA, Evo-0) — disrupt the VLM's pretrained vision-language alignment and require costly re-tuning to recover language reasoning.
FALCON's bet: inject rich 3D spatial tokens, but route them to the action head rather than the VLM, preserving language alignment while still exposing geometry to the controller.
- VLM backbone: Kosmos-2 (~1.6B params).
-
Embodied Spatial Model (ESM): ~1.0B params, built on VGGT (Wang et al., 2025a). Tokenizes images via DINO into visual tokens T_vis, concatenates a learnable camera token t_cam, then runs N alternating cross- and self-attention blocks
(T_spl, t̂_cam) = E_spl(T_vis, t_cam). Outputs M×Ds spatial tokens plus refined camera tokens used by depth/pose heads (these are auxiliary supervision targets, not action targets). - Spatial-Enhanced Action Head: max-pools spatial tokens → adapter MLP D → element-wise added to the VLM's semantic action token t̂_act. The fused vector goes to either an MLP or LSTM action predictor that emits a chunk of C 7-DoF actions (Δxyz + Euler + gripper).
- Camera pose P ∈ R^7 (intrinsics + normalized extrinsics) → MLP encoder produces ground-truth camera token t_gt-cam.
- Depth map + validity mask → normalized, concatenated with mask, passed through a 14×14 conv stack to produce depth tokens T_dpt aligned with image tokens.
-
Stochastic injection during training:
(T_spl, t̂_cam) = E_spl(T_vis + b_d·T_dpt, b_p·t_gt-cam + (1−b_p)·t_cam)whereb_d, b_p ∼ Bernoulli(p=0.66). This gives FALCON modality transferability — same model handles RGB-only, RGB-D, and RGB-D+pose without retraining.
L = Σ_{i=t}^{t+C-1} MSE(â_i,pose, a_i,pose) + λ · BCE(â_i,gripper, a_i,gripper)
ESM auxiliary losses follow VGGT (depth, point map, pose).
- Stage 1 (alignment): freeze VLM, action predictor, ESM; train only the lightweight adapter D, with zero-initialized final linear so spatial tokens contribute negligibly at the start. LR 1e-4, batch 128, no warmup.
- Stage 2 (joint refinement): unfreeze VLM and adapter; ESM and action predictor stay frozen. The VLM implicitly learns to incorporate spatial cues into its semantic features without disrupting its pretraining.
Performed before VLA post-training, on 16× A100, ~2 days. Random 1–12 frames per scene, 24 images per batch, AdamW with split LRs (1e-6 backbone, 1e-5 heads), Bernoulli p=0.66.
Main FALCON training on 32× A100. Pre-training of the 2D Kosmos-VLA-2D uses LR 2e-5, batch 128, 0.25-epoch warmup on CALVIN / 2.5k-step warmup on OXE. Real-world Stage 1 uses batch 512.
| Method | Setting | 1 | 2 | 3 | 4 | 5 | Avg. Len. |
|---|---|---|---|---|---|---|---|
| GR-1 | ABCD→D | 94.9 | 89.6 | 84.4 | 78.9 | 73.1 | 4.21 |
| UP-VLA | ABCD→D | 96.2 | 92.1 | 87.9 | 84.2 | 81.2 | 4.42 |
| RoboVLM | ABCD→D | 96.7 | 93.0 | 89.9 | 86.5 | 82.6 | 4.49 |
| FALCON | ABCD→D | 97.2 | 93.3 | 90.3 | 88.0 | 84.0 | 4.53 |
| 3D Diffuser Actor | ABC→D (zero-shot) | 93.8 | 80.3 | 66.2 | 53.3 | 41.2 | 3.35 |
| GR-1 | ABC→D | 85.4 | 71.2 | 59.6 | 49.7 | 40.1 | 3.06 |
| UP-VLA | ABC→D | 92.8 | 86.5 | 81.5 | 76.9 | 69.9 | 4.08 |
| RoboVLM | ABC→D | 98.0 | 93.6 | 85.4 | 77.8 | 70.4 | 4.25 |
| Seer-Large | ABC→D | 96.3 | 91.6 | 86.1 | 80.3 | 74.0 | 4.28 |
| FALCON | ABC→D | 98.4 | 94.5 | 88.6 | 82.5 | 75.5 | 4.40 |
Notable: FALCON beats 3D Diffuser Actor (which uses ground-truth point clouds) by +1.05 Avg. Len. in zero-shot ABC→D, and beats 3DDP by +4.13 — strong evidence that implicit spatial-token injection beats explicit point-cloud input.
| Method | Spoon→Towel | Carrot→Plate | Stack Block | Eggplant→Basket | Avg |
|---|---|---|---|---|---|
| RT-1-X | 0.0 | 4.2 | 0.0 | 0.0 | 1.1 |
| OpenVLA | 0.0 | 0.0 | 0.0 | 4.1 | 1.0 |
| Octo-Base | 12.5 | 8.3 | 0.0 | 43.1 | 16.0 |
| RoboVLM | 45.8 | 20.8 | 4.2 | 79.2 | 37.5 |
| SpatialVLA | 16.7 | 25.0 | 29.2 | 100.0 | 42.7 |
| FALCON | 62.5 | 41.7 | 20.8 | 100.0 | 56.3 |
| Method | Pick Coke | Move Near | Open/Close | Drawer-Apple | Avg |
|---|---|---|---|---|---|
| OpenVLA | 16.3 | 46.2 | 35.6 | 0.0 | 24.5 |
| TraceVLA | 28.0 | 53.7 | 57.0 | 0.0 | 34.7 |
| RT-2-X (55B!) | 78.7 | 77.9 | 25.0 | 3.7 | 46.3 |
| RoboVLM | 77.3 | 61.7 | 43.5 | 24.1 | 51.7 |
| SpatialVLA | 86.0 | 77.9 | 57.4 | 0.0 | 55.3 |
| FALCON | 90.7 | 79.2 | 39.8 | 41.7 | 62.9 |
The Drawer-Apple column is striking: most baselines collapse to 0%; FALCON reaches 41.7%, with even the 55B closed-source RT-2-X at only 3.7%. This is the spatial-reasoning task hardest for 2D-only methods.
- Base Tasks (9 task suites with cluttered scenes): FALCON 70.0% mean S.R., +25.6 over SpatialVLA's 44.4%.
- Few-shot: FALCON beats second-best by 27.5% in Simple setting and 27% in Unseen-Average. On the "open drawer + place bread" Unseen-Object split, FALCON 80% vs near-zero for the rest.
- Spatial-understanding tasks: RoboVLM struggles with object-scale variations (collisions on big blocks, premature releases on small ones); FALCON robust across scales.
| Variant | Where spatial tokens go | Avg. Len. (ABCD→D) | Avg. Len. (ABC→D) |
|---|---|---|---|
| FALCON-VLM-tokens | Concatenated into VLM input | 4.00 | 3.79 |
| Cross-Attention | VLM acts as query, spatial tokens key/value | 3.98 | 3.68 |
| FiLM-Gated | Spatial tokens generate γ, β for action token | 4.04 | 3.76 |
| FALCON (Element-wise add) | Action head only | 4.08 | 3.91 |
Two takeaways: (1) injecting spatial tokens into the VLM input disrupts pretrained alignment — Avg. Len. drops 3.91 → 3.79 in zero-shot. (2) Among action-head fusion methods, parameter-free element-wise addition wins over cross-attention and FiLM.
| Method | Train modality | Test modality | Avg. Len. (ABCD→D) | Avg. Len. (ABC→D) |
|---|---|---|---|---|
| Kosmos-VLA (no ESM) | RGB | RGB | 4.01 | 3.48 |
| Kosmos-VLA (no ESM) | RGB-D | RGB-D | 4.05 | 3.98 |
| FALCON | RGB | RGB | 4.08 | 3.91 |
| FALCON | RGB | RGB+D at test | 4.09 | 3.95 |
| FALCON | RGB-D | RGB-D | 4.09 | 3.97 |
| FALCON | RGB-D | RGB at test | 4.07 | 3.95 |
FALCON trained RGB-only matches Kosmos-VLA trained RGB-D, and adding test-time depth gives small further gains. The ESM gracefully degrades when 3D inputs disappear at test time — the central modality-transferability claim. Real-world: depth + pose at test bumps the cup-height-change task 60% → 80%.
| Method | Depth in | Camera in | δ<1.25 ↑ | Abs. Rel ↓ |
|---|---|---|---|---|
| VGGT | – | – | 91.33 | 8.53 |
| ESM | ✗ | ✗ | 90.91 | 8.61 |
| ESM | ✓ | ✗ | 99.79 | 0.91 |
| ESM | ✓ | ✓ | 99.47 | 0.87 |
ESM matches VGGT with image-only input and improves dramatically when 3D modalities are available — confirming the spatial tokens contain real geometric information rather than learned correlations.
The paper does not include a dedicated "Limitations" section. Implicit limitations from the design:
- ESM training requires 16 A100 × 2 days plus VGGT-style multi-task supervision (depth, point map, pose) — not trivially reproducible.
- Spatial-token injection is post-training by design (Sec. A.3): the authors deliberately avoid pre-training with spatial tokens to keep optimization tractable, but this also means spatial competence is bolted onto an existing VLA rather than learned end-to-end.
- All real-world tasks are tabletop manipulation; the framework's generalization to mobile or whole-body settings is not evaluated.
FALCON is one of three ICLR 2026 papers triangulating "where does the 3D signal enter a VLA?":
- Inside the VLM via pretraining: SP-VLA (point/box/trajectory grounding pretraining of the VLM).
- As an alignment loss against a frozen 3D encoder: Spatial Forcing.
- Into the action head as tokens, with the VLM untouched: FALCON.
Empirically FALCON's design choice is well supported — the FALCON-VLM-tokens ablation (CALVIN ABC→D 3.79 Avg. Len.) shows VLM-side injection actively hurts zero-shot generalization, while head-side injection gives 3.91. Versus point-cloud-input baselines (PointVLA, 3D Diffuser Actor) FALCON wins on CALVIN ABC→D despite using only RGB at training time, with optional depth/pose only as a test-time bonus.
The Embodied Spatial Model is also a notable contribution beyond the VLA paradigm — it generalizes VGGT to ingest optional depth/pose without architectural change, opening a path for any 2D foundation model to acquire 3D priors. The Bernoulli-stochastic conditioning trick is the practical machinery enabling that.
- OpenReview: https://openreview.net/forum?id=fzmittHfq3
- PDF: https://openreview.net/pdf?id=fzmittHfq3
- Project page: https://falcon-vla.github.io/
- Spatial Forcing — alignment-loss approach
- SP-VLA — VLM-side pretraining approach
- PA3FF — natively 3D part-aware features
- GR00T N1.5
- π0.6
← Back to ICLR-2026