ICLR 2026 Spatially Guided - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Training β Spatial pretraining + dual-system action expert Trend tag: Spatial / 3D for VLA Affiliations: Shanghai AI Lab, HKUST, SUSTech, Fudan Internal name: ST4VLA (released as part of the InternVLA-M1 framework)
flowchart LR
subgraph S1[Stage 1: Spatial Grounding Pre-training]
Web[LLaVA-OneVision + RefCOCO + COCO-ReM] --> VLM
Robot[RoboRefIt + A0 + ST4VLA Manip Data] --> VLM
VLM[Qwen2.5-VL-3B-Instruct<br/>Box / Point / Trajectory QA<br/>~3M data, 2.3M spatial]
end
subgraph S2[Stage 2: Spatially Guided Action Post-training]
VLM --> Plan[VLM Planner / System 2]
Prompt["Spatial prompt:<br/>'Figure out how to execute it,<br/>then locate the key object'"]
Prompt --> Plan
Plan --> QFmr[Querying Transformer<br/>8.7 MB, k-layer cross-attn<br/>gradient decay 0.5]
QFmr --> Actor[DiT Actor / System 1<br/>+ DINOv2 visual encoder]
State[Optional state] -.-> Actor
Actor --> Act[Actions]
end
Naive fine-tuning of a VLM into a VLA collapses its spatial-grounding capability β the paper measures RefCOCO-g [email protected] dropping to near-random by 20k steps of action-only training (Fig. 3a). Naive co-training with spatial+action data partially preserves perception but oscillates. The root cause is gradient subspace misalignment between the spatial-grounding objective and the action-prediction objective: the paper measures Projection-Space Similarity (PSS, Raghu et al. 2017) between the two objectives' gradients and finds vanilla co-training yields PSS only 0.25.
SP-VLA's bet: separate where and what to act (spatial priors in a VLM) from how to act (embodiment-specific control via a DiT actor), and explicitly align their optimization dynamics via spatial prompting + a gradient-decayed querying transformer.
- System 2 (VLM Planner): Qwen2.5-VL-3B-Instruct as the multimodal encoder for spatial+semantic priors.
- System 1 (DiT Actor): compact diffusion transformer (Peebles & Xie 2023) plus a DINOv2 visual encoder for the actor's own observations. Trained on robot demonstration data.
- Connector: a lightweight querying transformer (~8.7 MB), implemented as a k-layer cross-attention module where learnable query tokens attend to k intermediate VLM layers (k=1 in the default configuration β only the final VLM layer feeds the actor). Maps variable-length VLM tokens into a fixed query-token set, stabilizing actor learning.
- Gradient decay: within the querying transformer, gradients flowing back from the DiT actor to the VLM are multiplied by 0.5 to attenuate disruption of the VLM's multimodal knowledge while still allowing co-adaptation.
The VLM is trained on a mixture of four QA categories:
- General QA (LLaVA-OneVision, InternVL3): captioning, VQA, OCR, knowledge grounding.
- Bounding Box QA (RefCOCO, ASv2, COCO-ReM, RoboRefIt, ST4VLA Manipulation Dataset).
- Point QA (Pixmo-Points filtered to β€10 pts/image, RoboPoint, RefSpatial, ST4VLA point subset). All coordinates in absolute coords aligned with Qwen2.5-VL SmartResize.
- Trajectory QA (A0 ManiSkill subset, ST4VLA trajectory subset, MolmoAct).
This produces a VLM that scores 60.5% on Where2Place (point), 83.4 [email protected] on RoboRefIt (box), 3.6 L2 on A0 ManiSkill (trajectory), versus 0%/69.0/(β) for the un-pretrained Qwen2.5-VL-3B-Instruct baseline (Table 6).
For every action-data instruction the model sees a spatial prompt appended:
"store all toys into the toy box" β "Identify all relevant toys and their spatial relationships to the container." "{instruction}. Figure out how to execute it, then locate the key object needed."
The prompt elicits the VLM's internal spatial reasoning before tokens reach the actor. Co-training continues with multimodal grounding data alongside action data, but now with the spatial prompt as the bridge. The VLM is also updated via next-token-prediction on the image-prompt pairs, while the DiT actor is updated by behavior cloning.
This raises Projection-Space Similarity from 0.25 to 0.42, indicating the two objectives now share a much larger gradient subspace β measurable evidence that the perception and action losses cooperate rather than fight.
LIBERO setup (Sec. B.2): per-suite fine-tuning on 8Γ A100, batch 128, action chunk 8, ~30k steps (~20 hours per suite), 500 trials per evaluation. 244K pick-and-place demos for the large-scale Isaac-Sim post-training set. 22 hours of teleop demos for long-horizon tasks. Real-world post-training uses 1K trajectories across 23 objects + 5 containers. (Detailed LR/batch tables not extracted from the body β partly in Appendix.)
| Models | TextVQA-OCR | RefCOCO-g IoU0.5 | Where2place (pt-Acc) | RoboRefIt (Acc0.5) | Google Robot VM/VA | WidowX VM |
|---|---|---|---|---|---|---|
| Vanilla VLA | β | β | β | β | 66.1 / 63.5 | 54.7 |
| Vanilla co-train | 20.5 | 47.1 | 21.4 | 66.7 | 70.2 / 66.5 | 61.1 |
| +Spatially Guided | 28.4 | 68.1 | 25.5 | 72.5 | 78.8 / 70.0 | 67.4 |
| +Spatially Pretrained | 28.6 | 71.2 | 25.5 | 74.3 | 84.6 / 75.9 | 73.2 |
The two stages are individually helpful but compound when stacked.
| Model | Pick Coke | Move Near | Open/Close Drawer | Drawer-Apple | Avg |
|---|---|---|---|---|---|
| RT-1 | 85.7 | 44.2 | 73.0 | 6.5 | 52.4 |
| OpenVLA | 18.0 | 56.3 | 63.0 | 0.0 | 34.3 |
| SpatialVLA | 86.0 | 77.9 | 57.4 | β | 75.1 |
| CogACT | 91.3 | 85.0 | 71.8 | 50.9 | 74.8 |
| GR00T N1.5* | 51.7 | 54.0 | 27.8 | 7.4 | 35.2 |
| Ο0 | 72.7 | 65.3 | 38.3 | β | 58.8 |
| Ο0-FAST | 75.3 | 67.5 | 42.9 | β | 61.9 |
| Magma | 83.7 | 65.4 | 56.0 | 6.4 | 52.9 |
| Vanilla VLA | 90.0 | 69.8 | 52.5 | 52.2 | 66.1 |
| ST4VLA | 97.3 | 98.0 | 65.3 | 77.8 | 84.6 |
| Model | Pick Coke | Move Near | Open/Close | Drawer-Apple | Avg |
|---|---|---|---|---|---|
| OpenVLA | 60.8 | 67.7 | 28.8 | 0.0 | 39.3 |
| SpatialVLA | 88.0 | 82.5 | 41.8 | β | 70.7 |
| CogACT | 89.6 | 80.8 | 28.3 | 46.6 | 61.3 |
| Ο0-FAST | 77.6 | 68.2 | 31.3 | β | 59.0 |
| GR00T N1.5 | 69.3 | 68.7 | 35.8 | 4.0 | 44.5 |
| Vanilla VLA | 92.3 | 80.3 | 50.1 | 31.4 | 63.5 |
| ST4VLA | 95.6 | 74.5 | 68.0 | 65.3 | 75.9 |
| Model | SpoonβTowel | CarrotβPlate | Stack Block | EggplantβBasket | Avg |
|---|---|---|---|---|---|
| OpenVLA | 4.2 | 0.0 | 0.0 | 12.5 | 4.2 |
| Octo-Small | 41.7 | 8.2 | 0.0 | 56.7 | 26.7 |
| CogACT | 71.7 | 50.8 | 15.0 | 67.5 | 51.3 |
| SpatialVLA | 16.7 | 25.0 | 29.2 | 100.0 | 42.7 |
| GR00T N1.5 | 75.3 | 54.3 | 57.0 | 61.3 | 61.9 |
| Ο0-FAST | 29.1 | 21.9 | 10.8 | 66.6 | 48.3 |
| Magma | 37.5 | 31.0 | 12.7 | 60.5 | 35.8 |
| Vanilla VLA | 56.6 | 63.3 | 27.0 | 71.8 | 54.7 |
| ST4VLA | 80.2 | 79.2 | 35.4 | 98.0 | 73.2 |
vs. GR00T N1.5 (the closest dual-system baseline): +11.3 on WidowX (73.2 vs 61.9). On Google Robot Visual Matching ST4VLA's 84.6 also tops the strongest prior reported here (CogACT 74.8). (The paper does not state explicit "+X vs previous-SOTA" margins for the Google splits.)
| Model | Spatial | Object | Goal | Long | Avg |
|---|---|---|---|---|---|
| OpenVLA | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 |
| SpatialVLA | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 |
| CoT-VLA | 87.5 | 91.6 | 87.6 | 69.0 | 83.9 |
| GR00T N1 | 94.4 | 97.6 | 93.0 | 90.6 | 93.9 |
| Ο0 | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 |
| Ο0-FAST | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 |
| Ο0.5-KI | 98.0 | 97.8 | 95.6 | 85.8 | 94.3 |
| Vanilla VLA | 98.8 | 98.0 | 81.4 | 88.0 | 91.6 |
| ST4VLA | 98.0 | 99.0 | 93.8 | 92.6 | 95.9 |
ST4VLA tops the average and the LIBERO-Long suite (longest-horizon), specifically targeted by the dual-system planner.
| Model | In dist. | New inst. | Similar dist. | New bg. | Unseen pos. | Unseen orient. | By attr. | By spatial | Avg |
|---|---|---|---|---|---|---|---|---|---|
| Ο0 | 45 | 32 | 25 | 27 | 18 | 32 | 37 | 31 | 31 |
| GR00T N1.5 | 78 | 46 | 40 | 47 | 20 | 40 | 59 | 53 | 48 |
| ST4VLA | 92 | 62 | 49 | 63 | 52 | 72 | 73 | 61 | 65 |
Largest gaps over GR00T N1.5: unseen object orientation (+32), unseen object position (+32), in-distribution (+14), new background (+16).
ST4VLA wins all four splits (in-dist, unseen object, new background, unseen instruction); average S.R. 75% vs Ο0 51% and GR00T N1.5 65% (extracted from Fig. 4).
ST4VLA tops Ο0 and GR00T N1.5 on every cell of {desktop sorting, drawer organization, sandwich making} Γ {in-dist, physical interference, task replanning}. Quantitative numbers from Fig. 5: ST4VLA averaging in the high 50s to high 60s; Ο0 in the 30sβ40s; GR00T N1.5 in the 50s.
| Pretraining Data | Where2Place pt-Acc | RoboRefIt IoU0.5 | A0 L2 | Google Robot VM/VA | WidowX VM |
|---|---|---|---|---|---|
| No additional pretraining | 0 | 69.0 | β | 66.1 / 63.5 | 54.9 |
| + General Grounding (LLaVA-OV, RefCOCO) | 30.7 | 74.9 | β | 72.6 / 70.3 | 65.2 |
| + Robotic Grounding (RoboRefIt + A0 + ST4VLA data) | 60.5 | 83.4 | 3.6 | 84.3 / 75.9 | 73.1 |
General-domain grounding alone is worth +6.5 / +6.8 / +10.3 on the three SimplerEnv splits; adding robot-specific spatial data adds another +11.7 / +5.6 / +7.9.
Vanilla VLA vs Vanilla co-train vs ST4VLA across training. ST4VLA preserves 70% of original RefCOCO-g performance while reaching 60% WidowX SR by 20k steps; vanilla VLA's RefCOCO-g drops to near-random; vanilla co-train oscillates. Quantified by PSS: vanilla co-train 0.25, ST4VLA 0.42.
Baselines saturate at lower SR even with much more compute β ST4VLA's gain is in the policy upper bound, not just convergence speed.
The paper does not contain a dedicated Limitations section but has a failure-case study (Sec. E.5). Failures cluster around:
- Incorrect grasps in cluttered scenes ("grasp by the right grasp pose").
- Container misidentification under similar-distractor conditions.
- The authors note: "While some failures may stem from sensor limitations, integrating additional modalities, such as depth sensing and proprioceptive feedback, could improve performance. We leave this as future work."
Implicit limitations:
- The spatial-prompting recipe still requires a hand-designed prompt template per task family.
- The dual-system design adds latency vs monolithic VLAs (no quantitative inference-speed reporting in the body).
- Stage 1 needs 2.3M+ spatial QA examples β a heavy data engineering investment.
SP-VLA is the VLM-side answer to the same "where does the 3D signal enter a VLA?" question that FALCON (action-head injection) and Spatial Forcing (alignment loss) attack from other angles. Where FALCON keeps the VLM frozen and routes spatial tokens past it, SP-VLA argues you should bake spatial competence into the VLM itself through pretraining, then use prompting at action time to surface it.
The paper's distinguishing methodological contribution is the PSS measurement of perception-action gradient alignment (0.25 β 0.42) β a quantitative diagnostic for why naive co-training fails, and a leverage point for design choices like the gradient-decayed querying transformer.
Versus other spatial / 3D contemporaries:
- Magma (Yang et al., 2025) β also uses spatial pretraining, but does not use spatial prompting at action time. SP-VLA outperforms it consistently on SimplerEnv (84.6 vs 52.9 Google-VM).
- SpatialVLA β uses learnable spatial embeddings injected into the VLM. SP-VLA's dual-system + spatial prompting exceeds it by 9.5+ on every SimplerEnv split.
- GR00T N1.5 (NVIDIA) β closest in dual-system spirit (high-level planner + low-level controller). SP-VLA's pretraining on point/box/trajectory QA gives it a measurable edge on SimplerEnv WidowX (+11.3) and on real-world unseen-orientation tasks (+32).
- Ο0/Ο0.5-KI β flow-matching monolithic VLAs that dominate LIBERO. SP-VLA matches them on LIBERO Avg (95.9 vs 94.3) and exceeds on the LIBERO-Long suite (92.6 vs 85.8), suggesting spatial priors specifically help long-horizon decomposition.
The paper's positioning quote: "strategically separates where and what to act from how to act." This is the core architectural philosophy now common across GR00T N1.5, Ο0.5, and OneTwoVLA β SP-VLA's specific contribution is the spatial-prompting + querying-transformer recipe to make the planner-to-actor link spatially grounded rather than just symbolic.
- OpenReview: https://openreview.net/forum?id=eKhOrQWAVJ
- PDF: https://openreview.net/pdf?id=eKhOrQWAVJ
- Code/data/models: https://internrobotics.github.io/internvla-m1.github.io
- Spatial Forcing β alignment-loss approach
- FALCON β action-head spatial-token approach
- PA3FF β natively 3D part-aware features
- GR00T N1.5
- Ο0.6
- MV-RoboBench
β Back to ICLR-2026