ICLR 2026 Spatially Guided - Heungwoo/research GitHub Wiki

SP-VLA / ST4VLA β€” Spatially Guided Training for VLAs

Venue: ICLR 2026 Category: VLA Training β€” Spatial pretraining + dual-system action expert Trend tag: Spatial / 3D for VLA Affiliations: Shanghai AI Lab, HKUST, SUSTech, Fudan Internal name: ST4VLA (released as part of the InternVLA-M1 framework)

Approach diagram

flowchart LR
  subgraph S1[Stage 1: Spatial Grounding Pre-training]
    Web[LLaVA-OneVision + RefCOCO + COCO-ReM] --> VLM
    Robot[RoboRefIt + A0 + ST4VLA Manip Data] --> VLM
    VLM[Qwen2.5-VL-3B-Instruct<br/>Box / Point / Trajectory QA<br/>~3M data, 2.3M spatial]
  end
  subgraph S2[Stage 2: Spatially Guided Action Post-training]
    VLM --> Plan[VLM Planner / System 2]
    Prompt["Spatial prompt:<br/>'Figure out how to execute it,<br/>then locate the key object'"]
    Prompt --> Plan
    Plan --> QFmr[Querying Transformer<br/>8.7 MB, k-layer cross-attn<br/>gradient decay 0.5]
    QFmr --> Actor[DiT Actor / System 1<br/>+ DINOv2 visual encoder]
    State[Optional state] -.-> Actor
    Actor --> Act[Actions]
  end
Loading

Problem

Naive fine-tuning of a VLM into a VLA collapses its spatial-grounding capability β€” the paper measures RefCOCO-g [email protected] dropping to near-random by 20k steps of action-only training (Fig. 3a). Naive co-training with spatial+action data partially preserves perception but oscillates. The root cause is gradient subspace misalignment between the spatial-grounding objective and the action-prediction objective: the paper measures Projection-Space Similarity (PSS, Raghu et al. 2017) between the two objectives' gradients and finds vanilla co-training yields PSS only 0.25.

SP-VLA's bet: separate where and what to act (spatial priors in a VLM) from how to act (embodiment-specific control via a DiT actor), and explicitly align their optimization dynamics via spatial prompting + a gradient-decayed querying transformer.

Detailed Method

Dual-system architecture

  • System 2 (VLM Planner): Qwen2.5-VL-3B-Instruct as the multimodal encoder for spatial+semantic priors.
  • System 1 (DiT Actor): compact diffusion transformer (Peebles & Xie 2023) plus a DINOv2 visual encoder for the actor's own observations. Trained on robot demonstration data.
  • Connector: a lightweight querying transformer (~8.7 MB), implemented as a k-layer cross-attention module where learnable query tokens attend to k intermediate VLM layers (k=1 in the default configuration β†’ only the final VLM layer feeds the actor). Maps variable-length VLM tokens into a fixed query-token set, stabilizing actor learning.
  • Gradient decay: within the querying transformer, gradients flowing back from the DiT actor to the VLM are multiplied by 0.5 to attenuate disruption of the VLM's multimodal knowledge while still allowing co-adaptation.

Stage 1 β€” Spatial Grounding Pre-training (~3M samples; >2.3M spatial)

The VLM is trained on a mixture of four QA categories:

  • General QA (LLaVA-OneVision, InternVL3): captioning, VQA, OCR, knowledge grounding.
  • Bounding Box QA (RefCOCO, ASv2, COCO-ReM, RoboRefIt, ST4VLA Manipulation Dataset).
  • Point QA (Pixmo-Points filtered to ≀10 pts/image, RoboPoint, RefSpatial, ST4VLA point subset). All coordinates in absolute coords aligned with Qwen2.5-VL SmartResize.
  • Trajectory QA (A0 ManiSkill subset, ST4VLA trajectory subset, MolmoAct).

This produces a VLM that scores 60.5% on Where2Place (point), 83.4 [email protected] on RoboRefIt (box), 3.6 L2 on A0 ManiSkill (trajectory), versus 0%/69.0/(–) for the un-pretrained Qwen2.5-VL-3B-Instruct baseline (Table 6).

Stage 2 β€” Spatially Guided Action Post-training

For every action-data instruction the model sees a spatial prompt appended:

"store all toys into the toy box" β†’ "Identify all relevant toys and their spatial relationships to the container." "{instruction}. Figure out how to execute it, then locate the key object needed."

The prompt elicits the VLM's internal spatial reasoning before tokens reach the actor. Co-training continues with multimodal grounding data alongside action data, but now with the spatial prompt as the bridge. The VLM is also updated via next-token-prediction on the image-prompt pairs, while the DiT actor is updated by behavior cloning.

This raises Projection-Space Similarity from 0.25 to 0.42, indicating the two objectives now share a much larger gradient subspace β€” measurable evidence that the perception and action losses cooperate rather than fight.

Hyperparameters

LIBERO setup (Sec. B.2): per-suite fine-tuning on 8Γ— A100, batch 128, action chunk 8, ~30k steps (~20 hours per suite), 500 trials per evaluation. 244K pick-and-place demos for the large-scale Isaac-Sim post-training set. 22 hours of teleop demos for long-horizon tasks. Real-world post-training uses 1K trajectories across 23 objects + 5 containers. (Detailed LR/batch tables not extracted from the body β€” partly in Appendix.)

Comprehensive Results

Effect of training strategy (vs base Vanilla VLA)

Models TextVQA-OCR RefCOCO-g IoU0.5 Where2place (pt-Acc) RoboRefIt (Acc0.5) Google Robot VM/VA WidowX VM
Vanilla VLA – – – – 66.1 / 63.5 54.7
Vanilla co-train 20.5 47.1 21.4 66.7 70.2 / 66.5 61.1
+Spatially Guided 28.4 68.1 25.5 72.5 78.8 / 70.0 67.4
+Spatially Pretrained 28.6 71.2 25.5 74.3 84.6 / 75.9 73.2

The two stages are individually helpful but compound when stacked.

SimplerEnv Google Robot β€” Visual Matching (paper-reported numbers; * = re-implemented)

Model Pick Coke Move Near Open/Close Drawer Drawer-Apple Avg
RT-1 85.7 44.2 73.0 6.5 52.4
OpenVLA 18.0 56.3 63.0 0.0 34.3
SpatialVLA 86.0 77.9 57.4 – 75.1
CogACT 91.3 85.0 71.8 50.9 74.8
GR00T N1.5* 51.7 54.0 27.8 7.4 35.2
Ο€0 72.7 65.3 38.3 – 58.8
Ο€0-FAST 75.3 67.5 42.9 – 61.9
Magma 83.7 65.4 56.0 6.4 52.9
Vanilla VLA 90.0 69.8 52.5 52.2 66.1
ST4VLA 97.3 98.0 65.3 77.8 84.6

SimplerEnv Google Robot β€” Variant Aggregation

Model Pick Coke Move Near Open/Close Drawer-Apple Avg
OpenVLA 60.8 67.7 28.8 0.0 39.3
SpatialVLA 88.0 82.5 41.8 – 70.7
CogACT 89.6 80.8 28.3 46.6 61.3
Ο€0-FAST 77.6 68.2 31.3 – 59.0
GR00T N1.5 69.3 68.7 35.8 4.0 44.5
Vanilla VLA 92.3 80.3 50.1 31.4 63.5
ST4VLA 95.6 74.5 68.0 65.3 75.9

SimplerEnv WidowX

Model Spoon→Towel Carrot→Plate Stack Block Eggplant→Basket Avg
OpenVLA 4.2 0.0 0.0 12.5 4.2
Octo-Small 41.7 8.2 0.0 56.7 26.7
CogACT 71.7 50.8 15.0 67.5 51.3
SpatialVLA 16.7 25.0 29.2 100.0 42.7
GR00T N1.5 75.3 54.3 57.0 61.3 61.9
Ο€0-FAST 29.1 21.9 10.8 66.6 48.3
Magma 37.5 31.0 12.7 60.5 35.8
Vanilla VLA 56.6 63.3 27.0 71.8 54.7
ST4VLA 80.2 79.2 35.4 98.0 73.2

vs. GR00T N1.5 (the closest dual-system baseline): +11.3 on WidowX (73.2 vs 61.9). On Google Robot Visual Matching ST4VLA's 84.6 also tops the strongest prior reported here (CogACT 74.8). (The paper does not state explicit "+X vs previous-SOTA" margins for the Google splits.)

LIBERO (per-suite, 500 trials each, 8Γ— A100, ~30k steps, ~20h)

Model Spatial Object Goal Long Avg
OpenVLA 84.7 88.4 79.2 53.7 76.5
SpatialVLA 88.2 89.9 78.6 55.5 78.1
CoT-VLA 87.5 91.6 87.6 69.0 83.9
GR00T N1 94.4 97.6 93.0 90.6 93.9
Ο€0 96.8 98.8 95.8 85.2 94.2
Ο€0-FAST 96.4 96.8 88.6 60.2 85.5
Ο€0.5-KI 98.0 97.8 95.6 85.8 94.3
Vanilla VLA 98.8 98.0 81.4 88.0 91.6
ST4VLA 98.0 99.0 93.8 92.6 95.9

ST4VLA tops the average and the LIBERO-Long suite (longest-horizon), specifically targeted by the dual-system planner.

Real-world Pick-and-Place (Franka FR3, 1K demos, 23 objects + 5 containers)

Model In dist. New inst. Similar dist. New bg. Unseen pos. Unseen orient. By attr. By spatial Avg
Ο€0 45 32 25 27 18 32 37 31 31
GR00T N1.5 78 46 40 47 20 40 59 53 48
ST4VLA 92 62 49 63 52 72 73 61 65

Largest gaps over GR00T N1.5: unseen object orientation (+32), unseen object position (+32), in-distribution (+14), new background (+16).

Large-scale simulated pick-and-place (Isaac-Sim, 200 tasks, 3000+ objects)

ST4VLA wins all four splits (in-dist, unseen object, new background, unseen instruction); average S.R. 75% vs Ο€0 51% and GR00T N1.5 65% (extracted from Fig. 4).

Long-horizon real-robot tasks (22h teleop demos, 3 tasks, 3 settings each)

ST4VLA tops Ο€0 and GR00T N1.5 on every cell of {desktop sorting, drawer organization, sandwich making} Γ— {in-dist, physical interference, task replanning}. Quantitative numbers from Fig. 5: ST4VLA averaging in the high 50s to high 60s; Ο€0 in the 30s–40s; GR00T N1.5 in the 50s.

Ablation Studies

Stage 1 pretraining data (Table 6)

Pretraining Data Where2Place pt-Acc RoboRefIt IoU0.5 A0 L2 Google Robot VM/VA WidowX VM
No additional pretraining 0 69.0 – 66.1 / 63.5 54.9
+ General Grounding (LLaVA-OV, RefCOCO) 30.7 74.9 – 72.6 / 70.3 65.2
+ Robotic Grounding (RoboRefIt + A0 + ST4VLA data) 60.5 83.4 3.6 84.3 / 75.9 73.1

General-domain grounding alone is worth +6.5 / +6.8 / +10.3 on the three SimplerEnv splits; adding robot-specific spatial data adds another +11.7 / +5.6 / +7.9.

Spatial-prompting ablation (Sec. 3.1, Fig. 3 + Table 1)

Vanilla VLA vs Vanilla co-train vs ST4VLA across training. ST4VLA preserves 70% of original RefCOCO-g performance while reaching 60% WidowX SR by 20k steps; vanilla VLA's RefCOCO-g drops to near-random; vanilla co-train oscillates. Quantified by PSS: vanilla co-train 0.25, ST4VLA 0.42.

Extended training to 100k steps (Fig. 6)

Baselines saturate at lower SR even with much more compute β€” ST4VLA's gain is in the policy upper bound, not just convergence speed.

Limitations (as stated)

The paper does not contain a dedicated Limitations section but has a failure-case study (Sec. E.5). Failures cluster around:

  • Incorrect grasps in cluttered scenes ("grasp by the right grasp pose").
  • Container misidentification under similar-distractor conditions.
  • The authors note: "While some failures may stem from sensor limitations, integrating additional modalities, such as depth sensing and proprioceptive feedback, could improve performance. We leave this as future work."

Implicit limitations:

  • The spatial-prompting recipe still requires a hand-designed prompt template per task family.
  • The dual-system design adds latency vs monolithic VLAs (no quantitative inference-speed reporting in the body).
  • Stage 1 needs 2.3M+ spatial QA examples β€” a heavy data engineering investment.

Significance & Positioning

SP-VLA is the VLM-side answer to the same "where does the 3D signal enter a VLA?" question that FALCON (action-head injection) and Spatial Forcing (alignment loss) attack from other angles. Where FALCON keeps the VLM frozen and routes spatial tokens past it, SP-VLA argues you should bake spatial competence into the VLM itself through pretraining, then use prompting at action time to surface it.

The paper's distinguishing methodological contribution is the PSS measurement of perception-action gradient alignment (0.25 β†’ 0.42) β€” a quantitative diagnostic for why naive co-training fails, and a leverage point for design choices like the gradient-decayed querying transformer.

Versus other spatial / 3D contemporaries:

  • Magma (Yang et al., 2025) β€” also uses spatial pretraining, but does not use spatial prompting at action time. SP-VLA outperforms it consistently on SimplerEnv (84.6 vs 52.9 Google-VM).
  • SpatialVLA β€” uses learnable spatial embeddings injected into the VLM. SP-VLA's dual-system + spatial prompting exceeds it by 9.5+ on every SimplerEnv split.
  • GR00T N1.5 (NVIDIA) β€” closest in dual-system spirit (high-level planner + low-level controller). SP-VLA's pretraining on point/box/trajectory QA gives it a measurable edge on SimplerEnv WidowX (+11.3) and on real-world unseen-orientation tasks (+32).
  • Ο€0/Ο€0.5-KI β€” flow-matching monolithic VLAs that dominate LIBERO. SP-VLA matches them on LIBERO Avg (95.9 vs 94.3) and exceeds on the LIBERO-Long suite (92.6 vs 85.8), suggesting spatial priors specifically help long-horizon decomposition.

The paper's positioning quote: "strategically separates where and what to act from how to act." This is the core architectural philosophy now common across GR00T N1.5, Ο€0.5, and OneTwoVLA β€” SP-VLA's specific contribution is the spatial-prompting + querying-transformer recipe to make the planner-to-actor link spatially grounded rather than just symbolic.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️