ICLR 2026 SP VLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture โ Efficiency Trend tag: Efficiency / real-time deployment Affiliation: Tsinghua University ยท CUHK ยท UIUC ยท Beihang University
flowchart LR
Hist[Action buffer S_A<br/>last n=6 actions] --> SCH[Action-aware Scheduler<br/>velocity check + intuitive ratio ฯ=0.5]
SCH -- "Deliberative path<br/>low speed or low intuitive ratio" --> TP[Spatio-Semantic Token Pruning<br/>velocity-adaptive prune ratio]
TP --> ENC[ViT Encoder]
ENC --> LLM[7B LLM<br/>e.g. Llama 2 / OpenVLA / CogACT]
LLM --> AH[Action Head<br/>D-tokenizer or diffusion]
SCH -- "Intuitive path<br/>v in v_min..v_max and N_G/N_A above ฯ" --> RIDGE["Lightweight Generator<br/>Ridge Regression on action buffer<br/>under 1K params"]
RIDGE --> CHK[Validity check]
AH --> ACT[Robot action]
CHK --> ACT
ACT -. feedback .-> Hist
VLA models are too compute-heavy for real-time deployment. OpenVLA is 7B+, RT-X is up to 55B; even on a 4090 inference runs at ~4 Hz. Existing acceleration work attacks single-step compute via quantization (QAIL), early-exit (DeeR-VLA), token caching (VLA-Cache), parallel decoding (PD-VLA, OpenVLA-OFT), or speculative decoding โ but all treat one frame in isolation.
SP-VLA argues two redundancies are systematically ignored:
- Temporal redundancy in the sequential action stream (most steps don't need full reasoning โ humans alternate "deliberate" vs "intuitive" motor control).
- Spatial redundancy in visual tokens (most patches are background โ but VLA-targeted token pruning has unique constraints).
Behavioral analysis. Across 50 pick-and-place trials, the authors observe the manipulator follows a 4-phase velocity profile: target โ grasp โ move โ place. Slow alignment, then high-speed translation, then careful action, etc. They argue VLA models have implicitly learned this kinematic pattern, so the action stream is naturally split into:
- Deliberative actions โ slow, precise (grasping, turning, placement). Need the full 7B VLA.
- Intuitive actions โ fast, ballistic (point-to-point translation). Can be approximated by a tiny model.
Scheduler logic. Let a^t_d = (a_x, a_y, a_z) be the per-step end-effector translational velocity. Define:
- An action a is intuitive if all components |a_i| > v_min (speed threshold).
- The lightweight model is allowed when (a) a_{t-1} โ [v_min, v_max] and (b) the ratio of recent VLA-generated actions in the buffer N_G / N_A > ฯ (default ฯ = 0.5).
LWM = 1 if both hold, else 0. (Equation 1.) This drives small-step, high-frequency model switching โ even within an "intuitive" segment, the VLA is invoked periodically to correct drift.
Lightweight generator: Ridge Regression on the action buffer S_A = {a_{t-n}, โฆ, a_{t-1}} (n=6).
- X = [T, 1] โ R^{nร2}, T = [0,โฆ,nโ1]แต; Y = action buffer; ฮฒ โ R^{2รโ}.
- Minimize J(ฮฒ) = โXฮฒ โ Yโยฒ + ฮปโฮฒโยฒ (Tikhonov / ridge).
- Closed-form: ฮฒ = (Xแต X + ฮปI)โปยน Xแต Y.
- Predict a_t = x_t ฮฒ*, x_t = [t 1]แต.
- Gripper state is NOT regressed (it's binary) โ instead reuse the tโ1 value, leaving binary state transitions to the VLA. Predicted intuitive actions go through a validity check before execution.
Insight from controlled experiments (Fig. 2b): Random token pruning degrades but doesn't destroy task performance โ there is spatial redundancy. But two surprising failure modes:
- Reordering tokens by semantic importance (no actual pruning) causes complete task failure โ relative position of tokens carries spatial meaning the auto-regressive VLA depends on.
- Pruning purely by semantic attention scores removes object-contour background tokens and also fails โ object contours are critical for spatial grounding.
So pruning must (a) preserve relative ordering and (b) explicitly retain edge / contour tokens.
Semantic-aware token importance. From last-encoder-layer attention: Q,K,V = X W_q,k,v Attn = Softmax(QKแต/โd_k) V AccuAttn = ยฝ (eแต โ I_M) vec(Attn). Select T_se = {x_i | AccuAttn_i > t_ks}.
Spatial-aware token importance. Apply Canny edge detector to the input image: X_s = Canny(X). Then T_sp = f_E(X_s) is the ordered set of edge-region tokens.
Order-preserving union: T_select = U(T_se, T_sp) โ the union, with original token positions retained.
Velocity-adaptive prune rate. Pruning is disabled for low-speed (deliberative) actions to protect precision. For higher speeds: T_r(v) = 1 if v < v_pmin, else 1 โ (v โ v_pmin) / (v_pmax โ v_pmin).
This couples the spatial dimension (how aggressively to prune) to the temporal mode โ fast intuitive actions can tolerate aggressive pruning; precise grasps cannot.
- Buffer size n = 6.
- Deliberation/intuition ratio threshold ฯ = 0.5.
- Velocity thresholds: v_min = 0.2, v_max = 0.5 (paper settings); token-pruning velocity threshold v_pmin = 0.5. Sensitivity study (Table 7) varies these ยฑ25% and finds accuracy robust to speed but sensitive to n and ฯ.
- Hardware: NVIDIA A100 GPUs for experiments/training; NVIDIA RTX 4090 (40GB) for frequency/latency measurements (paper Sec. on Frequency and Latency).
| Method | Goal | Object | Spatial | Long | Avg | Speedup | FLOPs % |
|---|---|---|---|---|---|---|---|
| OpenVLA | 75.40 | 86.20 | 83.80 | 53.00 | 74.60 | 1.00ร | 100 |
| SparseVLM | 74.20 | 84.00 | 83.40 | 52.80 | 73.60 | 1.33ร | 75.55 |
| FoPru + R | 59.80 | 81.20 | 71.60 | 26.20 | 59.70 | 1.31ร | 77.20 |
| PruMerge + R | 0.00 | 0.00 | 0.00 | 0.00 | 0.00 | 1.36ร | 73.63 |
| FastVLM + R + S | 73.20 | 77.00 | 79.80 | 36.60 | 66.65 | 1.16ร | 86.22 |
| VisionZip + R + S | 46.00 | 47.40 | 34.20 | 4.60 | 33.05 | 1.21ร | 81.95 |
| Ours (Speed-priority) | 73.60 | 82.40 | 80.00 | 51.60 | 71.90 | 1.50ร | 66.51 |
| Ours (Acc-priority) | 75.40 | 85.60 | 84.40 | 54.20 | 74.90 | 1.35ร | 73.64 |
"+R" = preserve relative token positions. "+S" = add Canny edges. SP-VLA is the only method to deliver acceleration with โฅ baseline accuracy. PruMerge is interesting โ it gets a 1.36ร speedup but drops to 0% on every suite (the model collapses).
Visual Matching split:
| Method | PickCan | MoveNear | Drawer | DrawerApple | Avg | Speedup | FLOPs % |
|---|---|---|---|---|---|---|---|
| CogACT | 91.30 | 85.00 | 71.80 | 50.90 | 74.80 | 1.00ร | 100 |
| Random Drop | 9.70 | 20.40 | 53.50 | 0.00 | 20.90 | 1.20ร | 58.50 |
| FastV | 92.60 | 81.40 | 69.80 | 52.40 | 74.10 | 1.21ร | 42.00 |
| VLA-Cache | 92.00 | 83.30 | 70.50 | 51.60 | 74.40 | 1.38ร | 80.10 |
| EfficientVLA | 93.30 | 81.30 | 68.20 | 53.80 | 74.20 | 1.93ร | 28.90 |
| Ours | 90.00 | 82.08 | 75.35 | 52.78 | 75.05 | 2.15ร | 38.15 |
Visual Aggregation split:
| Method | PickCan | MoveNear | Drawer | DrawerApple | Avg | Speedup |
|---|---|---|---|---|---|---|
| CogACT | 89.60 | 80.80 | 28.30 | 46.60 | 61.30 | 1.00ร |
| EfficientVLA | 93.20 | 75.80 | 26.90 | 49.20 | 61.20 | 1.91ร |
| Ours | 86.18 | 77.33 | 55.29 | 41.80 | 65.16 | 2.09ร |
Note +27 pp on Drawer in Visual-Aggregation while still 1.81ร faster โ interpreted as error-correction (the scheduler/lightweight generator smooths trajectories where CogACT alone falters).
WidowX split (bottom of Table 2; paper labels it "WindowX"):
| Method | PutSpoon | PutCarrot | StackBlock | PutEggplant | Avg | Speedup |
|---|---|---|---|---|---|---|
| CogACT | 71.70 | 50.80 | 15.00 | 67.50 | 51.30 | 1.00ร |
| Ours | 70.83 | 54.17 | 29.17 | 75.00 | 57.29 | 2.41ร |
Stack Block jumps from 15 โ 29 with 2.54ร acceleration.
| Setting | CogACT freq | SP-VLA freq | CogACT latency | SP-VLA latency |
|---|---|---|---|---|
| SimplerEnv Visual-Matching avg | 3.77 Hz | 8.06 Hz | 0.27 s | 0.13 s |
| SimplerEnv WindowX avg | 3.96 Hz | 8.69 Hz | 0.25 s | 0.12 s |
โ 2.2ร frequency improvement, ~50% latency reduction.
The ICLR camera-ready / arXiv v3 paper (2506.12723v3, Oct 2025) reports no real-robot experiment table โ all benchmarks are simulation (LIBERO + SimplerEnv). The project's GitHub README separately states a real Franka Panda result of ~2.5ร end-to-end inference acceleration with only a 1% success-rate drop, but provides no per-task breakdown, FLOPs, or latency numbers. Treat any detailed real-robot figures with caution until the source table is located.
Audit note: a previously listed "Real Franka Research 3 (Table 4)" table with per-task success rates (80/74/77 vs 78/74/76), FLOPs 35.55, 0.27/0.13 s latency, "150 trajectories per task," and "20 morning / 10 noon / 20 evening" lighting splits was not found in any reachable source and has been removed as a likely fabrication.
| Variant | Goal | Object | Spatial | Long | Avg | Speedup |
|---|---|---|---|---|---|---|
| Full SP-VLA | 75.40 | 85.60 | 84.40 | 54.20 | 74.90 | 1.35ร |
| w/o Pruning | 74.40 | 84.20 | 84.00 | 53.30 | 73.98 | 1.27ร |
| w/o Scheduling | 77.31 | 81.80 | 79.00 | 48.00 | 71.52 | 1.21ร |
| w/o Canny (no edge tokens) | 33.60 | 39.00 | 22.00 | 1.10 | 23.93 | 1.35ร |
Without Canny edge tokens the model effectively collapses (74.9 โ 23.9), confirming the central insight that VLA spatial perception relies on object-contour tokens.
- Token pruning identifies ~22.7% redundancy on LIBERO-Spatial.
- Model scheduling identifies ~21% redundancy on LIBERO-Object.
- Intuitive-action proportion grows with task length: 18% on LIBERO-Spatial โ 28% on LIBERO-Long โ corresponding speedups 1.18ร โ 1.39ร.
The temporal and spatial dimensions are not orthogonal โ different tasks expose redundancy along different axes โ so combining both is multiplicative, not additive.
The authors flag a single explicit limitation:
- Intuitive-action generation is preliminary. Currently they only "lightweight-ify" the VLA via Ridge Regression; they have not achieved a complete behavioral separation between deliberative and intuitive generation. Making the VLA more human-like by structurally separating these two modes (rather than scheduling between a big and tiny model) is flagged as the key future direction.
Other implicit caveats:
- Hyperparameter v_min, v_max are device-dependent. The 1/4-3/4 max-task-speed heuristic (v_min = 0.2, v_max = 0.5 in the paper's settings) is validated only in simulation; the sensitivity study (Table 7) shows accuracy is robust to ยฑ25% speed perturbation but sensitive to buffer size n and the intuitive-action proportion ฯ.
- Lossless acceleration claim is benchmark-specific. "1.5ร lossless on LIBERO" is the Speed-priority preset with 71.9 vs 74.6 baseline (2.7 pp drop) โ not strictly lossless. Authors' "Acc" preset gets 74.9 (no drop) at 1.35ร.
- No real-robot evaluation in the paper. v3 reports only LIBERO + SimplerEnv (simulation); the lone real Franka claim lives in the repo README without a results table.
- Lightweight generator works because intuitive segments are approximately linear; on highly nonlinear / contact-rich intuitive segments (rare in pick-and-place but common in dexterous manipulation), Ridge Regression may break down.
- vs token-pruning-only methods (Action-aware Dynamic Pruning, FastV, VLA-Cache, EfficientVLA, SparseVLM, FoPru, FastVLM, VisionZip, PruMerge): SP-VLA's spatio-semantic-with-edge pruning prevents the catastrophic failure the paper documents in pure semantic pruners (PruMerge โ 0%, VisionZip โ 33%). Combined with scheduling it gets 2.4ร SimplerEnv speedup at +0.25 pp accuracy, which no single-axis pruner achieves.
- vs quantization (AutoQVLA, QAIL): orthogonal โ could be stacked.
- vs parallel decoding (OpenVLA-OFT, PD-VLA, FASTER): these accelerate inside one VLA forward pass; SP-VLA accelerates between forward passes by skipping them. Composable.
- vs hierarchical dual-system VLAs (ฯ0.5, Hi-Robot, Helix, OneTwoVLA): dual-system designs route reasoning to a separate model; SP-VLA borrows the System-1 / System-2 metaphor but applies it to action-typing rather than reasoning. The "lightweight generator" is essentially a System-1 motor primitive (ridge regression) for intuitive arm trajectories, while the full VLA acts as System-2 for grasp-class moments.
- vs early-exit (DeeR-VLA): DeeR exits earlier in the LLM stack on easy frames; SP-VLA skips the LLM entirely on intuitive frames, which is a stronger speedup ceiling.
- First systematic treatment of temporal ร spatial redundancy in VLA inference โ the action-type indicator + edge-aware pruning ablation reveals previously unappreciated structural facts about VLA computation.
- Action-aware Dynamic Pruning โ token-level efficiency
- AutoQVLA โ channel-aware quantization
- FASTER
- OneTwoVLA โ adaptive reasoning, kindred conceptual structure
- Survey: VLA & Manipulation
โ Back to ICLR-2026