ICLR 2026 SP VLA - Heungwoo/research GitHub Wiki

SP-VLA โ€” Joint Model Scheduling and Token Pruning for VLA Acceleration

Venue: ICLR 2026 Category: VLA Architecture โ€” Efficiency Trend tag: Efficiency / real-time deployment Affiliation: Tsinghua University ยท CUHK ยท UIUC ยท Beihang University

Approach diagram

flowchart LR
  Hist[Action buffer S_A<br/>last n=6 actions] --> SCH[Action-aware Scheduler<br/>velocity check + intuitive ratio ฯ„=0.5]
  SCH -- "Deliberative path<br/>low speed or low intuitive ratio" --> TP[Spatio-Semantic Token Pruning<br/>velocity-adaptive prune ratio]
  TP --> ENC[ViT Encoder]
  ENC --> LLM[7B LLM<br/>e.g. Llama 2 / OpenVLA / CogACT]
  LLM --> AH[Action Head<br/>D-tokenizer or diffusion]
  SCH -- "Intuitive path<br/>v in v_min..v_max and N_G/N_A above ฯ„" --> RIDGE["Lightweight Generator<br/>Ridge Regression on action buffer<br/>under 1K params"]
  RIDGE --> CHK[Validity check]
  AH --> ACT[Robot action]
  CHK --> ACT
  ACT -. feedback .-> Hist
Loading

Problem

VLA models are too compute-heavy for real-time deployment. OpenVLA is 7B+, RT-X is up to 55B; even on a 4090 inference runs at ~4 Hz. Existing acceleration work attacks single-step compute via quantization (QAIL), early-exit (DeeR-VLA), token caching (VLA-Cache), parallel decoding (PD-VLA, OpenVLA-OFT), or speculative decoding โ€” but all treat one frame in isolation.

SP-VLA argues two redundancies are systematically ignored:

  • Temporal redundancy in the sequential action stream (most steps don't need full reasoning โ€” humans alternate "deliberate" vs "intuitive" motor control).
  • Spatial redundancy in visual tokens (most patches are background โ€” but VLA-targeted token pruning has unique constraints).

Detailed Method

3.1 Action-Type-Aware Model Scheduling

Behavioral analysis. Across 50 pick-and-place trials, the authors observe the manipulator follows a 4-phase velocity profile: target โ†’ grasp โ†’ move โ†’ place. Slow alignment, then high-speed translation, then careful action, etc. They argue VLA models have implicitly learned this kinematic pattern, so the action stream is naturally split into:

  • Deliberative actions โ€” slow, precise (grasping, turning, placement). Need the full 7B VLA.
  • Intuitive actions โ€” fast, ballistic (point-to-point translation). Can be approximated by a tiny model.

Scheduler logic. Let a^t_d = (a_x, a_y, a_z) be the per-step end-effector translational velocity. Define:

  • An action a is intuitive if all components |a_i| > v_min (speed threshold).
  • The lightweight model is allowed when (a) a_{t-1} โˆˆ [v_min, v_max] and (b) the ratio of recent VLA-generated actions in the buffer N_G / N_A > ฯ„ (default ฯ„ = 0.5).

LWM = 1 if both hold, else 0. (Equation 1.) This drives small-step, high-frequency model switching โ€” even within an "intuitive" segment, the VLA is invoked periodically to correct drift.

Lightweight generator: Ridge Regression on the action buffer S_A = {a_{t-n}, โ€ฆ, a_{t-1}} (n=6).

  • X = [T, 1] โˆˆ R^{nร—2}, T = [0,โ€ฆ,nโˆ’1]แต€; Y = action buffer; ฮฒ โˆˆ R^{2ร—โ„“}.
  • Minimize J(ฮฒ) = โ€–Xฮฒ โˆ’ Yโ€–ยฒ + ฮปโ€–ฮฒโ€–ยฒ (Tikhonov / ridge).
  • Closed-form: ฮฒ = (Xแต€ X + ฮปI)โปยน Xแต€ Y.
  • Predict a_t = x_t ฮฒ*, x_t = [t 1]แต€.
  • Gripper state is NOT regressed (it's binary) โ€” instead reuse the tโˆ’1 value, leaving binary state transitions to the VLA. Predicted intuitive actions go through a validity check before execution.

3.2 Spatio-Semantic Dual-Aware Token Pruning

Insight from controlled experiments (Fig. 2b): Random token pruning degrades but doesn't destroy task performance โ€” there is spatial redundancy. But two surprising failure modes:

  1. Reordering tokens by semantic importance (no actual pruning) causes complete task failure โ†’ relative position of tokens carries spatial meaning the auto-regressive VLA depends on.
  2. Pruning purely by semantic attention scores removes object-contour background tokens and also fails โ†’ object contours are critical for spatial grounding.

So pruning must (a) preserve relative ordering and (b) explicitly retain edge / contour tokens.

Semantic-aware token importance. From last-encoder-layer attention: Q,K,V = X W_q,k,v Attn = Softmax(QKแต€/โˆšd_k) V AccuAttn = ยฝ (eแต€ โŠ— I_M) vec(Attn). Select T_se = {x_i | AccuAttn_i > t_ks}.

Spatial-aware token importance. Apply Canny edge detector to the input image: X_s = Canny(X). Then T_sp = f_E(X_s) is the ordered set of edge-region tokens.

Order-preserving union: T_select = U(T_se, T_sp) โ€” the union, with original token positions retained.

Velocity-adaptive prune rate. Pruning is disabled for low-speed (deliberative) actions to protect precision. For higher speeds: T_r(v) = 1 if v < v_pmin, else 1 โˆ’ (v โˆ’ v_pmin) / (v_pmax โˆ’ v_pmin).

This couples the spatial dimension (how aggressively to prune) to the temporal mode โ€” fast intuitive actions can tolerate aggressive pruning; precise grasps cannot.

Hyperparameters

  • Buffer size n = 6.
  • Deliberation/intuition ratio threshold ฯ„ = 0.5.
  • Velocity thresholds: v_min = 0.2, v_max = 0.5 (paper settings); token-pruning velocity threshold v_pmin = 0.5. Sensitivity study (Table 7) varies these ยฑ25% and finds accuracy robust to speed but sensitive to n and ฯ„.
  • Hardware: NVIDIA A100 GPUs for experiments/training; NVIDIA RTX 4090 (40GB) for frequency/latency measurements (paper Sec. on Frequency and Latency).

Comprehensive Results

LIBERO (Table 1) โ€” base = OpenVLA; 130 tasks, 2000 trajectories

Method Goal Object Spatial Long Avg Speedup FLOPs %
OpenVLA 75.40 86.20 83.80 53.00 74.60 1.00ร— 100
SparseVLM 74.20 84.00 83.40 52.80 73.60 1.33ร— 75.55
FoPru + R 59.80 81.20 71.60 26.20 59.70 1.31ร— 77.20
PruMerge + R 0.00 0.00 0.00 0.00 0.00 1.36ร— 73.63
FastVLM + R + S 73.20 77.00 79.80 36.60 66.65 1.16ร— 86.22
VisionZip + R + S 46.00 47.40 34.20 4.60 33.05 1.21ร— 81.95
Ours (Speed-priority) 73.60 82.40 80.00 51.60 71.90 1.50ร— 66.51
Ours (Acc-priority) 75.40 85.60 84.40 54.20 74.90 1.35ร— 73.64

"+R" = preserve relative token positions. "+S" = add Canny edges. SP-VLA is the only method to deliver acceleration with โ‰ฅ baseline accuracy. PruMerge is interesting โ€” it gets a 1.36ร— speedup but drops to 0% on every suite (the model collapses).

SimplerEnv (Table 2) โ€” base = CogACT

Visual Matching split:

Method PickCan MoveNear Drawer DrawerApple Avg Speedup FLOPs %
CogACT 91.30 85.00 71.80 50.90 74.80 1.00ร— 100
Random Drop 9.70 20.40 53.50 0.00 20.90 1.20ร— 58.50
FastV 92.60 81.40 69.80 52.40 74.10 1.21ร— 42.00
VLA-Cache 92.00 83.30 70.50 51.60 74.40 1.38ร— 80.10
EfficientVLA 93.30 81.30 68.20 53.80 74.20 1.93ร— 28.90
Ours 90.00 82.08 75.35 52.78 75.05 2.15ร— 38.15

Visual Aggregation split:

Method PickCan MoveNear Drawer DrawerApple Avg Speedup
CogACT 89.60 80.80 28.30 46.60 61.30 1.00ร—
EfficientVLA 93.20 75.80 26.90 49.20 61.20 1.91ร—
Ours 86.18 77.33 55.29 41.80 65.16 2.09ร—

Note +27 pp on Drawer in Visual-Aggregation while still 1.81ร— faster โ€” interpreted as error-correction (the scheduler/lightweight generator smooths trajectories where CogACT alone falters).

WidowX split (bottom of Table 2; paper labels it "WindowX"):

Method PutSpoon PutCarrot StackBlock PutEggplant Avg Speedup
CogACT 71.70 50.80 15.00 67.50 51.30 1.00ร—
Ours 70.83 54.17 29.17 75.00 57.29 2.41ร—

Stack Block jumps from 15 โ†’ 29 with 2.54ร— acceleration.

Frequency / latency on RTX 4090 (Table 4; appendix Tables 5โ€“6)

Setting CogACT freq SP-VLA freq CogACT latency SP-VLA latency
SimplerEnv Visual-Matching avg 3.77 Hz 8.06 Hz 0.27 s 0.13 s
SimplerEnv WindowX avg 3.96 Hz 8.69 Hz 0.25 s 0.12 s

โ‰ˆ 2.2ร— frequency improvement, ~50% latency reduction.

Real robot (Franka Panda)

The ICLR camera-ready / arXiv v3 paper (2506.12723v3, Oct 2025) reports no real-robot experiment table โ€” all benchmarks are simulation (LIBERO + SimplerEnv). The project's GitHub README separately states a real Franka Panda result of ~2.5ร— end-to-end inference acceleration with only a 1% success-rate drop, but provides no per-task breakdown, FLOPs, or latency numbers. Treat any detailed real-robot figures with caution until the source table is located.

Audit note: a previously listed "Real Franka Research 3 (Table 4)" table with per-task success rates (80/74/77 vs 78/74/76), FLOPs 35.55, 0.27/0.13 s latency, "150 trajectories per task," and "20 morning / 10 noon / 20 evening" lighting splits was not found in any reachable source and has been removed as a likely fabrication.

Ablation Studies (Table 2 โ€” LIBERO)

Component ablation

Variant Goal Object Spatial Long Avg Speedup
Full SP-VLA 75.40 85.60 84.40 54.20 74.90 1.35ร—
w/o Pruning 74.40 84.20 84.00 53.30 73.98 1.27ร—
w/o Scheduling 77.31 81.80 79.00 48.00 71.52 1.21ร—
w/o Canny (no edge tokens) 33.60 39.00 22.00 1.10 23.93 1.35ร—

Without Canny edge tokens the model effectively collapses (74.9 โ†’ 23.9), confirming the central insight that VLA spatial perception relies on object-contour tokens.

Acceleration source per dimension (Table 8 in appendix A.5)

  • Token pruning identifies ~22.7% redundancy on LIBERO-Spatial.
  • Model scheduling identifies ~21% redundancy on LIBERO-Object.
  • Intuitive-action proportion grows with task length: 18% on LIBERO-Spatial โ†’ 28% on LIBERO-Long โ†’ corresponding speedups 1.18ร— โ†’ 1.39ร—.

The temporal and spatial dimensions are not orthogonal โ€” different tasks expose redundancy along different axes โ€” so combining both is multiplicative, not additive.

Limitations (as stated in Appendix A.6)

The authors flag a single explicit limitation:

  1. Intuitive-action generation is preliminary. Currently they only "lightweight-ify" the VLA via Ridge Regression; they have not achieved a complete behavioral separation between deliberative and intuitive generation. Making the VLA more human-like by structurally separating these two modes (rather than scheduling between a big and tiny model) is flagged as the key future direction.

Other implicit caveats:

  • Hyperparameter v_min, v_max are device-dependent. The 1/4-3/4 max-task-speed heuristic (v_min = 0.2, v_max = 0.5 in the paper's settings) is validated only in simulation; the sensitivity study (Table 7) shows accuracy is robust to ยฑ25% speed perturbation but sensitive to buffer size n and the intuitive-action proportion ฯ„.
  • Lossless acceleration claim is benchmark-specific. "1.5ร— lossless on LIBERO" is the Speed-priority preset with 71.9 vs 74.6 baseline (2.7 pp drop) โ€” not strictly lossless. Authors' "Acc" preset gets 74.9 (no drop) at 1.35ร—.
  • No real-robot evaluation in the paper. v3 reports only LIBERO + SimplerEnv (simulation); the lone real Franka claim lives in the repo README without a results table.
  • Lightweight generator works because intuitive segments are approximately linear; on highly nonlinear / contact-rich intuitive segments (rare in pick-and-place but common in dexterous manipulation), Ridge Regression may break down.

Significance & Positioning

  • vs token-pruning-only methods (Action-aware Dynamic Pruning, FastV, VLA-Cache, EfficientVLA, SparseVLM, FoPru, FastVLM, VisionZip, PruMerge): SP-VLA's spatio-semantic-with-edge pruning prevents the catastrophic failure the paper documents in pure semantic pruners (PruMerge โ†’ 0%, VisionZip โ†’ 33%). Combined with scheduling it gets 2.4ร— SimplerEnv speedup at +0.25 pp accuracy, which no single-axis pruner achieves.
  • vs quantization (AutoQVLA, QAIL): orthogonal โ€” could be stacked.
  • vs parallel decoding (OpenVLA-OFT, PD-VLA, FASTER): these accelerate inside one VLA forward pass; SP-VLA accelerates between forward passes by skipping them. Composable.
  • vs hierarchical dual-system VLAs (ฯ€0.5, Hi-Robot, Helix, OneTwoVLA): dual-system designs route reasoning to a separate model; SP-VLA borrows the System-1 / System-2 metaphor but applies it to action-typing rather than reasoning. The "lightweight generator" is essentially a System-1 motor primitive (ridge regression) for intuitive arm trajectories, while the full VLA acts as System-2 for grasp-class moments.
  • vs early-exit (DeeR-VLA): DeeR exits earlier in the LLM stack on easy frames; SP-VLA skips the LLM entirely on intuitive frames, which is a stronger speedup ceiling.
  • First systematic treatment of temporal ร— spatial redundancy in VLA inference โ€” the action-type indicator + edge-aware pruning ablation reveals previously unappreciated structural facts about VLA computation.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ