ICML 2026 See What Matters - Heungwoo/research GitHub Wiki

See What Matters: Differentiable Grid Sample Pruning for Generalizable VLA — Continuous token resampling, not discarding

Venue: ICML 2026 (Poster) Category: Efficiency Traction (2026-06): 1 citation (arXiv)

Motivation and performance of GridS: dense uniform tokens (2×256) carry heavy redundancy, while GridS resamples to 2×16 salient tokens at 6.25% compute and boosts OOD success (Figure 1 from Wang et al., 2026)

Problem

Vision-Language-Action (VLA) models are bottlenecked by the sheer volume of visual tokens emitted by their vision encoders, making real-time deployment costly. Existing token-pruning methods face a fundamental trade-off: aggressive compression via discrete pruning inevitably discards critical geometric details (contact points, object boundaries), causing severe performance degradation. This forces a compromise that caps the achievable compression rate and the resulting speedup. The authors argue the fix is to rethink compression as a geometry-aware, continuous token resampling operation inside the vision encoder rather than discrete selection.

Method

The Differentiable Grid Sampler (GridS) is a plug-and-play module inserted between the vision encoder and the transformer, jointly optimized with the VLA policy during standard fine-tuning (no auxiliary supervision or ground-truth attention maps). The pipeline has two stages:

  1. Global Coordinate Prediction. The dense feature map is aggregated via global average pooling into a context vector z; a lightweight MLP then predicts K sets of continuous, normalized coordinates P = σ(MLP(z)) ∈ [0,1]^{K×2}, with K ≪ H×W. Predicting continuous coordinates (rather than discrete grid indices) avoids quantization error.
  2. Differentiable Bilinear Sampling with Geometry Injection. Each predicted coordinate is resampled via bilinear interpolation of its four nearest grid neighbors (weights ω₁…ω₄), achieving sub-patch accuracy while remaining fully differentiable for end-to-end training.

The GridS pipeline: global average pooling → MLP coordinate prediction → differentiable bilinear sampling (Figure 3 / x1 from Wang et al., 2026)

Because downstream fine-tuning is already required for VLAs, GridS adds no extra pipeline steps and is compatible with both auto-regressive and flow-matching policies (π₀, π₀.₅, SmolVLA).

Results

On LIBERO (Table 1), π₀ + GridS with only 16 visual tokens (vs. 256 baseline) reaches 96.0% avg SR (+1.6) while cutting FLOPs from 216.0G to 51.65G; with just 4 tokens it still scores 95.5% (+1.1). π₀.₅ + GridS-16 reaches 97.7% (+1.0). This validates the lowest feasible visual-token count reported to date with no degradation, yielding a 76% FLOPs reduction. On the ALOHA benchmark, π₀ + GridS (16 tokens) matches or beats the 256-token baseline. In real-world SO100 experiments (SmolVLA policy, RTX 3090, 3 tasks, 21 OOD scenarios), GridS reduces inference latency by ~2.3s and lifts OOD success rate by +28.6% over the baseline, with the largest gains on the hardest Stack Cubes task. Analysis ("Why GridS Works") shows bilinear sampling produces a diagonal-dominant, near-orthogonal token set, with information-retention dropping gracefully from clean ALOHA (0.9994) to noisy real-world (0.8610) scenes.

Significance

GridS reframes VLA token compression from lossy discrete pruning to learnable continuous resampling, breaking the compression-vs-accuracy trade-off. Operating at <10% of the original token budget while improving OOD robustness, it offers a practical, drop-in path to real-time VLA deployment.

Links

← Back to ICML-2026