ICML 2026 See What Matters - Heungwoo/research GitHub Wiki
See What Matters: Differentiable Grid Sample Pruning for Generalizable VLA — Continuous token resampling, not discarding
Venue: ICML 2026 (Poster) Category: Efficiency Traction (2026-06): 1 citation (arXiv)

Problem
Vision-Language-Action (VLA) models are bottlenecked by the sheer volume of visual tokens emitted by their vision encoders, making real-time deployment costly. Existing token-pruning methods face a fundamental trade-off: aggressive compression via discrete pruning inevitably discards critical geometric details (contact points, object boundaries), causing severe performance degradation. This forces a compromise that caps the achievable compression rate and the resulting speedup. The authors argue the fix is to rethink compression as a geometry-aware, continuous token resampling operation inside the vision encoder rather than discrete selection.
Method
The Differentiable Grid Sampler (GridS) is a plug-and-play module inserted between the vision encoder and the transformer, jointly optimized with the VLA policy during standard fine-tuning (no auxiliary supervision or ground-truth attention maps). The pipeline has two stages:
- Global Coordinate Prediction. The dense feature map is aggregated via global average pooling into a context vector z; a lightweight MLP then predicts K sets of continuous, normalized coordinates
P = σ(MLP(z)) ∈ [0,1]^{K×2}, with K ≪ H×W. Predicting continuous coordinates (rather than discrete grid indices) avoids quantization error. - Differentiable Bilinear Sampling with Geometry Injection. Each predicted coordinate is resampled via bilinear interpolation of its four nearest grid neighbors (weights ω₁…ω₄), achieving sub-patch accuracy while remaining fully differentiable for end-to-end training.

Because downstream fine-tuning is already required for VLAs, GridS adds no extra pipeline steps and is compatible with both auto-regressive and flow-matching policies (π₀, π₀.₅, SmolVLA).
Results
On LIBERO (Table 1), π₀ + GridS with only 16 visual tokens (vs. 256 baseline) reaches 96.0% avg SR (+1.6) while cutting FLOPs from 216.0G to 51.65G; with just 4 tokens it still scores 95.5% (+1.1). π₀.₅ + GridS-16 reaches 97.7% (+1.0). This validates the lowest feasible visual-token count reported to date with no degradation, yielding a 76% FLOPs reduction. On the ALOHA benchmark, π₀ + GridS (16 tokens) matches or beats the 256-token baseline. In real-world SO100 experiments (SmolVLA policy, RTX 3090, 3 tasks, 21 OOD scenarios), GridS reduces inference latency by ~2.3s and lifts OOD success rate by +28.6% over the baseline, with the largest gains on the hardest Stack Cubes task. Analysis ("Why GridS Works") shows bilinear sampling produces a diagonal-dominant, near-orthogonal token set, with information-retention dropping gracefully from clean ALOHA (0.9994) to noisy real-world (0.8610) scenes.
Significance
GridS reframes VLA token compression from lossy discrete pruning to learnable continuous resampling, breaking the compression-vs-accuracy trade-off. Operating at <10% of the original token budget while improving OOD robustness, it offers a practical, drop-in path to real-time VLA deployment.
Links
- arXiv: 2605.11817
- ICML 2026: https://icml.cc/virtual/2026/poster/65010
← Back to ICML-2026