ICML 2026 SCALE - Heungwoo/research GitHub Wiki

SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models — training-free, verifier-free, single-pass test-time scaling

Venue: ICML 2026 (Poster) Category: Efficiency Traction (2026-06): 0 citations (arXiv)

Problem

Vision-Language-Action (VLA) models are a promising paradigm for general-purpose robotic control, and test-time scaling (TTS) has drawn attention as a way to improve robustness beyond what training provides. But existing TTS methods for VLAs are impractical to deploy: they require additional training, external verifiers, and multiple forward passes. They also intervene only at action decoding while keeping the visual representation fixed. SCALE argues this is insufficient under perceptual ambiguity, where reconsidering how to perceive is as important as deciding what to do.

Method

SCALE is a simple inference strategy that jointly modulates visual perception and action based on "self-uncertainty," inspired by uncertainty-driven exploration in Active Inference theory. Crucially it requires no additional training, no verifier, and only a single forward pass.

The core idea is uncertainty-conditioned adaptation in both channels:

  • Under high self-uncertainty, SCALE broadens exploration in both perception ("looking") and action ("execution") — exploring more when the model is unsure.
  • Under low uncertainty (confidence), SCALE focuses on exploitation — committing to the perceived state and chosen action.

This yields adaptive execution across varying conditions while preserving single-pass efficiency.

flowchart LR
    O[Observation] --> U{Self-uncertainty<br/>estimate}
    U -- high --> E[Broaden exploration:<br/>perception + action]
    U -- low --> X[Exploit:<br/>commit perception + action]
    E --> A[Single forward pass]
    X --> A
    A --> ACT[Action]
Loading

Results

Experiments on simulated and real-world benchmarks demonstrate that SCALE improves state-of-the-art VLAs and outperforms existing TTS methods while maintaining single-pass efficiency. (Per-benchmark numbers are not available from the abstract-only source; the arXiv v1 full text was not yet rendered at the time of writing.)

Significance

SCALE reframes test-time scaling for VLAs as a question of perception as well as action: rather than re-decoding actions multiple times with a verifier, it conditions a single forward pass on the model's own uncertainty to decide when to explore versus exploit. By dropping the requirements for extra training, verifiers, and multiple passes, it targets the practical deployability gap that has limited TTS adoption in real robotic control, grounding the mechanism in Active Inference's uncertainty-driven exploration.

Links

← Back to ICML-2026

⚠️ **GitHub.com Fallback** ⚠️