ICLR 2026 Verifier Free Sampling - Heungwoo/research GitHub Wiki

MG-Select — Verifier-free Test-Time Sampling for VLA

Venue: ICLR 2026 Authors: Suhyeok Jang (KAIST), Dongyoung Kim (KAIST / RLWRLD), Changyeon Kim (KAIST); Youngsuk Kim (SNU); Jinwoo Shin (KAIST / RLWRLD) Category: VLA Inference / Test-time scaling Trend tag: Self-confidence / classifier-free-guidance-style sampling

Approach diagram

flowchart LR
  Obs[Observation o_t, q_t<br/>+ Instruction I] --> VLA[Autoregressive VLA π_θ<br/>e.g. π0-FAST / OpenVLA]
  VLA -->|temperature τ=0.5| Cand[N candidates ã^1..ã^N<br/>parallel sampling]
  Obs --> Ref["Same VLA with masked inputs<br/>π_θ given (o, ∅, I) — state-mask<br/>π_θ given (o, q, ∅) — text-mask<br/>τ_ref=4.0"]
  Cand --> KL["Token-level KL(Q ‖ P)<br/>per-candidate"]
  Ref --> KL
  KL --> Agg[Aggregate: First-5 tokens<br/>FAST-tokenizer aware]
  Agg --> Pick[argmax → a*]
  Pick --> Robot[Execute]
Loading

Problem

Autoregressive VLAs (OpenVLA, π0-FAST) are bottlenecked by single-shot greedy decoding on high-precision tasks like grasping and placement. Existing test-time scaling for VLAs (e.g., RoboMonkey; Nakamoto et al. 2024) trains a separate verifier — costly and brittle on unseen prompts/objects. MG-Select asks whether the VLA can score its own candidates using only its internal distribution.

Detailed Method

Framework (Sec. 3.1). Two stages: (1) sample N candidates in parallel at temperature τ > 0, (2) pick a* = argmax_ã C_ã for some confidence C.

Condition-masking distributional confidence (Sec. 3.2). Define token-level confidence C_i = KL(Q_i || P_i) where P_i = π_θ(· | o_t, q_t, I, a_{<i}) is the standard conditional and Q_i is a reference obtained by masking conditions:

  • Text-masking: KL_text = KL(π_θ(· | o_t, q_t, ∅, a_{<i}) || P_i) (Eq. 1)
  • State-masking: KL_state = KL(π_θ(· | o_t, ∅, I, a_{<i}) || P_i) (Eq. 2)
  • Both: KL_both = KL(π_θ(· | o_t, ∅, ∅, a_{<i}) || P_i) (Eq. 3)

Per-candidate score C_ã = Σ_{i ∈ I} C_i aggregated over a token index set I. The optimal masking variant is task-dependent: state-masking on SIMPLER-WidowX (pure pick-and-place — the model knows what to do without instructions), text-masking on RoboCasa (24 distinct tasks — needs the instruction).

Joint training strategy (Sec. 3.3). Standard fine-tuning produces conditional distributions, so naive condition-masking yields garbage. MG-Select fine-tunes with dropout on conditions: randomly sample one of {(q_t, I), (q_t, ∅), (∅, I), (∅, ∅)} per training step. This teaches the model both conditional and unconditional distributions, sharpening the reference distribution. Trained variant is denoted MG-Select*.

Aggregation choice. Aggregating over all tokens hurts; truncating to the first ~5 tokens (where FAST-tokenized actions encode low-frequency / coarse motion components) works best. Authors hypothesize this aligns with FAST's variable-length low-to-high-frequency token ordering.

Single-prefill deployment. Since prefill is per-step in VLAs, naive Best-of-N pays N× prefill cost. MG-Select shares one prefill across N decodes, reducing latency by 45% at N=4 vs. vanilla MG-Select (Fig. 3).

Comprehensive Results

RoboCasa (Table 1). 24 tasks × 50 trials × 3 seeds; 8 pick-and-place (PnP) tasks reported separately:

Model 30-demo PnP 30-demo All 100-demo PnP 100-demo All 300-demo PnP 300-demo All
GR00T N1 0.4 17.4 2.2 32.1 22.6 49.6
π0-FAST† 5.3 30.9 17.0 40.2 43.2 61.2
+ MG-Select 7.2 32.0 22.6 43.7 46.5 61.3
+ MG-Select* 14.2 34.6 31.0 48.1 46.9 62.9

At 30 demos PnP: 5.3 → 14.2 is a 168% relative gain. Test-time-only MG-Select (no joint training) is positive but joint training amplifies it substantially in the low-data regime.

SIMPLER-WidowX (Table 2). 24 trials × 4 PnP tasks, π0-FAST trained on BridgeData V2:

Model Spoon-on-Towel Carrot-on-Plate Stack Cubes Eggplant-in-Basket Avg
RoboVLM 29.2 25.0 12.5 58.3 31.3
SpatialVLA 16.7 25.0 29.2 100.0 42.7
π0-FAST† 66.7 70.8 41.7 8.3 46.9
+ MG-Select* 69.4 75.0 43.1 13.9 50.3

LIBERO (Table 6). 4 suites × 10 tasks × 50 trials × 3 seeds:

Model Spatial Object Goal Long Avg
OpenVLA† 85.2 63.7 75.5 52.5 69.2
+ MG-Select* 81.7 72.5 73.6 55.4 70.8
π0-FAST† 97.4 95.4 95.6 79.6 92.0
+ MG-Select* 97.2 98.0 94.5 82.7 93.1

Largest gains on the lowest-base-rate suites (LIBERO-Object for OpenVLA; LIBERO-Long for π0-FAST), confirming the precision-critical hypothesis.

Real-world DROID (Franka Research 3). 60 demos / task (4 objects × 15). ID: 24 trials/task (Table 4). OOD: 16 trials/task (Table 3). Note: ID rows use MG-Select* (joint-trained); OOD rows use plain MG-Select (no joint training).

Setting π0-FAST-DROID + MG-Select Δ
ID Box-to-Bowl 41.7 58.3 +16.6
ID Box-to-Plate 37.5 54.2 +16.7
ID Basket-to-Bowl 45.8 50.0 +4.2
ID Plate-to-Basket 25.0 29.2 +4.2
ID average 37.5 47.9 +28% relative
OOD Pick-up-Tape 56.3 68.8 +12.5
OOD Cup-out-of-Bowl 50.0 75.0 +25.0
OOD average 53.1 71.9 +35% relative

Ablation Studies

All on RoboCasa, 100 demos, τ_sampling = 0.5 (Table 5):

(a) Inference strategy. Greedy 28.5 / Sampling 27.6 / Uniform-KL (Kang et al. 2025) 30.0 / Likelihood 30.5 / MG-Select 31.0. Likelihood selection alone helps modestly; KL against a non-task-aware uniform helps similarly; condition-masking KL is best.

(b) Number of candidates N. PnP SR rises 27.6 (N=1) → 30.0 (N=2) → 31.0 (N=4) → 30.0 (N=8) → 30.7 (N=16); N=4 is the peak and authors fix it as the cost-effective point. (Table 5b only sweeps N up to 16.)

(c) Masking variant. On RoboCasa (24-task): text-masking 31.0 / state-masking 30.1 / both 29.7. Text-masking wins on multi-task benches because task identity is the salient signal.

(d) Joint training is additive. Joint-IL only: 28.5; MG-Select only: 22.6; both: 31.0.

(e) Reference temperature. Naive τ_ref = 1.0 (the trained masked distribution as-is): only 28.8. The masked distribution is too "peaked"; flattening to τ_ref = 4.0 gives 31.0. Going further to τ_ref = 8.0 drops to 30.0.

(f) Aggregation strategy. Sum 26.1, Avg 24.7, First-5 31.0, First-10 26.6 (the only four entries in Table 5f). Truncating to early FAST tokens is critical — naive summation over all tokens performs worst, as late high-frequency tokens introduce noise into the confidence score.

Inference latency (Fig. 3, LIBERO-Object). Vanilla MG-Select N=4 is ~3× slower than single-action inference; single-prefill MG-Select N=4 has ~45% lower latency and remains comparable to single-action inference across N.

Limitations stated by authors

The paper does not contain a Limitations section. From the experiments and design, implicit constraints are:

  • Evaluation is dominated by pick-and-place tasks. All sim benches (RoboCasa PnP subset, SIMPLER-WidowX, LIBERO suites) and all real-world tasks are pick-and-place. The "high-precision" claim is anchored to grasping/release moments.
  • Autoregressive VLA only. Method explicitly relies on π_θ producing token-level categorical distributions over a discrete vocabulary V. Diffusion-based VLAs without explicit token distributions are out of scope.
  • Aggregation index I is hand-tuned. First-5 is empirically optimal for FAST tokens but is tied to FAST's low-to-high-frequency ordering; other tokenizers may need re-tuning.
  • τ_ref tuning required. A free hyperparameter (4.0 nominally) that, if set to 1.0, degrades performance noticeably.
  • Masking variant is task-suite-dependent. Text vs. state vs. both must be chosen per benchmark.

Significance & Positioning

MG-Select transports classifier-free guidance logic from generative modeling into VLA test-time scaling: train the same model with both conditional and unconditional distributions; score candidates by how much they exploit the conditioning. The most striking result is the 168% relative gain on RoboCasa-30-demos PnP — exactly the low-data, low-precision regime where existing VLAs are weakest.

Compared to:

  • RoboMonkey — external verifier, additional training, doesn't generalize to unseen prompts; MG-Select is verifier-free and bench-agnostic.
  • Compose Your Policies — composes multiple policies; MG-Select composes one policy against itself.
  • Nakamoto et al. 2024 — offline-RL value function as verifier; MG-Select removes this dependency entirely.
  • Kang et al. 2025 (LLM self-certainty) — direct inspiration; MG-Select adapts the uniform-reference KL to a task-aware condition-masked reference.

The single-prefill deployment trick is independently valuable for any best-of-N method on autoregressive VLAs.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️