ICLR 2026 Verifier Free Sampling - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Suhyeok Jang (KAIST), Dongyoung Kim (KAIST / RLWRLD), Changyeon Kim (KAIST); Youngsuk Kim (SNU); Jinwoo Shin (KAIST / RLWRLD) Category: VLA Inference / Test-time scaling Trend tag: Self-confidence / classifier-free-guidance-style sampling
flowchart LR
Obs[Observation o_t, q_t<br/>+ Instruction I] --> VLA[Autoregressive VLA π_θ<br/>e.g. π0-FAST / OpenVLA]
VLA -->|temperature τ=0.5| Cand[N candidates ã^1..ã^N<br/>parallel sampling]
Obs --> Ref["Same VLA with masked inputs<br/>π_θ given (o, ∅, I) — state-mask<br/>π_θ given (o, q, ∅) — text-mask<br/>τ_ref=4.0"]
Cand --> KL["Token-level KL(Q ‖ P)<br/>per-candidate"]
Ref --> KL
KL --> Agg[Aggregate: First-5 tokens<br/>FAST-tokenizer aware]
Agg --> Pick[argmax → a*]
Pick --> Robot[Execute]
Autoregressive VLAs (OpenVLA, π0-FAST) are bottlenecked by single-shot greedy decoding on high-precision tasks like grasping and placement. Existing test-time scaling for VLAs (e.g., RoboMonkey; Nakamoto et al. 2024) trains a separate verifier — costly and brittle on unseen prompts/objects. MG-Select asks whether the VLA can score its own candidates using only its internal distribution.
Framework (Sec. 3.1). Two stages: (1) sample N candidates in parallel at temperature τ > 0, (2) pick a* = argmax_ã C_ã for some confidence C.
Condition-masking distributional confidence (Sec. 3.2). Define token-level confidence C_i = KL(Q_i || P_i) where P_i = π_θ(· | o_t, q_t, I, a_{<i}) is the standard conditional and Q_i is a reference obtained by masking conditions:
- Text-masking:
KL_text = KL(π_θ(· | o_t, q_t, ∅, a_{<i}) || P_i)(Eq. 1) - State-masking:
KL_state = KL(π_θ(· | o_t, ∅, I, a_{<i}) || P_i)(Eq. 2) - Both:
KL_both = KL(π_θ(· | o_t, ∅, ∅, a_{<i}) || P_i)(Eq. 3)
Per-candidate score C_ã = Σ_{i ∈ I} C_i aggregated over a token index set I. The optimal masking variant is task-dependent: state-masking on SIMPLER-WidowX (pure pick-and-place — the model knows what to do without instructions), text-masking on RoboCasa (24 distinct tasks — needs the instruction).
Joint training strategy (Sec. 3.3). Standard fine-tuning produces conditional distributions, so naive condition-masking yields garbage. MG-Select fine-tunes with dropout on conditions: randomly sample one of {(q_t, I), (q_t, ∅), (∅, I), (∅, ∅)} per training step. This teaches the model both conditional and unconditional distributions, sharpening the reference distribution. Trained variant is denoted MG-Select*.
Aggregation choice. Aggregating over all tokens hurts; truncating to the first ~5 tokens (where FAST-tokenized actions encode low-frequency / coarse motion components) works best. Authors hypothesize this aligns with FAST's variable-length low-to-high-frequency token ordering.
Single-prefill deployment. Since prefill is per-step in VLAs, naive Best-of-N pays N× prefill cost. MG-Select shares one prefill across N decodes, reducing latency by 45% at N=4 vs. vanilla MG-Select (Fig. 3).
RoboCasa (Table 1). 24 tasks × 50 trials × 3 seeds; 8 pick-and-place (PnP) tasks reported separately:
| Model | 30-demo PnP | 30-demo All | 100-demo PnP | 100-demo All | 300-demo PnP | 300-demo All |
|---|---|---|---|---|---|---|
| GR00T N1 | 0.4 | 17.4 | 2.2 | 32.1 | 22.6 | 49.6 |
| π0-FAST† | 5.3 | 30.9 | 17.0 | 40.2 | 43.2 | 61.2 |
| + MG-Select | 7.2 | 32.0 | 22.6 | 43.7 | 46.5 | 61.3 |
| + MG-Select* | 14.2 | 34.6 | 31.0 | 48.1 | 46.9 | 62.9 |
At 30 demos PnP: 5.3 → 14.2 is a 168% relative gain. Test-time-only MG-Select (no joint training) is positive but joint training amplifies it substantially in the low-data regime.
SIMPLER-WidowX (Table 2). 24 trials × 4 PnP tasks, π0-FAST trained on BridgeData V2:
| Model | Spoon-on-Towel | Carrot-on-Plate | Stack Cubes | Eggplant-in-Basket | Avg |
|---|---|---|---|---|---|
| RoboVLM | 29.2 | 25.0 | 12.5 | 58.3 | 31.3 |
| SpatialVLA | 16.7 | 25.0 | 29.2 | 100.0 | 42.7 |
| π0-FAST† | 66.7 | 70.8 | 41.7 | 8.3 | 46.9 |
| + MG-Select* | 69.4 | 75.0 | 43.1 | 13.9 | 50.3 |
LIBERO (Table 6). 4 suites × 10 tasks × 50 trials × 3 seeds:
| Model | Spatial | Object | Goal | Long | Avg |
|---|---|---|---|---|---|
| OpenVLA† | 85.2 | 63.7 | 75.5 | 52.5 | 69.2 |
| + MG-Select* | 81.7 | 72.5 | 73.6 | 55.4 | 70.8 |
| π0-FAST† | 97.4 | 95.4 | 95.6 | 79.6 | 92.0 |
| + MG-Select* | 97.2 | 98.0 | 94.5 | 82.7 | 93.1 |
Largest gains on the lowest-base-rate suites (LIBERO-Object for OpenVLA; LIBERO-Long for π0-FAST), confirming the precision-critical hypothesis.
Real-world DROID (Franka Research 3). 60 demos / task (4 objects × 15). ID: 24 trials/task (Table 4). OOD: 16 trials/task (Table 3). Note: ID rows use MG-Select* (joint-trained); OOD rows use plain MG-Select (no joint training).
| Setting | π0-FAST-DROID | + MG-Select | Δ |
|---|---|---|---|
| ID Box-to-Bowl | 41.7 | 58.3 | +16.6 |
| ID Box-to-Plate | 37.5 | 54.2 | +16.7 |
| ID Basket-to-Bowl | 45.8 | 50.0 | +4.2 |
| ID Plate-to-Basket | 25.0 | 29.2 | +4.2 |
| ID average | 37.5 | 47.9 | +28% relative |
| OOD Pick-up-Tape | 56.3 | 68.8 | +12.5 |
| OOD Cup-out-of-Bowl | 50.0 | 75.0 | +25.0 |
| OOD average | 53.1 | 71.9 | +35% relative |
All on RoboCasa, 100 demos, τ_sampling = 0.5 (Table 5):
(a) Inference strategy. Greedy 28.5 / Sampling 27.6 / Uniform-KL (Kang et al. 2025) 30.0 / Likelihood 30.5 / MG-Select 31.0. Likelihood selection alone helps modestly; KL against a non-task-aware uniform helps similarly; condition-masking KL is best.
(b) Number of candidates N. PnP SR rises 27.6 (N=1) → 30.0 (N=2) → 31.0 (N=4) → 30.0 (N=8) → 30.7 (N=16); N=4 is the peak and authors fix it as the cost-effective point. (Table 5b only sweeps N up to 16.)
(c) Masking variant. On RoboCasa (24-task): text-masking 31.0 / state-masking 30.1 / both 29.7. Text-masking wins on multi-task benches because task identity is the salient signal.
(d) Joint training is additive. Joint-IL only: 28.5; MG-Select only: 22.6; both: 31.0.
(e) Reference temperature. Naive τ_ref = 1.0 (the trained masked distribution as-is): only 28.8. The masked distribution is too "peaked"; flattening to τ_ref = 4.0 gives 31.0. Going further to τ_ref = 8.0 drops to 30.0.
(f) Aggregation strategy. Sum 26.1, Avg 24.7, First-5 31.0, First-10 26.6 (the only four entries in Table 5f). Truncating to early FAST tokens is critical — naive summation over all tokens performs worst, as late high-frequency tokens introduce noise into the confidence score.
Inference latency (Fig. 3, LIBERO-Object). Vanilla MG-Select N=4 is ~3× slower than single-action inference; single-prefill MG-Select N=4 has ~45% lower latency and remains comparable to single-action inference across N.
The paper does not contain a Limitations section. From the experiments and design, implicit constraints are:
- Evaluation is dominated by pick-and-place tasks. All sim benches (RoboCasa PnP subset, SIMPLER-WidowX, LIBERO suites) and all real-world tasks are pick-and-place. The "high-precision" claim is anchored to grasping/release moments.
-
Autoregressive VLA only. Method explicitly relies on
π_θproducing token-level categorical distributions over a discrete vocabularyV. Diffusion-based VLAs without explicit token distributions are out of scope. -
Aggregation index
Iis hand-tuned. First-5 is empirically optimal for FAST tokens but is tied to FAST's low-to-high-frequency ordering; other tokenizers may need re-tuning. - τ_ref tuning required. A free hyperparameter (4.0 nominally) that, if set to 1.0, degrades performance noticeably.
- Masking variant is task-suite-dependent. Text vs. state vs. both must be chosen per benchmark.
MG-Select transports classifier-free guidance logic from generative modeling into VLA test-time scaling: train the same model with both conditional and unconditional distributions; score candidates by how much they exploit the conditioning. The most striking result is the 168% relative gain on RoboCasa-30-demos PnP — exactly the low-data, low-precision regime where existing VLAs are weakest.
Compared to:
- RoboMonkey — external verifier, additional training, doesn't generalize to unseen prompts; MG-Select is verifier-free and bench-agnostic.
- Compose Your Policies — composes multiple policies; MG-Select composes one policy against itself.
- Nakamoto et al. 2024 — offline-RL value function as verifier; MG-Select removes this dependency entirely.
- Kang et al. 2025 (LLM self-certainty) — direct inspiration; MG-Select adapts the uniform-reference KL to a task-aware condition-masked reference.
The single-prefill deployment trick is independently valuable for any best-of-N method on autoregressive VLAs.
- OpenReview: https://openreview.net/forum?id=UD4Rw8MOEK
- RoboMonkey — best-of-N with external scoring
- Compose Your Policies
- Survey: VLA & Manipulation
← Back to ICLR-2026