CoRL 2025 RoboMonkey - Heungwoo/research GitHub Wiki

RoboMonkey — Test-Time Scaling for VLAs

Venue: CoRL 2025 · Authors: Kwok, Agia, Sinha, Foutter, Li, Stoica, Mirhoseini, Pavone — Stanford / UC Berkeley / NVIDIA · arXiv: 2506.17811 Category: Inference Recipe Trend tag: Test-time scaling

Approach diagram

flowchart LR
  O[Observation] --> VLA[VLA base policy]
  VLA -- N samples --> P[Gaussian perturbation<br/>+ majority voting]
  P --> A1[Proposal 1]
  P --> A2[Proposal 2]
  P --> AN[Proposal K]
  A1 & A2 & AN --> V[VLM verifier<br/>Bradley-Terry trained]
  V -- best --> EXEC[Execute]
Loading

Problem

VLAs spend all their compute in one forward pass, deterministically. LLMs at the same scale routinely use best-of-N sampling with a reward model to gain quality at inference time — this idea hadn't transferred to VLAs.

Method

RoboMonkey samples N action candidates from a base VLA (OpenVLA in all experiments), applies Gaussian perturbation + majority voting to build an action-proposal distribution, then scores proposals with a VLM-based verifier and executes the highest-scoring one. The verifier is trained on a synthetic data pipeline: candidate actions are clustered to K representatives and pairwise preferences are assigned by RMSE to the ground-truth action, with the reward model fit using a Bradley-Terry loss (modified for preference levels). Action error vs. number of samples follows an approximate power law (log e ≈ log a + b·log k), giving a controllable inference-time scaling lever. Serving uses an SGLang extension for OpenVLA, reported at ~41.3% lower latency than naive policy sampling.

Results

  • Real-world OOD (WidowX): 60% success vs. 35% for OpenVLA and 30% for V-GPS — a 25-point absolute gain.
  • SIMPLER (in-distribution): 46.3% average success, +7.8% over OpenVLA.
  • LIBERO-Long (fine-tuning): co-fine-tuning the VLA and verifier adds ~7% over fine-tuning the VLA alone.

Demonstrates that test-time scaling translates to manipulation — not just language.

Significance

Brings LLM inference-scaling lore (best-of-N, verifier-guided decoding) to robotics. Together with Streaming Flow Policy and DemoSpeedup, RoboMonkey establishes CoRL 2025's trend of cheap inference-time levers that work without retraining the base model.

Links

Related pages

← Back to CoRL-2025

⚠️ **GitHub.com Fallback** ⚠️