CoRL 2025 RoboMonkey - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 · Authors: Kwok, Agia, Sinha, Foutter, Li, Stoica, Mirhoseini, Pavone — Stanford / UC Berkeley / NVIDIA · arXiv: 2506.17811 Category: Inference Recipe Trend tag: Test-time scaling
flowchart LR
O[Observation] --> VLA[VLA base policy]
VLA -- N samples --> P[Gaussian perturbation<br/>+ majority voting]
P --> A1[Proposal 1]
P --> A2[Proposal 2]
P --> AN[Proposal K]
A1 & A2 & AN --> V[VLM verifier<br/>Bradley-Terry trained]
V -- best --> EXEC[Execute]
VLAs spend all their compute in one forward pass, deterministically. LLMs at the same scale routinely use best-of-N sampling with a reward model to gain quality at inference time — this idea hadn't transferred to VLAs.
RoboMonkey samples N action candidates from a base VLA (OpenVLA in all experiments), applies Gaussian perturbation + majority voting to build an action-proposal distribution, then scores proposals with a VLM-based verifier and executes the highest-scoring one. The verifier is trained on a synthetic data pipeline: candidate actions are clustered to K representatives and pairwise preferences are assigned by RMSE to the ground-truth action, with the reward model fit using a Bradley-Terry loss (modified for preference levels). Action error vs. number of samples follows an approximate power law (log e ≈ log a + b·log k), giving a controllable inference-time scaling lever. Serving uses an SGLang extension for OpenVLA, reported at ~41.3% lower latency than naive policy sampling.
- Real-world OOD (WidowX): 60% success vs. 35% for OpenVLA and 30% for V-GPS — a 25-point absolute gain.
- SIMPLER (in-distribution): 46.3% average success, +7.8% over OpenVLA.
- LIBERO-Long (fine-tuning): co-fine-tuning the VLA and verifier adds ~7% over fine-tuning the VLA alone.
Demonstrates that test-time scaling translates to manipulation — not just language.
Brings LLM inference-scaling lore (best-of-N, verifier-guided decoding) to robotics. Together with Streaming Flow Policy and DemoSpeedup, RoboMonkey establishes CoRL 2025's trend of cheap inference-time levers that work without retraining the base model.
← Back to CoRL-2025