RSS 2026 Betting for Sim to Real Performance Evaluation - Heungwoo/research GitHub Wiki
Betting for Sim-to-Real Performance Evaluation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Datasets and Benchmarks · paper #90 Authors: Yujia Chen, Zaid Mahboob, Bowen Weng arXiv: 2604.24018 · program page
Summary compiled from the arXiv paper (v1; note the preprint reverses the first-author order vs the RSS version); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Four experiment panels: (a) win-rate heatmaps of Kelly-betting estimators vs the Monte Carlo baseline across synthetic "real" distributions, learning rates η, and rounds T; (b) average wealth vs error improvement, empirically supporting the wealth-growth diagnostic (Theorem 3, red region = no-edge null); (c) round-by-round placement-error estimates for the SO-ARM101 pick-and-place study (inset photo); (d) round-by-round win rates for the Unitree G1 sim-to-sim velocity-tracking study.
Problem
Estimating a robot's real-world expected performance (success probability, tracking error, etc.) is expensive when physical trials are scarce and safety-limited. Existing sim-to-real estimation tools work by variance reduction (importance sampling) or bias correction (prediction-powered inference, learned control variates); this paper asks whether a fundamentally different mechanism — sequential betting on real outcomes using simulator-derived side information — can provably beat plain Monte Carlo.
Method
The authors formalize sim-to-real evaluation as a Kelly-style betting game: before each real trial, a "gambler" stakes a fraction b_t of wealth on a prediction informed by accumulated simulator samples; wealth grows/shrinks with the realized payoff, and a bet-weighted estimator replaces the Monte Carlo mean. Theorems 1-2 show the log-optimal (Kelly) strategy approximately coincides with the minimum-MSE estimator (provably outperforming Monte Carlo under stated conditions); Theorem 3 shows wealth growth itself is a statistically valid diagnostic that the betting strategy has an "edge" (no-edge null H0: E[Y_t|F_{t-1}] ≤ 0). A practical approximated-Kelly algorithm aggregates a bank of simulator "experts," with just two hyperparameters (learning rate η, Kelly fraction λ).
Results
Synthetic studies (Beta/truncated-normal/mixture families; 100 seeds; T up to 300): ideal Kelly dominates Monte Carlo across all settings, and the largest simulator bank (Sim_172) generally performs best while an intentionally biased bank (Sim_17_biased) fails — matching theory. Real pick-and-place with an imitation-trained SO-ARM101 (OptiTrack ground truth from 119 trials, 102 mm bucket displacement): purely synthetic distribution banks still beat Monte Carlo estimation of placement error, illustrating that the estimator needs distributional signal, not task-matched simulators. A Unitree G1 sim-to-sim velocity-tracking study (552-trial ground truth, T = 30) shows domain-randomized MuJoCo expert banks overtaking synthetic banks in later rounds.
Significance
A statistically principled, simulator-agnostic alternative to variance-reduction/bias-correction for the evaluation-efficiency problem that plagues real-robot benchmarking — directly relevant to Review-VLA-Evaluation and to certification-style testing debates; code released (Bet4Sim2Real).
← Back to RSS 2026 survey · RSS-2026-Papers · Home