RSS 2026 Beyond Binary Success - Heungwoo/research GitHub Wiki

Beyond Binary Success: Sample-Efficient and Statistically Rigorous Robot Policy Comparison

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 1 · paper #76 Authors: David Snyder, Apurva Badithela, Nikolai Matni, George J. Pappas, Anirudha Majumdar, Masha Itkina, Haruki Nishimura arXiv: 2603.13616 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

The N-SCORE evaluation context (Figure 1 of arXiv 2603.13616, © the authors)

Figure 1 frames the problem in three panels: (left) policy comparison arises from design choices — action tokenization, data mixture, architecture, vision encoder — that pit policy π_A against π_B; (middle) hardware evaluation runs multi-stage trials under a sequential, any-time stopping procedure; (right) N-SCORE delivers comparisons that go beyond binary metrics (partial credit, continuous progress scores), minimize expected trial count, and carry a statistical validity guarantee P[π0 > π1] ≤ α.

Problem

Real-robot evaluation is usually limited to 10–60 rollouts per policy, and comparisons rarely carry statistical guarantees; the state-of-the-art sequential procedure (STEP) is rigorous but only handles binary success, wasting the information in partial-credit rubrics, episodic reward, or trajectory-smoothness metrics that practitioners increasingly use.

Method

The paper introduces N-SCORE, a sequential test built on safe, anytime-valid inference (SAVI). Evidence is aggregated multiplicatively as a test (super)martingale X_{n+1} = (1 + ξ_n(r_{1,n} − r_{0,n}))·X_n over paired progress scores, with the null rejected once X exceeds 1/α*; Ville's inequality gives Type-1 error control at any stopping time (Theorem 1), so evaluators can stop as soon as evidence suffices. The key technical piece is online optimization of the betting fraction ξ_n using kernel-density-estimation-style nonparametric representations of the reward-difference distribution (a family N-SCORE_k indexed by a bandwidth-like parameter k), which handles discrete partial credit and fully nonparametric continuous metrics in one framework — unlike STEP (binary-only) or the parametric θ-SAVI baseline.

Results

Validation spans over 4,500 hardware rollouts and 2,000 high-fidelity simulation rollouts, including RoboArena/DROID crowd-sourced evaluations and the LBM 1.0 study. On LBM 1.0, partial-credit metrics with sequential testing cut simulation evaluation burden by ~70% (a nearly 1,400-sample reduction versus the 2,000-rollout batch), an over-50% improvement versus sequential binary methods like STEP; on hardware the reduction is ~45% versus batch and 24–30% versus STEP. On synthetic nonparametric data, N-SCORE∞ beats the WSR baseline by ~15% in time-to-decision with ~5 points higher power. Fine-grained progress metrics consistently separate competing policies faster than binary success.

Significance

Gives the field a statistically rigorous way to certify "policy A beats policy B" with far fewer physical rollouts — and a quantified argument for rubric-based scoring over binary success. Directly relevant to the evaluation-methodology thread in Review-VLA-Evaluation and the LBM co-training comparisons discussed in Review-LBM-Cotraining.

← Back to RSS 2026 survey · RSS-2026-Papers · Home