RSS 2026 Robometer - Heungwoo/research GitHub Wiki

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #140 Authors: Anthony Liang, Yigit Korkmaz, Jiahui Zhang, Minyoung Hwang, Abrar Anwar, Sidhant Kaushik, Aditya Shah, Alex S. Huang, Luke Zettlemoyer, Dieter Fox, Yu Xiang, Anqi Li, Andreea Bobu, Abhishek Gupta, Stephen Tu, Erdem Biyik, Jesse Zhang arXiv: 2603.02115 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Robometer overview (Figure 1 of arXiv 2603.02115, © the authors)

Figure 1: RBM-1M (left) mixes reward-labeled/successful trajectories with unlabeled failures across 1M episodes and 21 embodiments; training (center) combines augmentations (trajectory rewind, fixed-length subsampling) with a dual objective over paired videos — task progress, trajectory comparisons, and task success; downstream (right), the reward model drives online/offline RL, data filtering and retrieval, and zero-shot failure detection, evaluated on 6 OOD scenes from 3 institutions.

Problem

General-purpose robot reward models are trained to predict absolute frame-level progress from expert demonstrations — labels that are trivial for successes (interpolate 0→1) but ill-defined for the abundant failed and suboptimal trajectories in real robot datasets, which therefore go unused, limiting scalability and generalization.

Method

Robometer adds inter-trajectory preference supervision to intra-trajectory progress supervision. Built on Qwen3-VL-4B-Instruct, it interleaves learned progress tokens within the first video (8 frames per trajectory) and appends a preference token after an optional second video; the composite loss L = L_pref + L_prog + L_succ combines a Bradley-Terry-style pairwise preference classifier, a 10-bin categorical progress head, and per-frame success prediction. Training pairs are curated from RBM-1M — a new dataset of over one million trajectories across 21 embodiments (Open-X, AgiBotWorld, Epic-Kitchens, RH20T, LIBERO, plus failed rollouts) — via three strategies: progress-based comparisons (expert vs unlabeled failure), instruction negatives (wrong-task videos get zero progress), and video rewind for synthetic "undoing" failures.

Results

On the authors' RBM-EVAL-OOD set (976 trajectories, 3 institutions, 6 embodiments, 3 unseen), Robometer tops all baselines: VOC Pearson r of 0.94 (vs e.g. RoboReward-4B 0.80) and Kendall-τa of 0.66 for ordering failed/suboptimal/successful trajectories vs RoboReward-4B's 0.50 and 8B's 0.47 — the paper summarizes this as ~14% average gain in rank correlation and 32% relative improvement in distinguishing successful from suboptimal executions. Downstream, policy learning yields 2.4–4.5x higher success than the best baseline per category: online DSRL improves π0 from 20% to 70% success within 10k timesteps (~45 min, 2.5x over RoboReward), offline RL on SO-101 gains 2.4x, and retrieval-filtered imitation data gives a 4.5x average improvement; it also posts the highest average F1 in zero-shot failure detection.

Significance

Shows that global preference constraints unlock the failure-heavy long tail of robot data for reward learning — a scalable complement to progress-only VLM reward models (VLAC, RoboReward, ReWiND) and a practical bridge to RL fine-tuning of generalist policies discussed in RL and Review-LBM-Cotraining.

← Back to RSS 2026 survey · RSS-2026-Papers · Home