CoRL 2026 TOPReward - Heungwoo/research GitHub Wiki

CoRL 2026 — TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

Venue: CoRL 2026 (Austin, TX, Nov 9–12). Paper: arXiv 2602.19313. Representative of: training-free VLM rewards — read internal token probabilities instead of asking the VLM to output a score. Companions: RL for VLA · CoRL 2026 survey.

TOPReward result highlights: token-probability rewards vs. baselines across benchmarks (figure from the authors, arXiv 2602.19313, © the authors)

1. Problem

Robot learning needs dense, instruction-conditioned feedback that separates real task progress from stalled, failed, or partial behavior. Existing rewards rely on manual annotations, task-specific demonstrations, or curated reward models that must be trained. A pretrained Video-Language Model (VLM) already "knows" whether a task looks complete — but asking it to emit a numerical score is unreliable, especially for open-source models. TOPReward asks: can we read that knowledge directly, with zero training?

2. Method

Instead of generating a score, TOPReward reads the VLM's internal token probability. It poses a binary completion query — "the above video shows a robot trajectory that completes the task {INSTRUCTION}; decide whether this is True" — and takes the log-probability of the "True" token as the raw signal: s_t = log p_θ("True" | τ_{1:t}, u).

  • Evaluate on uniformly-spaced trajectory prefixes, then min-max normalize per episode to [0,1].
  • Convert to dense per-step rewards via Δ_t = clip(τ·exp(s_t − s_{t−1}), 0, δ_max).

No reward model, no progress labels, no fine-tuning — purely the frozen VLM's hidden likelihoods.

3. Results

Benchmarks: ManiRewardBench (130 tasks, 497 success / 156 failure episodes, 4 platforms — Franka Emika, SO-100/101, single-arm YAM, bimanual YAM, with subtask-level temporal annotations) and Open X-Embodiment (39 datasets, 780 episodes). Models: Qwen3-VL-8B, Molmo2-8B (open-source), Gemini-2.5-Pro (proprietary). Main metric is Value-Order Correlation (VOC).

  • On Open-X, Qwen3-VL-8B reaches 0.857 VOC vs. the GVL training-free baseline's 0.194 (+0.663); Molmo2-8B 0.417 vs. −0.016. GVL collapses to near-zero on open-source backbones; TOPReward does not.
  • On ManiRewardBench, Qwen3-VL holds 0.94–0.95 VOC across all four datasets. Success-detection ROC-AUC improves too (Qwen3-VL 0.654 vs. GVL 0.519).
  • Policy improvement: advantage-weighted BC (AWR) on six real SO-100 tasks using TOPReward advantages beats plain BC on all six (e.g. "place doll in box" 10/10 vs. 7/10).

Overall: substantially beats prior training-free VLM reward methods, and is competitive with a trained reward-model baseline.

4. Why it matters

It turns any off-the-shelf VLM into a zero-shot dense reward with no reward-model training, and — crucially — works on open-source backbones where prompt-a-score methods fail. That makes instruction-conditioned RL/AWR feedback cheap and portable across robots.

Limitations (reviewer): reward is sensitive to how the instruction is phrased; per-episode min-max normalization blocks absolute cross-trajectory progress comparison; quality is capped by the backbone VLM's video understanding and fine-grained spatial reasoning.

5. Links

← Back to CoRL 2026 survey · Home