CVPR 2026 Robo Dopamine - Heungwoo/research GitHub Wiki

Robo-Dopamine โ€” General Process Reward Modeling for High-Precision Robotic Manipulation

Venue: CVPR 2026 Category: RL for VLA / Reward modeling Trend tag: Trend 3 Affiliations: Peking University + BAAI + University of Sydney + CAS Institute of Automation (code released under the FlagOpen org)

Approach diagram

flowchart LR
  INIT["initial state image"] --> PRM
  GOAL["goal state image"] --> PRM
  CURR["current state image"] --> PRM
  GRM["General Reward Model (GRM)<br/>multi-view, step-aware"] --> PROG["step-wise progress estimate"]
  PROG --> RWD["dense reward (Policy-Invariant Reward Shaping)"]
  RWD --> RL["Dopamine-RL fine-tune"]
  RL --> POL["improved policy"]
Loading

Problem

RL fine-tuning of VLAs needs a dense reward signal. Designing reward functions per task is brittle; using sparse success/failure signals is sample-inefficient. Existing process-reward models are task-specific and do not transfer.

Method

The method (Dopamine-Reward) learns a general-purpose, step-aware process reward model from multi-view inputs, motivated by two failures of prior PRMs: (1) lack of step-aware understanding plus reliance on single-view perception, and (2) theoretically-unsound reward shaping that induces a "semantic trap" misguiding policy optimization.

  • General Reward Model (GRM) โ€” trained on a 3,400+ hour dataset. It ingests initial-state, goal-state, and before/after current-state images across multiple camera views plus the task text, and predicts relative progress/regress between state pairs. Two components: Step-wise Reward Discretization (structural understanding of manipulation steps) and Multi-Perspective Reward Fusion (overcomes single-view occlusion limits).
  • Dopamine-RL โ€” a policy-learning framework using Policy-Invariant Reward Shaping: a theoretically-sound shaping that supplies dense rewards for efficient self-improvement without altering the optimal policy, thereby avoiding the semantic trap. (Note: "policy-invariant" here refers to optimal-policy-preserving shaping, not merely "scoring states.")

Results

Headline result: after one-shot adaptation of the GRM from a single expert trajectory, Dopamine-RL drives a policy from near-zero to 95% success with only ~150 online rollouts (โ‰ˆ1 hour of real-robot interaction), while retaining strong cross-task generalization. The GRM is reported to reach state-of-the-art accuracy in reward assessment.

Significance

Robo-Dopamine is the first general process-reward model for manipulation โ€” analogous to PRMs in the LLM reasoning literature, but grounded in visual state. It decouples RL fine-tuning from per-task reward engineering, which has been the main blocker for RL-for-VLA at scale. Sister paper: Stage-Aware RL (stage-conditioned reward shaping).

Links

Related pages

โ† Back to CVPR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ