CVPR 2026 Robo Dopamine - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: RL for VLA / Reward modeling Trend tag: Trend 3 Affiliations: Peking University + BAAI + University of Sydney + CAS Institute of Automation (code released under the FlagOpen org)
flowchart LR
INIT["initial state image"] --> PRM
GOAL["goal state image"] --> PRM
CURR["current state image"] --> PRM
GRM["General Reward Model (GRM)<br/>multi-view, step-aware"] --> PROG["step-wise progress estimate"]
PROG --> RWD["dense reward (Policy-Invariant Reward Shaping)"]
RWD --> RL["Dopamine-RL fine-tune"]
RL --> POL["improved policy"]
RL fine-tuning of VLAs needs a dense reward signal. Designing reward functions per task is brittle; using sparse success/failure signals is sample-inefficient. Existing process-reward models are task-specific and do not transfer.
The method (Dopamine-Reward) learns a general-purpose, step-aware process reward model from multi-view inputs, motivated by two failures of prior PRMs: (1) lack of step-aware understanding plus reliance on single-view perception, and (2) theoretically-unsound reward shaping that induces a "semantic trap" misguiding policy optimization.
- General Reward Model (GRM) โ trained on a 3,400+ hour dataset. It ingests initial-state, goal-state, and before/after current-state images across multiple camera views plus the task text, and predicts relative progress/regress between state pairs. Two components: Step-wise Reward Discretization (structural understanding of manipulation steps) and Multi-Perspective Reward Fusion (overcomes single-view occlusion limits).
- Dopamine-RL โ a policy-learning framework using Policy-Invariant Reward Shaping: a theoretically-sound shaping that supplies dense rewards for efficient self-improvement without altering the optimal policy, thereby avoiding the semantic trap. (Note: "policy-invariant" here refers to optimal-policy-preserving shaping, not merely "scoring states.")
Headline result: after one-shot adaptation of the GRM from a single expert trajectory, Dopamine-RL drives a policy from near-zero to 95% success with only ~150 online rollouts (โ1 hour of real-robot interaction), while retaining strong cross-task generalization. The GRM is reported to reach state-of-the-art accuracy in reward assessment.
Robo-Dopamine is the first general process-reward model for manipulation โ analogous to PRMs in the LLM reasoning literature, but grounded in visual state. It decouples RL fine-tuning from per-task reward engineering, which has been the main blocker for RL-for-VLA at scale. Sister paper: Stage-Aware RL (stage-conditioned reward shaping).
- arXiv: 2512.23703
- Code:
FlagOpen/Robo-Dopamine
โ Back to CVPR-2026