ICML 2026 VLAC - Heungwoo/research GitHub Wiki

VLAC: A Generalist Pair-wise Progress Critic Model for Vision-Language-Action Robots — dense intrinsic rewards from pair-wise progress understanding

Venue: ICML 2026 (Poster) Category: RL for VLA Affiliations: Shanghai AI Laboratory (InternRobotics)

Problem

Vision-Language-Action (VLA) models have substantially improved robotic perception and manipulation, but they still struggle to adapt in dynamic, open-ended, real-world environments. The core gap the authors identify is a missing feedback loop: VLA agents lack a reliable signal of task progress and a corresponding mechanism for self-improvement. Without a dependable way to judge whether a given action moved the robot closer to completing a task, models cannot autonomously refine their behaviour in the wild, where environments shift and demonstrations are scarce. This makes online reinforcement learning on physical robots brittle, since reward functions are typically sparse, hand-engineered, or unavailable.

Method

VLAC is proposed as a unified, generalist vision-language action-critic model. Its central idea is a scalable and generalizable pair-wise progress understanding formulation: rather than scoring a single state in isolation, the model compares two points along a trajectory and predicts the difference in task progress between them. By learning to estimate these pair-wise progress deltas, VLAC converts ordinary visual observations and language goals into a continuous, dense measure of how much closer the robot has come to finishing the task.

This pair-wise critic is trained on diverse data sources so that progress understanding transfers across many tasks and environments. The same model is then folded into a reinforcement learning loop: VLAC autonomously evaluates task progress and emits intrinsic dense rewards, removing the need for externally specified reward functions. Because the critic and the action-generation capability live in a single architecture, the model can both propose robust actions and judge their effect, closing the perceive–act–evaluate loop for real-world learning.

flowchart LR
    A[Observation t_i] --> C[VLAC pair-wise critic]
    B[Observation t_j] --> C
    L[Language goal] --> C
    C --> D[Predicted progress delta]
    D --> E[Dense intrinsic reward]
    E --> F[Real-world RL loop]
    F --> G[Robust VLA actions]
    G --> A

Results

The authors report that VLAC generalizes effectively across diverse tasks and environments. Leveraging its pair-wise progress understanding, the model provides reliable dense rewards that can stand in for hand-crafted reward signals, enabling autonomous task-progress evaluation during reinforcement learning. The abstract frames this generalization and the quality of the dense-reward signal as the primary empirical findings, demonstrating that a single critic can serve heterogeneous real-world manipulation settings rather than being tuned per task.

Significance

VLAC reframes reward design for embodied agents: instead of engineering a reward per task, a generalist critic learns progress and supplies dense intrinsic feedback automatically. By unifying action generation and progress evaluation in one model, it offers a practical path to self-improving VLA robots that can keep learning in dynamic, open-ended environments where external supervision is impractical. This positions pair-wise progress estimation as a reusable building block for real-world reinforcement learning on top of pretrained VLA backbones.

Links

← Back to ICML-2026