ICML 2026 ReLAM - Heungwoo/research GitHub Wiki
ReLAM โ Anticipated keypoint subgoals that auto-generate dense rewards for visual manipulation RL
Venue: ICML 2026 (Poster) Category: RL for VLA Traction (2026-06): 0 citations (arXiv)

Problem
Reward design is a persistent bottleneck for visual reinforcement learning in robotic manipulation. In simulation, rewards are conventionally a function of the distance to a target position, but such precise positional information is usually unavailable in real-world visual settings due to sensory and perceptual limitations. ReLAM's goal is to recover an implicit, dense, geometry-aligned reward signal directly from action-free video demonstrations, without ground-truth state.
Method
ReLAM (Reward Learning with Anticipation Model) implicitly infers spatial distances through keypoints extracted from images, and runs in two stages.
- Stage 1 โ Anticipation model over keypoints. Rather than generating images, ReLAM learns a keypoint-generation model (following ATM, Wen et al. 2024). For each demo video it takes the first frame, applies a grounded SAM model (Zhang et al. 2024) to obtain task-relevant segmentations, then uses a point-tracking model (Karaev et al. 2024) to follow each object pixel's trajectory across the video. Pixels with negligible motion are removed by a threshold, and the final representative keypoints are chosen by Farthest Point Sampling (FPS, Eq. 1). The anticipation model takes the current task state and the desired final goal and proposes an intermediate, easy-to-reach keypoint-based subgoal โ a structured curriculum aligned with the task's geometric objective.
- Stage 2 โ Hierarchical RL. A continuous, distance-based reward derived from the anticipated subgoals trains a low-level, goal-conditioned policy under a hierarchical RL framework. The combined reward is
r = r_dense + r_success + I(l_s โค ฮธ_s)(Eq. 4). The paper proves a sub-optimality bound (Eq. 10): the value gap from optimal is bounded bykยท(ฮต_ฯ + 2ฮต_A/M), linking policy and anticipation-model errors.

Results
Evaluated on Meta-World (online RL, Sawyer arm) and ManiSkill (offline RL via IQL, Panda arm), with 256ร256 third-person images and 100 action-free demo trajectories per task. Baselines: Diffusion Reward (DR), DACfO, Image Subgoal (IS), and an unachievable ground-truth Oracle. Final success rates (%): Drawer Open โ DR 80.0 / DACfO 71.3 / IS 4.0 / Oracle 93.3 / ReLAM 100.0; Door Open โ 100.0 / 70.0 / 24.7 / 100.0 / 100.0; Button Press Wall โ 60.0 / 47.8 / 12.7 / 76.7 / 75.8; Push Cube โ 69.3 / 76.7 / 46.7 / 92.7 / 89.3; Pick Cube โ 68.0 / 78.7 / 60.7 / 90.7 / 88.0. ReLAM matches or beats all real-world-achievable baselines and approaches the Oracle, while reaching high success rates with far fewer interaction steps.
Significance
ReLAM shows that abstract, trackable keypoints can stand in for the privileged positional state that visual RL normally lacks, converting cheap action-free video into structured dense rewards with a provable sub-optimality guarantee. This is a practical recipe for long-horizon manipulation where reward engineering and ground-truth state are the main obstacles.
Links
- arXiv: 2509.22402
- ICML 2026: https://icml.cc/virtual/2026/poster/66341
โ Back to ICML-2026