ICML 2026 ReLAM - Heungwoo/research GitHub Wiki

ReLAM โ€” Anticipated keypoint subgoals that auto-generate dense rewards for visual manipulation RL

Venue: ICML 2026 (Poster) Category: RL for VLA Traction (2026-06): 0 citations (arXiv)

ReLAM generates keypoint subgoals with an anticipation model and turns them into rewards for a goal-conditioned policy (Figure 1 from Anonymous et al., 2026)

Problem

Reward design is a persistent bottleneck for visual reinforcement learning in robotic manipulation. In simulation, rewards are conventionally a function of the distance to a target position, but such precise positional information is usually unavailable in real-world visual settings due to sensory and perceptual limitations. ReLAM's goal is to recover an implicit, dense, geometry-aligned reward signal directly from action-free video demonstrations, without ground-truth state.

Method

ReLAM (Reward Learning with Anticipation Model) implicitly infers spatial distances through keypoints extracted from images, and runs in two stages.

  • Stage 1 โ€” Anticipation model over keypoints. Rather than generating images, ReLAM learns a keypoint-generation model (following ATM, Wen et al. 2024). For each demo video it takes the first frame, applies a grounded SAM model (Zhang et al. 2024) to obtain task-relevant segmentations, then uses a point-tracking model (Karaev et al. 2024) to follow each object pixel's trajectory across the video. Pixels with negligible motion are removed by a threshold, and the final representative keypoints are chosen by Farthest Point Sampling (FPS, Eq. 1). The anticipation model takes the current task state and the desired final goal and proposes an intermediate, easy-to-reach keypoint-based subgoal โ€” a structured curriculum aligned with the task's geometric objective.
  • Stage 2 โ€” Hierarchical RL. A continuous, distance-based reward derived from the anticipated subgoals trains a low-level, goal-conditioned policy under a hierarchical RL framework. The combined reward is r = r_dense + r_success + I(l_s โ‰ค ฮธ_s) (Eq. 4). The paper proves a sub-optimality bound (Eq. 10): the value gap from optimal is bounded by kยท(ฮต_ฯ€ + 2ฮต_A/M), linking policy and anticipation-model errors.

ReLAM training framework: keypoint selection, subgoal generation, and reward computation (Figure 2 from Anonymous et al., 2026)

Results

Evaluated on Meta-World (online RL, Sawyer arm) and ManiSkill (offline RL via IQL, Panda arm), with 256ร—256 third-person images and 100 action-free demo trajectories per task. Baselines: Diffusion Reward (DR), DACfO, Image Subgoal (IS), and an unachievable ground-truth Oracle. Final success rates (%): Drawer Open โ€” DR 80.0 / DACfO 71.3 / IS 4.0 / Oracle 93.3 / ReLAM 100.0; Door Open โ€” 100.0 / 70.0 / 24.7 / 100.0 / 100.0; Button Press Wall โ€” 60.0 / 47.8 / 12.7 / 76.7 / 75.8; Push Cube โ€” 69.3 / 76.7 / 46.7 / 92.7 / 89.3; Pick Cube โ€” 68.0 / 78.7 / 60.7 / 90.7 / 88.0. ReLAM matches or beats all real-world-achievable baselines and approaches the Oracle, while reaching high success rates with far fewer interaction steps.

Significance

ReLAM shows that abstract, trackable keypoints can stand in for the privileged positional state that visual RL normally lacks, converting cheap action-free video into structured dense rewards with a provable sub-optimality guarantee. This is a practical recipe for long-horizon manipulation where reward engineering and ground-truth state are the main obstacles.

Links

โ† Back to ICML-2026