CoRL 2024 ReKep - Heungwoo/research GitHub Wiki

ReKep โ€” Relational Keypoint Constraints for Manipulation

Venue: CoRL 2024 ยท Authors: Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, Li Fei-Fei โ€” Stanford ยท arXiv: 2409.01652 Category: Affordance / Keypoint grounding (training-free) Trend tag: VLM-as-constraint-writer

Approach diagram

flowchart LR
  RGBD[RGB-D observation<br/>+ language instruction] --> KP[DINOv2 features + SAM masks<br/>propose 3D keypoint candidates<br/>on meaningful regions]
  KP --> VLM[GPT-4o<br/>writes ReKep constraints<br/>as Python functions over keypoints]
  VLM --> Split[Sub-goal constraints<br/>+ path constraints<br/>per stage]
  Split --> Opt[Hierarchical constrained optimization<br/>SDF collision + IK, real-time]
  Opt --> Act[SE/3/ end-effector pose sequence<br/>perception-action loop]
Loading

Problem

Specifying manipulation tasks in a form a robot can optimize against usually requires task-specific reward design, demonstrations, or learned policies โ€” none of which generalize zero-shot to new objects, multi-stage goals, or new embodiments. ReKep asks whether a large VLM can write the task specification itself, as a structured, optimizable object grounded in the scene, with no task-specific training or environment models.

Method

Representation. A Relational Keypoint Constraint is a Python function that maps a set of 3D keypoints in the scene to a scalar cost (= 0 satisfied, > 0 violated). A task is a sequence of stages; each stage carries two constraint types:

  • Sub-goal constraints โ€” what relational configuration of keypoints must hold at the end of the stage (e.g., spout above cup).
  • Path constraints โ€” what must hold throughout the stage (e.g., keep the cup upright while transporting).

Grounding pipeline (training-free).

  1. A large vision model proposes keypoint candidates: DINOv2 features cluster fine-grained meaningful regions, with SAM masks restricting candidates to object parts.
  2. The RGB image overlaid with numbered keypoints plus the instruction is fed to GPT-4o, which emits the staged ReKep constraints as Python programs referencing keypoints by index.

Solving. Given the constraints, a hierarchical constrained optimization solves for robot actions as a sequence of SE(3) end-effector poses: an outer loop picks the next sub-goal, an inner loop solves the path subproblem subject to path constraints, collision (SDF), and reachability (inverse kinematics). It runs in a closed perception-action loop at real-time frequency, so it is reactive to disturbances. No policy weights, reward models, or object pose estimators are trained.

Results

The system runs on two real platforms โ€” a wheeled single-arm mobile manipulator and a stationary bimanual setup โ€” and is reported qualitatively (no task-specific data or models). Demonstrated behaviors span multi-stage (pouring tea), in-the-wild (stowing books, taping a box, recycling), bimanual (collaborative folding, shoe packing), and reactive manipulation. ReKep emphasizes the breadth and zero-shot nature of these demonstrations rather than benchmark success-rate tables.

Significance

ReKep is a foundational training-free VLM-grounding approach: instead of using a VLM as a policy or planner, it uses GPT-4o as a constraint writer that emits an optimizable, keypoint-relational task spec, with vision foundation models (DINOv2 + SAM) supplying the grounding. This template โ€” VLM priors as the specification/adaptation layer over an explicit optimizer rather than as the controller โ€” is the lineage that later affordance and bimanual pipelines build on or react against. VLBiMan explicitly benchmarks against ReKep, arguing demonstration-conditioning avoids ReKep's brittle prompt-engineering and unreliable open-loop grasps; the VLA Architectures review situates this constraint-writer paradigm against end-to-end VLA policies. ReKep's reliance on correct keypoint proposal and on GPT-4o producing valid constraint code is its main fragility.

Links

Related pages

  • VLBiMan โ€” demonstration-conditioned bimanual pipeline that benchmarks against ReKep
  • VLA & Manipulation Survey โ€” situates VLM-grounding among VLA trends

โ† Back to Home

โš ๏ธ **GitHub.com Fallback** โš ๏ธ