CoRL 2024 ReKep - Heungwoo/research GitHub Wiki
Venue: CoRL 2024 ยท Authors: Wenlong Huang, Chen Wang, Yunzhu Li, Ruohan Zhang, Li Fei-Fei โ Stanford ยท arXiv: 2409.01652 Category: Affordance / Keypoint grounding (training-free) Trend tag: VLM-as-constraint-writer
flowchart LR
RGBD[RGB-D observation<br/>+ language instruction] --> KP[DINOv2 features + SAM masks<br/>propose 3D keypoint candidates<br/>on meaningful regions]
KP --> VLM[GPT-4o<br/>writes ReKep constraints<br/>as Python functions over keypoints]
VLM --> Split[Sub-goal constraints<br/>+ path constraints<br/>per stage]
Split --> Opt[Hierarchical constrained optimization<br/>SDF collision + IK, real-time]
Opt --> Act[SE/3/ end-effector pose sequence<br/>perception-action loop]
Specifying manipulation tasks in a form a robot can optimize against usually requires task-specific reward design, demonstrations, or learned policies โ none of which generalize zero-shot to new objects, multi-stage goals, or new embodiments. ReKep asks whether a large VLM can write the task specification itself, as a structured, optimizable object grounded in the scene, with no task-specific training or environment models.
Representation. A Relational Keypoint Constraint is a Python function that maps a set of 3D keypoints in the scene to a scalar cost (= 0 satisfied, > 0 violated). A task is a sequence of stages; each stage carries two constraint types:
- Sub-goal constraints โ what relational configuration of keypoints must hold at the end of the stage (e.g., spout above cup).
- Path constraints โ what must hold throughout the stage (e.g., keep the cup upright while transporting).
Grounding pipeline (training-free).
- A large vision model proposes keypoint candidates: DINOv2 features cluster fine-grained meaningful regions, with SAM masks restricting candidates to object parts.
- The RGB image overlaid with numbered keypoints plus the instruction is fed to GPT-4o, which emits the staged ReKep constraints as Python programs referencing keypoints by index.
Solving. Given the constraints, a hierarchical constrained optimization solves for robot actions as a sequence of SE(3) end-effector poses: an outer loop picks the next sub-goal, an inner loop solves the path subproblem subject to path constraints, collision (SDF), and reachability (inverse kinematics). It runs in a closed perception-action loop at real-time frequency, so it is reactive to disturbances. No policy weights, reward models, or object pose estimators are trained.
The system runs on two real platforms โ a wheeled single-arm mobile manipulator and a stationary bimanual setup โ and is reported qualitatively (no task-specific data or models). Demonstrated behaviors span multi-stage (pouring tea), in-the-wild (stowing books, taping a box, recycling), bimanual (collaborative folding, shoe packing), and reactive manipulation. ReKep emphasizes the breadth and zero-shot nature of these demonstrations rather than benchmark success-rate tables.
ReKep is a foundational training-free VLM-grounding approach: instead of using a VLM as a policy or planner, it uses GPT-4o as a constraint writer that emits an optimizable, keypoint-relational task spec, with vision foundation models (DINOv2 + SAM) supplying the grounding. This template โ VLM priors as the specification/adaptation layer over an explicit optimizer rather than as the controller โ is the lineage that later affordance and bimanual pipelines build on or react against. VLBiMan explicitly benchmarks against ReKep, arguing demonstration-conditioning avoids ReKep's brittle prompt-engineering and unreliable open-loop grasps; the VLA Architectures review situates this constraint-writer paradigm against end-to-end VLA policies. ReKep's reliance on correct keypoint proposal and on GPT-4o producing valid constraint code is its main fragility.
- arXiv: 2409.01652 ยท PMLR (CoRL 2024): proceedings.mlr.press/v270/huang25g
- Project page: rekep-robot.github.io
- Code: github.com/huangwl18/ReKep
- VLBiMan โ demonstration-conditioned bimanual pipeline that benchmarks against ReKep
- VLA & Manipulation Survey โ situates VLM-grounding among VLA trends
โ Back to Home