ICLR 2026 HVD Offline RL - Heungwoo/research GitHub Wiki

HVD — Hierarchical value-decomposed offline RL for whole-body control

Venue: ICLR 2026 · Authors: Zhilong Zhang, Yunpeng Mei, Xinghao Du, Hongjie Cao, Haonan Wang, Pengyuan Min, Chenyu Wang, Pengfei Chen, Chenbo Xin, Yijie Wang, Wenyu Luo, Yihao Sun, Yidi Wang, Lei Yuan, Gang Wang, Yang Yu (LAMDA, Nanjing University et al.) · OpenReview · Category: Whole-Body Control · Trend tag: Offline RL from suboptimal data

Approach diagram

flowchart LR
  Data[WB-50 dataset<br/>50h teleop + rollouts<br/>imperfect, reward-annotated] --> Label[Reward labeling over<br/>collected trajectories]
  Label --> Sel[Principled offline data<br/>selection over suboptimal data]
  Sel --> HVD[Hierarchical Q-value decomposition<br/>along kinematic structure]
  HVD --> Net[Unified transformer policy<br/>multi-modal, multi-task]
  Net --> WBC[Whole-body control<br/>high-DoF tasks]
Loading

Problem

Scaling imitation learning to high-DoF whole-body robots runs into two coupled difficulties: (1) collected offline data is suboptimal / imperfect, so naive behavior cloning inherits the mistakes, and (2) the high-dimensional action space makes learning and credit assignment hard. Standard offline RL methods, and policy-level decompositions, struggle on both fronts.

Method

HVD (Hierarchical Value-Decomposed offline RL) introduces hierarchy directly into the Q-value function rather than the policy. The action space of whole-body control is decomposed into hierarchical components aligned with the robot's kinematic structure, and value is assessed per component — giving fine-grained, component-specific credit assignment while still maintaining a unified policy network. This reduces the effective learning complexity in high-DoF systems.

To handle imperfect data, HVD performs reward labeling over collected trajectories, providing consistent supervision across suboptimal demonstrations and enabling principled offline data selection. The policy is realized with a transformer-based architecture supporting multi-modal and multi-task learning. The paper also introduces WB-50, a 50-hour dataset of teleoperated and policy-rollout trajectories with reward annotations and natural imperfections.

Results

  • HVD significantly outperforms existing baselines in success rate across complex whole-body tasks.
  • Ablations confirm gains stem not only from the offline RL training paradigm but specifically from the hierarchical value-decomposition structure.

(Specific per-task numbers are reported in the full paper; code at LAMDA-RL/HVD.)

Significance

Most hierarchical control methods decompose the policy; HVD instead decomposes the value function along the kinematic tree, which is what unlocks accurate credit assignment in high-DoF whole-body control while keeping a single unified policy. Combined with reward-labeled offline data selection and the new WB-50 benchmark, it offers a scalable route to learning whole-body skills from imperfect demonstration data.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️