RSS 2026 Offline Policy Evaluation for Manipulation - Heungwoo/research GitHub Wiki
Offline Policy Evaluation for Manipulation Policies via Discounted Liveness Formulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #154 Authors: Hao Wang, Joshua Bowden, Colton Crosby, Somil Bansal (USC, Stanford) arXiv: 2605.11479 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 shows three case-study episodes (LIBERO pick-and-place with π0, Robomimic square-peg insertion with a diffusion policy, real Franka cloth folding from human teleop) with the learned value on the right: the value trends toward −1 as the task progresses, spikes up at setbacks (bowl slips, grasp slips), and comes back down when the policy recovers — all learned from sparse episode-level labels only.
Problem
Offline evaluation of manipulation policies is hard because rewards are sparse (success/time-out only), rollouts are finite (so failures and truncations are indistinguishable, causing truncation bias that makes standard Bellman-based estimators systematically underestimate performance), and modern policies exhibit recovery behaviors producing non-monotonic task progression.
Method
The paper recasts policy evaluation as a discounted liveness problem: the value of a state is the expected minimum of a target function l(s) (+1 non-goal, −1 goal) along future rollouts, yielding the Bellman operator Ṽ(s) = (1−γ) + γ·min{l(s), Ṽ(s′)}, which is proven to be a contraction and whose fixed point is a deliberately conservative (pessimistic) estimate; values map monotonically to steps-to-completion (1−2γᵏ). A two-stage bootstrapping algorithm then reduces truncation bias: stage 1 computes well-defined values for predecessors of observed goal states; stage 2 bootstraps l(s) with these anchors, so the min operator propagates observed-success values backward through timed-out episodes that overlap successful ones — timed-out episodes are only penalized if they neither reach a goal nor intersect a successful episode. Value networks take SigLIP2 image embeddings plus proprioception and are trained with prioritized experience replay.
Results
Across three case studies — LIBERO-Spatial pick-and-place evaluated on π0 (250-step time-out), Robomimic Square peg insertion with a diffusion policy (200 steps), and real Franka cloth folding from 150 human teleop episodes (300 steps) — the method beats TD(0), Monte Carlo, distributional MC, and the no-bootstrap ablation on the success and composite metrics while staying competitive on the failure metric (200-episode training sets, 50/50 success/time-out, 5 seeds). On cloth folding, the success metric exceeds the next-best baseline by more than 2×. The value function shows emergent predictive ability (a poor grasp raises the value before the slip occurs), and ablations show performance improves with dataset size and holds across SigLIP2/CLIP/DINOv2 encoders.
Significance
A theoretically grounded fix for the truncation-bias problem that plagues offline evaluation of recovery-capable generalist policies — increasingly relevant as VLA developers rely on offline value estimates (cf. π*0.6-style value learning) for deployment decisions. Related wiki threads: RL · Review-Dexterous-Manipulation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home