RSS 2026 Visual Verification Enables Inference time Steering - Heungwoo/research GitHub Wiki
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 1 · paper #79 Authors: Mingtong Zhang, Dhruv Shah (Princeton University) arXiv: 2606.18247 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: Overview of VERITAS. A pre-trained generalist policy (left) acts as a stochastic generator that samples multiple short-horizon action chunks per decision step; a gradient-free visual verifier scores the candidates and the best-of-N action is executed ("Inference-Time Steering"). Successful verifier-guided rollouts are logged as "Verified Rollouts" and reused to fine-tune the policy ("Policy Improvement"), forming a self-improvement flywheel with minimal human supervision.
Problem
Robot foundation models are trained almost entirely on human-expert demonstrations, so improving them scales linearly with human labor. The paper asks how to improve generalist policies without collecting more human data, by letting the robot practice and learn from its own experience.
Method
The paper proposes VERITAS, a generator–verifier framework. A pre-trained generalist policy is treated as a "generator" that samples diverse candidate action chunks; a gradient-free "visual verifier" (VLM-based or heuristic) scores candidates on task alignment and physical plausibility, and the highest-scoring action is executed via best-of-N selection — steering the policy at inference time with no parameter updates. Verifier-approved rollouts are then logged and used as on-policy supervision to fine-tune the base policy for offline improvement. Experiments use π0-Bridge in SimplerEnv simulation and π0-DROID / π0.5-DROID on a real FR3-DROID platform, benchmarked against a V-GPS-DROID learned-value baseline.
Results
Across simulation and real-world settings (3 policies, 1160 total evaluation episodes), verifier-guided execution improved success rates by an average of 12.6% in simulation and 35% in real-world deployment with no fine-tuning. For offline improvement, VERITAS fine-tuning raised average success by 9.7% over the base π0-Bridge policy across 4 SimplerEnv tasks (960 episodes), with the largest gain on Stack Blocks (31.3% → 59.2%, +27.9%). Post-training on verified autonomous rollouts matched the efficiency of learning from human-expert demonstrations, without any human intervention.
Significance
Positions inference-time verification as a scalable, policy- and verifier-agnostic mechanism for autonomous policy improvement, complementing test-time-scaling and co-training trends in Review-LBM-Cotraining and Review-Human-Video-Transfer.
← Back to RSS 2026 survey · RSS-2026-Papers · Home