ICML 2026 Temporal Difference Calibration in Sequential - Heungwoo/research GitHub Wiki
Temporal Difference Calibration in Sequential Tasks — TD value estimation as a principled calibration mechanism for VLAs
Venue: ICML 2026 (Poster) Category: Analysis-Insight Affiliations: Shelly Francis-Meretzki, Mirco Mutti, Yaniv Romano, Aviv Tamar Traction (2026-06): 0 citations (arXiv)

Problem
Reliable uncertainty quantification matters for VLA models in sequential tasks, yet assessing and improving calibration in this setting is largely unexplored—especially when only partial trajectories are observed. In an episodic task, task-success confidence is produced along the episode, but success is only determined at the very end, creating a temporal mismatch between when confidence is emitted and when ground truth is revealed.
Method
The paper formulates sequential calibration for episodic tasks and introduces a sequential extension of the Brier score. The central theoretical result (Theorem 5.1) shows that for binary outcomes, the risk minimizer of the sequential Brier score coincides with the VLA policy's value function: the optimal future-reward predictor equals the expected terminal reward given history, i.e. accumulated past reward plus the policy's Q-value. This connection bridges uncertainty calibration and reinforcement learning, enabling temporal-difference (TD) value estimation as a principled calibration mechanism over time (the method is referred to as TDQC). Two practical applications follow:
- Early stopping via conformal prediction: treat 1 − f_θ as a failure predictor and raise a failure flag when 1 − f_θ exceeds a time-varying threshold δ_t, calibrated to a significance level α using a held-out set of successful trajectories—stopping low-success trajectories early to avoid damage.
- Test-time guided action search: the learned value predictor (a by-product of TDQC) scores actions sampled from the policy to guide action selection.
The methods are trained with TD losses and compared against predictors trained with binary cross-entropy (BCE), over sequences of features or single-step action probabilities. Four VLAs are evaluated—OpenVLA, UniVLA, π0, and π0-FAST—spanning different action parameterizations (e.g., OpenVLA's discretized token probabilities), plus a Franka real-robot dataset collected with π0-FAST.
Results
The paper gives no single headline number, but reports: across all (VLA model, benchmark) settings, TD-based methods consistently achieve lower sequential Brier scores than BCE-trained predictors (Figure 1, averaged over 21 random seeds). TD calibration improves performance over the state of the art on both simulated and real-robot data, and—contrary to recent findings using other calibration techniques—the VLA's single-step action probabilities, once TD-calibrated, yield competitive uncertainty estimates. As a by-product, using the TDQC value predictor to guide action selection yields a 15% increase in success rate for OpenVLA on the LIBERO benchmark.
Significance
The work draws a clean equivalence between a proper scoring rule (sequential Brier) and the RL value function, giving VLA calibration a principled, RL-grounded objective rather than ad-hoc binary-success predictors. It turns calibration into a usable runtime tool—early failure detection and test-time action guidance—which matters as VLAs move toward safety-critical deployment.
Links
- arXiv: 2604.20472
- ICML 2026: https://icml.cc/virtual/2026/poster/63834
← Back to ICML-2026