RSS 2026 TAIL Safe - Heungwoo/research GitHub Wiki

TAIL-Safe: Task-Agnostic Safety Monitoring for Imitation Learning Policies

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #207 Authors: Riad Ahmed, Momotaz Begum arXiv: 2605.01195 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

TAIL-Safe pipeline overview (Figure 1 of arXiv 2605.01195, © the authors)

Figure 1. Overview of TAIL-Safe. Top-left: a Gaussian Splatting pipeline builds a digital twin (~20 min: 5 min capture + 15 min reconstruction), aligned to the robot via Umeyama's algorithm. Top-right: the simulator generates safe and unsafe trajectories under perturbations to train WeightNet (score fusion) and Q-ValueNet (success prediction). Middle/bottom: at deployment TAIL-Safe monitors Q(s,a) in real time, staying inactive while Q > 0 and steering the system back to safety via gradient-based recovery as Q approaches zero.

Problem

Imitation-learning policies (flow-matching, diffusion) can fail even within their training distribution due to sensitivity to initial conditions and compounding drift, making field deployment unsafe. Safe deployment requires knowing, for a trained policy, the set of states from which it is guaranteed to complete the learned task.

Method

TAIL-Safe learns a Lipschitz-continuous Q-value function mapping state–action pairs to a safety score built from three task-agnostic criteria — visibility, recognizability, and graspability. The zero-superlevel set of Q defines a Control Invariant Set; when the nominal policy proposes an action outside it, Nagumo's theorem is used to compute a recovery action by gradient ascent on Q, steering back to safety. Training data is gathered from a photorealistic Gaussian-Splatting digital twin (~20 min to build) that generates ~500 rollouts per task (~40% failures) without risking hardware, feeding a WeightNet (score fusion) and Q-ValueNet (success prediction). The monitor runs at 20 Hz, with recovery taking 3–5 iterations (mean 2.3).

Results

On a Franka Emika robot across two tabletop tasks, flow-matching policies without safety succeed only ~20–25% of the time under run-time perturbations, but reach 100% success when guided by TAIL-Safe. Compared to baselines, TAIL-Safe's Q(s,a) attains AUROC 0.999 with 100% recovery success at 2.8 ms per step, versus a Learned CBF (AUROC 0.987 but only 6.9% recovery) and an ensemble (0% recovery, whose mean recovery is actively harmful, ΔQ = −0.86). A WeightNet score fusion is shown to outperform equal weighting, and recovery achieves 100% success with 99.3% state-level accuracy.

Significance

Provides a task-agnostic run-time safety watchdog with formal invariance guarantees for otherwise-brittle IL policies, addressing reliability concerns central to Review-Realtime-Execution and safe deployment of learned manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home