ICML 2026 Embodied Interpretability - Heungwoo/research GitHub Wiki
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models — Interventional attribution to diagnose when VLAs rely on spurious visual cues
Venue: ICML 2026 (Poster) Category: Analysis-Insight Traction (2026-06): 0 citations (arXiv)

Problem
Vision–Language–Action (VLA) policies frequently fail under distribution shift, which suggests their decisions may rest on spurious visual correlations rather than task-relevant causes. Standard interpretability tools for VLAs — attention scores and token norms — are correlational and may not reflect what actually drives an action. The paper reframes visual–action attribution as an interventional estimation problem: rather than asking which regions a model attends to, it asks which regions causally change the predicted action when intervened upon.
Method
The framework introduces two quantities, grounded in structural causal models and Markov-blanket reasoning:
- Interventional Significance Score (ISS): an interventional masking procedure that estimates the causal influence of a visual region on the action by comparing a full-information policy against a counterfactual (intervened) policy via a divergence. To isolate the per-step causal effect without compounding trajectory-divergence errors, ISS is evaluated under teacher forcing, conditioning both distributions on the ground-truth action. The authors show ISS admits unbiased estimation and characterize when action prediction error (MSE) is a valid proxy for causal influence (a proxy for KL divergence).
- Nuisance Mass Ratio (NMR / nmr@k): a scalar measure of how much attribution mass falls on task-irrelevant features. Built on a Causal Spatial Partition of the input into critical vs. nuisance regions and a Regional Mass Ratio, nmr@k captures the density of "important" tokens landing on nuisance regions among the top-k.
Interventions come in two flavors: soft (inject Gaussian noise into nuisance regions to test robustness while preserving causal mechanisms) and hard (textural, geometric, and patch perturbations to test fidelity).
Results
Experiments use π₀.₅ evaluated in the RLBench simulator across 5 random seeds over 41 tasks (seen set S and unseen sets U₁, U₂), 25 trials each.
- NMR predicts generalization: the Pearson correlation between nmr@k and task success rate is strongest at NMR@10, reaching a maximal negative correlation of −0.77 — higher nuisance attribution corresponds to lower success.
- ISS is more faithful and robust than baselines: against attention score (ATT) and token norm (NORM), ISS occupies the optimal region — simultaneously maximizing saliency-map cosine similarity (Δ-similarity, higher is better) and minimizing action MSE (Δ-action, lower is better) under perturbation, the best trade-off among the compared explanation methods.

Significance
The work provides a principled, interventional alternative to correlational saliency for embodied policies, with theory (unbiasedness, an MSE-as-KL-proxy justification) backing the estimators. Practically, NMR offers a simple diagnostic for causal misalignment that predicts how a VLA will generalize under distribution shift — a cheap pre-deployment signal for spotting policies that have latched onto spurious visual cues.
Links
- arXiv: 2605.00321
- ICML 2026: https://icml.cc/virtual/2026/poster/63997
← Back to ICML-2026