ICML 2026 Embodied Interpretability - Heungwoo/research GitHub Wiki

Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models — Interventional attribution to diagnose when VLAs rely on spurious visual cues

Venue: ICML 2026 (Poster) Category: Analysis-Insight Traction (2026-06): 0 citations (arXiv)

Interventional attribution for VLA policies: estimating the causal influence of visual regions on action predictions (Figure 1 from Zhang et al., 2026)

Problem

Vision–Language–Action (VLA) policies frequently fail under distribution shift, which suggests their decisions may rest on spurious visual correlations rather than task-relevant causes. Standard interpretability tools for VLAs — attention scores and token norms — are correlational and may not reflect what actually drives an action. The paper reframes visual–action attribution as an interventional estimation problem: rather than asking which regions a model attends to, it asks which regions causally change the predicted action when intervened upon.

Method

The framework introduces two quantities, grounded in structural causal models and Markov-blanket reasoning:

  • Interventional Significance Score (ISS): an interventional masking procedure that estimates the causal influence of a visual region on the action by comparing a full-information policy against a counterfactual (intervened) policy via a divergence. To isolate the per-step causal effect without compounding trajectory-divergence errors, ISS is evaluated under teacher forcing, conditioning both distributions on the ground-truth action. The authors show ISS admits unbiased estimation and characterize when action prediction error (MSE) is a valid proxy for causal influence (a proxy for KL divergence).
  • Nuisance Mass Ratio (NMR / nmr@k): a scalar measure of how much attribution mass falls on task-irrelevant features. Built on a Causal Spatial Partition of the input into critical vs. nuisance regions and a Regional Mass Ratio, nmr@k captures the density of "important" tokens landing on nuisance regions among the top-k.

Interventions come in two flavors: soft (inject Gaussian noise into nuisance regions to test robustness while preserving causal mechanisms) and hard (textural, geometric, and patch perturbations to test fidelity).

Results

Experiments use π₀.₅ evaluated in the RLBench simulator across 5 random seeds over 41 tasks (seen set S and unseen sets U₁, U₂), 25 trials each.

  • NMR predicts generalization: the Pearson correlation between nmr@k and task success rate is strongest at NMR@10, reaching a maximal negative correlation of −0.77 — higher nuisance attribution corresponds to lower success.
  • ISS is more faithful and robust than baselines: against attention score (ATT) and token norm (NORM), ISS occupies the optimal region — simultaneously maximizing saliency-map cosine similarity (Δ-similarity, higher is better) and minimizing action MSE (Δ-action, lower is better) under perturbation, the best trade-off among the compared explanation methods.

Relationship between nmr@k and success rate across tasks and seeds; NMR@10 gives the strongest negative correlation (Figure 3 from Zhang et al., 2026)

Significance

The work provides a principled, interventional alternative to correlational saliency for embodied policies, with theory (unbiasedness, an MSE-as-KL-proxy justification) backing the estimators. Practically, NMR offers a simple diagnostic for causal misalignment that predicts how a VLA will generalize under distribution shift — a cheap pre-deployment signal for spotting policies that have latched onto spurious visual cues.

Links

← Back to ICML-2026