CoRL 2026 SAE VLA - Heungwoo/research GitHub Wiki
CoRL 2026 — Sparse Autoencoders Reveal Interpretable & Steerable Features in VLA Models
Venue: CoRL 2026 (Austin, TX, Nov 9–12) · Stanford (Schwager Lab). Paper: arXiv 2603.19183. Representative of: mechanistic interpretability of VLAs — look inside the model, find steerable motion/semantic features, change robot behavior. Companions: VLA Architectures · VLA Attention · CoRL 2026 survey.

1. Problem
VLA models generalize across objects, scenes, and instructions, but when and why they do so is largely a black box. The paper asks whether a VLA's internal representations contain reusable, interpretable structure — and whether that structure can be identified and used to control behavior, rather than only prompted through language.
2. Method
Train Sparse Autoencoders (SAEs) on the hidden-layer activations of a VLA — primarily π₀.₅ (PaliGemma VL backbone + action expert), probing eight layers (PaliGemma and action-expert layers 0/5/11/17). SAEs use a TopK architecture with AuxK auxiliary loss, per-sample normalization and a learned pre-bias, learning a sparse dictionary over activations. Discovered latents correspond to motion primitives (grasp onset, transport phase, pre-grasp alignment, goal-placement approach) and semantic concepts. The authors propose a metric distinguishing general transferable primitives from episode-specific memorizations. Steering takes an SAE decoder column as a direction v and perturbs activations as y′ = y + α·v, broadcast across the sequence; amplification induces the feature's behavior, ablation removes it.
3. Results
- Interpretability: human raters found sampled features interpretable ~79% overall, up to 90% (π₀.₅ LIBERO, PaliGemma layer 5).
- Feature classification: the large majority of features are episode-specific memorizations (e.g. ~97% on π₀.₅ LIBERO, ~89% on DROID), leaving a minority of general transferable primitives.
- Steering (LIBERO sim + real DROID): amplifying general/semantic features induces consistent behavior — e.g. steering a pre-grasp-alignment feature makes the robot hover above targets instead of picking them up; steering a transport feature makes it skip grasps and move straight to the goal. Ablating these features destroys task performance. Steering also drives behaviors that are hard to elicit via language prompts.
4. Why it matters
This is direct mechanistic evidence that VLAs learn reusable internal features linking perception, language, and action across tasks and scenes — not just per-episode lookups. SAE features give a concrete, causal handle for interpreting and steering robot policies, a complement to prompt- or fine-tune-based control.
Limitations (reviewer): most features classify as episode-specific memorization, so the transferable set is small; steering effects are demonstrated qualitatively on a limited feature set; analysis centers on π₀.₅, so generality across VLA families is only partly established.
5. Links
- arXiv 2603.19183
- Survey: CoRL 2026 · Related: VLA Architectures · VLA Attention
← Back to CoRL 2026 survey · Home