ICLR 2026 Align Then Steer - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Affiliation: China Telecom Institute of AI Β· Tsinghua Β· CUHK Shenzhen Β· Northwestern Polytechnical Category: VLA Training β Adaptation / fine-tuning for diffusion- and flow-based VLAs Trend tag: Cross-embodiment / data efficiency / classifier guidance
flowchart LR
subgraph S1[Stage 1: Unified Action Latent Space]
PT[Pre-training action chunks<br/>OXE / DROID / ALOHA / Kuka] --> V1[InfoVAE V_phi<br/>Transformer enc/dec<br/>latent dim 512]
V1 -. "KL to N(0,I)" .-> Prior["Pretrain latent prior q_phi(z)"]
Adapt[Adaptation action chunks<br/>RoboTwin / ManiSkill / RealMan] --> V2[InfoVAE V_psi<br/>same architecture]
V2 -. "reverse KL to q_phi(z)" .-> Mode[Mode-seeking embed<br/>into one mode of q_phi]
end
subgraph S2[Stage 2: Classifier Guidance Steers Fine-tuning]
Mode --> E[Frozen E_psi encoder]
Noisy[a^k_t:t+h noisy chunk] --> E
Clean[a^0_t:t+h clean chunk] --> E
E --> Dist["βE_psi(a^k) β E_psi(a^0)βΒ²"]
Dist --> Grad["g = -β_a βz_hat β zβΒ²"]
Grad --> Diff["Diffusion-VLA loss + β(1βαΎ±_k)·λ·g"]
Grad --> Flow["Flow-VLA loss + ((1βΟ)/Ο)·λ·g"]
end
Diff --> POL[Adapted RDT / DP]
Flow --> POL2[Adapted Οβ]
Pre-trained VLAs (Οβ, RDT-1B, Diffusion Policy) face a distribution mismatch when fine-tuned on a new embodiment or task: degrees-of-freedom, end-effector representation, joint torques and chunk length all change between pre-training data (typically single-arm 6-DoF from Open X-Embodiment / DROID) and adaptation data (e.g., dual-arm 7-DoF RealMan with single-DoF grippers). Direct fine-tuning either overfits the small target set or drifts away from the pre-training visuomotor prior. Methods that operate at the parameter level (LoRA, dynamic layer activation) or use kinematics retargeting do not directly close the action distribution gap, and diffusion/flow-based VLAs are especially sensitive because their output is a distribution over action chunks, not a point estimate.
The authors instantiate two InfoVAEs (Zhao et al. 2017), not vanilla VAEs, because vanilla VAEs suffer from posterior collapse and inaccurate amortized inference. Each VAE has a Transformer encoder/decoder. Two learnable tokens ΞΌ_token and Ξ£_token are prepended to the encoder input (motion-VAE style, Γ la Petrovich et al. 2021), and the decoder uses H zero embeddings as queries with the latent z as key/value via cross-attention.
Pre-training VAE V_phi (encoder E_phi, decoder D_phi):
-
Trained on the cross-embodiment pre-training corpus with the standard InfoVAE objective:
L(Ο) = E_q[log p_Ο(a|z)] β (1βΞ±)Β·D_KL(q_Ο(z|a) β p(z)) β (Ξ±βΞ»β1)Β·D_KL(q_Ο(z) β p(z))where p(z) = N(0, I), and the second KL term is replaced by MMD for tractable optimisation (per Zhao et al. 2017). Uses the same chunk length H as the pre-training VLA.
Adaptation VAE V_psi (encoder E_psi, decoder D_psi):
-
Same architecture and variant. Trained on the small adaptation set with a reverse KL to the learned pre-training latent distribution q_Ο(z):
L(Ο) = E_q[log p_Ο(Γ£|z)] β (1βΞ±)Β·D_KL(q_Ο(z|Γ£) β q_Ο(z)) β (Ξ±βΞ»β1)Β·D_KL(q_Ο(z) β q_Ο(z)) -
q_Ο(z) is approximated as N(ΞΌ_Ο, Ξ£_Ο) with statistics aggregated over all pre-training action latents. The mode-seeking property of reverse KL is the key: it pushes q_Ο into one mode of q_Ο, producing a unified, hierarchically structured latent space.
-
Uses the adaptation chunk length L (which may differ from H).
t-SNE visualisations in Fig. 15 confirm that InfoVAE produces a compact latent space where adaptation embeddings (blue) sit inside pre-training modes (green); vanilla VAE produces poorly aligned, gapped distributions.
A guidance function is constructed in the unified latent Z. Following Carvalho et al. 2023, the classifier is an energy-based model:
p_Ο(y | Γ’^k_t:t+h) = (1/Z_Ο) Β· exp(ββE_Ο(Γ’^k_t:t+h) β E_Ο(a^0_t:t+h)βΒ²)
with guidance gradient:
g = β_Γ’ log p_Ο(y|Γ’) β ββ_Γ’ βE_Ο(Γ’^k) β E_Ο(a^0)βΒ²
This gradient is folded into the training objective of the underlying generative head:
-
Diffusion-based VLAs (RDT, DP):
L(ΞΈ) = E[βΞ΅ β Ξ΅_ΞΈ(a^k; k, o_t, l) + β(1 β αΎ±_k)·λ·gβΒ²] -
Flow-based VLAs (Οβ):
L(ΞΈ) = E[βv_ΞΈ(a^Ο; Ο, o_t, l) + ((1βΟ)/Ο)·λ·g β (a^0 β Ξ΅)βΒ²]
The encoder E_psi is frozen during this stage; it serves purely as a distance metric. Crucially, since guidance enters at training time only, inference cost is zero β unlike test-time gradient-based steering (DynaGuide) that runs iterative backprop per step. The recommended guidance scale is Ξ» = 2 (Table 13).
| Component | RDT-1B | Οβ | Diffusion Policy |
|---|---|---|---|
| InfoVAE latent dim | 512 | 512 | 512 |
| Temporal input length | 64 | 50 | 14 |
| Input dim | 128 | 32 | 14 |
| Optimizer | Adam | Adam | Adam |
| LR (VAE) | 1e-4 | 1e-4 | 1e-4 |
| Weight decay | 1e-4 | 1e-4 | 1e-4 |
| Batch size (VAE) | 64 | 64 | 64 |
| Chunk size (adaptation) | 64 | 50 | 8 |
| Batch size (fine-tune) | 64 | 24 | 128 |
| LR (fine-tune) | 1e-4 | 2.5e-5 | 1e-4 |
| Training duration | 100k steps | 60k steps | 300 epochs |
| Step-1 VAE wall-clock | ~12 h | ~12 h | ~12 h |
| Step-2 VAE wall-clock | < 0.5 h | < 0.5 h | < 0.5 h |
| Compute | 4Γ A100 | 4Γ A100 | 1Γ A100 |
Real-robot fine-tuning runs 120k steps on 8Γ A100 with batch 6/GPU at 20 Hz inference on a single 4090.
| Task | RDT-1B | RDT + ATE (Ξ) | Οβ | Οβ + ATE (Ξ) |
|---|---|---|---|---|
| Block Hammer Beat | 52 | 71 (+19) | 38 | 44 (+6) |
| Block Handover | 69 | 91 (+22) | 80 | 92 (+12) |
| Blocks Stack (Easy) | 10 | 31 (+21) | 30 | 50 (+20) |
| Blocks Stack (Hard) | 1 | 7 (+6) | 8 | 7 (β1) |
| Bottle Adjust | 53 | 37 (β16) | 39 | 45 (+6) |
| Container Place | 34 | 55 (+21) | 56 | 59 (+3) |
| Diverse Bottles Pick | 18 | 24 (+6) | 20 | 40 (+20) |
| Dual Bottles Pick (Easy) | 76 | 87 (+11) | 48 | 85 (+37) |
| Dual Bottles Pick (Hard) | 39 | 58 (+19) | 52 | 55 (+3) |
| Dual Shoes Place | 6 | 9 (+3) | 22 | 24 (+2) |
| Empty Cup Place | 22 | 61 (+39) | 32 | 36 (+4) |
| Mug Hanging (Easy) | 6 | 6 (0) | 11 | 27 (+16) |
| Mug Hanging (Hard) | 1 | 1 (0) | 4 | 4 (0) |
| Pick Apple Messy | 35 | 39 (+4) | 18 | 11 (β7) |
| Put Apple Cabinet | 20 | 45 (+25) | 34 | 55 (+21) |
| Shoe Place | 44 | 43 (β1) | 52 | 58 (+6) |
| Tool Adjust | 54 | 42 (β12) | 70 | 69 (β1) |
| Average | 31.8 | 41.6 (+9.8) | 36.1 | 44.8 (+8.7) |
Sample-efficiency: RDT + ATE surpasses RDT baseline at 70k steps vs. 90k for the baseline. ATE on Οβ at 25 demos/task averages 29.0% vs. baseline 9.2% (3.15Γ gain, Table 12).
| Task | RDT-1B | RDT + ATE |
|---|---|---|
| Push Cube | 65.2 | 78.4 (+13.2) |
| Pick Cube | 7.6 | 14.8 (+7.2) |
| Average | 36.4 | 46.6 (+10.2) |
| Method | Success rate |
|---|---|
| Fine-tuned Οβ | 0.78 |
| + Offline DSRL | 0.43 |
| + ATE (Ours) | 0.88 |
Four long-horizon tasks at 60k / 90k / 120k steps. Tasks include Cook Bun, Pick Bun, Make Sandwich, Use Toaster. Reported in Fig. 3:
- Cook Bun: Οβ ~15% vs. ATE 100% at 90k steps.
- Pick Bun (single-arm): ATE 50% @ 90k β 70% @ 120k.
- Overall average at 120k: Οβ 16.7% vs. ATE 58.1%.
Make Yogurt Bowl (multi-tool dual-arm, 80 demos): Οβ 15% β ATE 25% (Table 2). Under visual distractors: Οβ 0% β ATE 20% (Table 3).
| Setting | Οβ (Cook / Pick / Sandwich) | Οβ + ATE |
|---|---|---|
| Low illumination (25 lux) | 0 / 0 / 40 | 60 / 40 / 40 |
| High illumination (65 lux) | 0 / 0 / 40 | 60 / 40 / 60 |
| Flickering (0β45 lux @ 2 Hz) | 0 / 0 / 20 | 20 / 0 / 20 |
| Visual Distractor | 30 / 30 / 40 | 45 / 75 / 50 |
| Spatial (~6.5 cm) | 0 / 40 / 20 | 40 / 60 / 40 |
| Human Disturbance | 0 / 10 / 40 | 5 / 45 / 50 |
- Two-step Info-VAE vs. single-step (Fig. 11): Single-step performs roughly on par with direct fine-tuning; two-step delivers the full gain. Confirms that alignment to the pre-training prior β not just having a latent at all β is the load-bearing piece.
- Guidance scale Ξ» (Table 13): Ξ» = 1 / 2 / 3 swept on four tasks. Ξ» = 2 is the sweet spot; Ξ» = 3 occasionally hurts because too-strong guidance disrupts chunk smoothness.
- Reverse-KL weight Ξ± (Table 14): Ξ± β {0.1, 1, 10}. Ξ± = 1 is best; Ξ± = 10 over-constrains the adaptation VAE (e.g., Block Handover 92% β 37%); Ξ± = 0.1 weakens alignment.
- InfoVAE vs. Vanilla VAE (Tables 15-16): Vanilla VAE already beats baseline (showing the steering module is modular), but InfoVAE consistently adds another 4-20 points (e.g., DP on Empty Cup: baseline 22 β vanilla 34 β InfoVAE 37; Οβ on Blocks Stack Easy: 30 β 31 β 50).
- Sample efficiency (Table 12): at 25 demos/task, Οβ averages 9.2%, ATE 29.0% (3.15Γ). With 25 demos ATE matches/exceeds Οβ with 50 demos on multiple tasks β roughly 2Γ sample efficiency.
- DynaGuide comparison (Table 18): ATE matches or exceeds the test-time guidance baseline, without inference overhead (DynaGuide requires iterative gradient steps at every denoising step).
The paper does not have a dedicated "Limitations" section, but the experiments reveal:
- A handful of tasks (Bottle Adjust, Tool Adjust, Pick Apple Messy, Shoe Place) regress with ATE β guidance can occasionally fight the pre-training prior on contact-rich tasks the prior covers well.
- Ξ± and Ξ» must be tuned (the sensitivity studies show a 1-order-of-magnitude window).
- Step-1 VAE training itself takes ~12 h on the pre-training corpus per backbone, though this is amortised across all downstream adaptation tasks.
- Currently requires access to the pre-training action distribution β not just the VLA checkpoint β to train V_Ο. For closed checkpoints this would block direct application.
ATE regularises adaptation in action-distribution space rather than weight space. Compared to:
- Robust Parameter Merging β weight-space interpolation; complementary.
- UniVLA / GR00T N1 β shared latent action space (VQ-VAE / discrete latents), used as an intermediate representation. ATE instead learns a probability prior over actions and uses it as a classifier-guidance signal.
- Latent Action Diffusion (Bauer et al. 2025) β separate VAEs + contrastive learning; ATE replaces contrastive alignment with the more principled reverse-KL mode-seeking mechanism.
- Οβ FAST / OpenVLA fine-tuning β depend on data scale; ATE is most effective in the low-data, cross-embodiment regime where direct fine-tuning collapses.
- Test-time steering (DynaGuide): ATE bakes guidance into training, paying zero inference cost (vs. DynaGuide's iterative gradient steps), echoing the inference-efficiency arguments in SP-VLA and Action-aware Dynamic Pruning.
Conceptually closest in spirit to classifier-guided diffusion (Dhariwal & Nichol 2021) and motion-planning diffusion (Carvalho et al. 2023), repurposed for the VLA adaptation problem. Plug-and-play across DP, RDT-1B and Οβ is the central practical claim.
- OpenReview: https://openreview.net/forum?id=T3i7Ifeatk
- Code: https://github.com/TeleHuman/Align-Then-Steer
- Robust Parameter Merging β weight-space alternative
- UniVLA β task-centric latent action space
- X-VLA
- Ο0.6
- Survey: VLA & Manipulation
β Back to ICLR-2026