ICLR 2026 Align Then Steer - Heungwoo/research GitHub Wiki

Align-Then-stEer (ATE) β€” Unified Latent Guidance for VLA Adaptation

Venue: ICLR 2026 Affiliation: China Telecom Institute of AI Β· Tsinghua Β· CUHK Shenzhen Β· Northwestern Polytechnical Category: VLA Training β€” Adaptation / fine-tuning for diffusion- and flow-based VLAs Trend tag: Cross-embodiment / data efficiency / classifier guidance

Approach diagram

flowchart LR
  subgraph S1[Stage 1: Unified Action Latent Space]
    PT[Pre-training action chunks<br/>OXE / DROID / ALOHA / Kuka] --> V1[InfoVAE V_phi<br/>Transformer enc/dec<br/>latent dim 512]
    V1 -. "KL to N(0,I)" .-> Prior["Pretrain latent prior q_phi(z)"]
    Adapt[Adaptation action chunks<br/>RoboTwin / ManiSkill / RealMan] --> V2[InfoVAE V_psi<br/>same architecture]
    V2 -. "reverse KL to q_phi(z)" .-> Mode[Mode-seeking embed<br/>into one mode of q_phi]
  end
  subgraph S2[Stage 2: Classifier Guidance Steers Fine-tuning]
    Mode --> E[Frozen E_psi encoder]
    Noisy[a^k_t:t+h noisy chunk] --> E
    Clean[a^0_t:t+h clean chunk] --> E
    E --> Dist["β€–E_psi(a^k) βˆ’ E_psi(a^0)β€–Β²"]
    Dist --> Grad["g = -βˆ‡_a β€–z_hat βˆ’ zβ€–Β²"]
    Grad --> Diff["Diffusion-VLA loss + √(1βˆ’αΎ±_k)·λ·g"]
    Grad --> Flow["Flow-VLA loss + ((1βˆ’Ο„)/Ο„)·λ·g"]
  end
  Diff --> POL[Adapted RDT / DP]
  Flow --> POL2[Adapted Ο€β‚€]
Loading

Problem

Pre-trained VLAs (Ο€β‚€, RDT-1B, Diffusion Policy) face a distribution mismatch when fine-tuned on a new embodiment or task: degrees-of-freedom, end-effector representation, joint torques and chunk length all change between pre-training data (typically single-arm 6-DoF from Open X-Embodiment / DROID) and adaptation data (e.g., dual-arm 7-DoF RealMan with single-DoF grippers). Direct fine-tuning either overfits the small target set or drifts away from the pre-training visuomotor prior. Methods that operate at the parameter level (LoRA, dynamic layer activation) or use kinematics retargeting do not directly close the action distribution gap, and diffusion/flow-based VLAs are especially sensitive because their output is a distribution over action chunks, not a point estimate.

Method (detailed)

Stage 1 β€” Two InfoVAEs build a unified action latent space

The authors instantiate two InfoVAEs (Zhao et al. 2017), not vanilla VAEs, because vanilla VAEs suffer from posterior collapse and inaccurate amortized inference. Each VAE has a Transformer encoder/decoder. Two learnable tokens ΞΌ_token and Ξ£_token are prepended to the encoder input (motion-VAE style, Γ  la Petrovich et al. 2021), and the decoder uses H zero embeddings as queries with the latent z as key/value via cross-attention.

Pre-training VAE V_phi (encoder E_phi, decoder D_phi):

  • Trained on the cross-embodiment pre-training corpus with the standard InfoVAE objective:

    L(Ο†) = E_q[log p_Ο†(a|z)] βˆ’ (1βˆ’Ξ±)Β·D_KL(q_Ο†(z|a) β€– p(z)) βˆ’ (Ξ±βˆ’Ξ»βˆ’1)Β·D_KL(q_Ο†(z) β€– p(z))

    where p(z) = N(0, I), and the second KL term is replaced by MMD for tractable optimisation (per Zhao et al. 2017). Uses the same chunk length H as the pre-training VLA.

Adaptation VAE V_psi (encoder E_psi, decoder D_psi):

  • Same architecture and variant. Trained on the small adaptation set with a reverse KL to the learned pre-training latent distribution q_Ο†(z):

    L(ψ) = E_q[log p_ψ(Γ£|z)] βˆ’ (1βˆ’Ξ±)Β·D_KL(q_ψ(z|Γ£) β€– q_Ο†(z)) βˆ’ (Ξ±βˆ’Ξ»βˆ’1)Β·D_KL(q_ψ(z) β€– q_Ο†(z))

  • q_Ο†(z) is approximated as N(ΞΌ_Ο†, Ξ£_Ο†) with statistics aggregated over all pre-training action latents. The mode-seeking property of reverse KL is the key: it pushes q_ψ into one mode of q_Ο†, producing a unified, hierarchically structured latent space.

  • Uses the adaptation chunk length L (which may differ from H).

t-SNE visualisations in Fig. 15 confirm that InfoVAE produces a compact latent space where adaptation embeddings (blue) sit inside pre-training modes (green); vanilla VAE produces poorly aligned, gapped distributions.

Stage 2 β€” Classifier guidance in the unified latent space

A guidance function is constructed in the unified latent Z. Following Carvalho et al. 2023, the classifier is an energy-based model:

p_ψ(y | Γ’^k_t:t+h) = (1/Z_ψ) Β· exp(βˆ’β€–E_ψ(Γ’^k_t:t+h) βˆ’ E_ψ(a^0_t:t+h)β€–Β²)

with guidance gradient:

g = βˆ‡_Γ’ log p_ψ(y|Γ’) ∝ βˆ’βˆ‡_Γ’ β€–E_ψ(Γ’^k) βˆ’ E_ψ(a^0)β€–Β²

This gradient is folded into the training objective of the underlying generative head:

  • Diffusion-based VLAs (RDT, DP): L(ΞΈ) = E[β€–Ξ΅ βˆ’ Ξ΅_ΞΈ(a^k; k, o_t, l) + √(1 βˆ’ αΎ±_k)·λ·gβ€–Β²]
  • Flow-based VLAs (Ο€β‚€): L(ΞΈ) = E[β€–v_ΞΈ(a^Ο„; Ο„, o_t, l) + ((1βˆ’Ο„)/Ο„)·λ·g βˆ’ (a^0 βˆ’ Ξ΅)β€–Β²]

The encoder E_psi is frozen during this stage; it serves purely as a distance metric. Crucially, since guidance enters at training time only, inference cost is zero β€” unlike test-time gradient-based steering (DynaGuide) that runs iterative backprop per step. The recommended guidance scale is Ξ» = 2 (Table 13).

Hyperparameters (Tables 9-10 + appendices)

Component RDT-1B Ο€β‚€ Diffusion Policy
InfoVAE latent dim 512 512 512
Temporal input length 64 50 14
Input dim 128 32 14
Optimizer Adam Adam Adam
LR (VAE) 1e-4 1e-4 1e-4
Weight decay 1e-4 1e-4 1e-4
Batch size (VAE) 64 64 64
Chunk size (adaptation) 64 50 8
Batch size (fine-tune) 64 24 128
LR (fine-tune) 1e-4 2.5e-5 1e-4
Training duration 100k steps 60k steps 300 epochs
Step-1 VAE wall-clock ~12 h ~12 h ~12 h
Step-2 VAE wall-clock < 0.5 h < 0.5 h < 0.5 h
Compute 4Γ— A100 4Γ— A100 1Γ— A100

Real-robot fine-tuning runs 120k steps on 8Γ— A100 with batch 6/GPU at 20 Hz inference on a single 4090.

Comprehensive Results

RoboTwin 1.0 β€” 17 dual-arm tasks (Table 1, 50 trajectories/task)

Task RDT-1B RDT + ATE (Ξ”) Ο€β‚€ Ο€β‚€ + ATE (Ξ”)
Block Hammer Beat 52 71 (+19) 38 44 (+6)
Block Handover 69 91 (+22) 80 92 (+12)
Blocks Stack (Easy) 10 31 (+21) 30 50 (+20)
Blocks Stack (Hard) 1 7 (+6) 8 7 (βˆ’1)
Bottle Adjust 53 37 (βˆ’16) 39 45 (+6)
Container Place 34 55 (+21) 56 59 (+3)
Diverse Bottles Pick 18 24 (+6) 20 40 (+20)
Dual Bottles Pick (Easy) 76 87 (+11) 48 85 (+37)
Dual Bottles Pick (Hard) 39 58 (+19) 52 55 (+3)
Dual Shoes Place 6 9 (+3) 22 24 (+2)
Empty Cup Place 22 61 (+39) 32 36 (+4)
Mug Hanging (Easy) 6 6 (0) 11 27 (+16)
Mug Hanging (Hard) 1 1 (0) 4 4 (0)
Pick Apple Messy 35 39 (+4) 18 11 (βˆ’7)
Put Apple Cabinet 20 45 (+25) 34 55 (+21)
Shoe Place 44 43 (βˆ’1) 52 58 (+6)
Tool Adjust 54 42 (βˆ’12) 70 69 (βˆ’1)
Average 31.8 41.6 (+9.8) 36.1 44.8 (+8.7)

Sample-efficiency: RDT + ATE surpasses RDT baseline at 70k steps vs. 90k for the baseline. ATE on Ο€β‚€ at 25 demos/task averages 29.0% vs. baseline 9.2% (3.15Γ— gain, Table 12).

ManiSkill3 (Table 5, 100 demos/task, RDT-1B backbone)

Task RDT-1B RDT + ATE
Push Cube 65.2 78.4 (+13.2)
Pick Cube 7.6 14.8 (+7.2)
Average 36.4 46.6 (+10.2)

LIBERO-10 comparison vs. offline DSRL (Table 17)

Method Success rate
Fine-tuned Ο€β‚€ 0.78
+ Offline DSRL 0.43
+ ATE (Ours) 0.88

Real-world RealMan dual-arm (160 demos/task; 8Γ— A100; 20 trials each)

Four long-horizon tasks at 60k / 90k / 120k steps. Tasks include Cook Bun, Pick Bun, Make Sandwich, Use Toaster. Reported in Fig. 3:

  • Cook Bun: Ο€β‚€ ~15% vs. ATE 100% at 90k steps.
  • Pick Bun (single-arm): ATE 50% @ 90k β†’ 70% @ 120k.
  • Overall average at 120k: Ο€β‚€ 16.7% vs. ATE 58.1%.

Make Yogurt Bowl (multi-tool dual-arm, 80 demos): Ο€β‚€ 15% β†’ ATE 25% (Table 2). Under visual distractors: Ο€β‚€ 0% β†’ ATE 20% (Table 3).

Real-world generalisation (Tables 7-8)

Setting Ο€β‚€ (Cook / Pick / Sandwich) Ο€β‚€ + ATE
Low illumination (25 lux) 0 / 0 / 40 60 / 40 / 40
High illumination (65 lux) 0 / 0 / 40 60 / 40 / 60
Flickering (0–45 lux @ 2 Hz) 0 / 0 / 20 20 / 0 / 20
Visual Distractor 30 / 30 / 40 45 / 75 / 50
Spatial (~6.5 cm) 0 / 40 / 20 40 / 60 / 40
Human Disturbance 0 / 10 / 40 5 / 45 / 50

Ablation Studies (Appendix D.5, J)

  1. Two-step Info-VAE vs. single-step (Fig. 11): Single-step performs roughly on par with direct fine-tuning; two-step delivers the full gain. Confirms that alignment to the pre-training prior β€” not just having a latent at all β€” is the load-bearing piece.
  2. Guidance scale Ξ» (Table 13): Ξ» = 1 / 2 / 3 swept on four tasks. Ξ» = 2 is the sweet spot; Ξ» = 3 occasionally hurts because too-strong guidance disrupts chunk smoothness.
  3. Reverse-KL weight Ξ± (Table 14): Ξ± ∈ {0.1, 1, 10}. Ξ± = 1 is best; Ξ± = 10 over-constrains the adaptation VAE (e.g., Block Handover 92% β†’ 37%); Ξ± = 0.1 weakens alignment.
  4. InfoVAE vs. Vanilla VAE (Tables 15-16): Vanilla VAE already beats baseline (showing the steering module is modular), but InfoVAE consistently adds another 4-20 points (e.g., DP on Empty Cup: baseline 22 β†’ vanilla 34 β†’ InfoVAE 37; Ο€β‚€ on Blocks Stack Easy: 30 β†’ 31 β†’ 50).
  5. Sample efficiency (Table 12): at 25 demos/task, Ο€β‚€ averages 9.2%, ATE 29.0% (3.15Γ—). With 25 demos ATE matches/exceeds Ο€β‚€ with 50 demos on multiple tasks β€” roughly 2Γ— sample efficiency.
  6. DynaGuide comparison (Table 18): ATE matches or exceeds the test-time guidance baseline, without inference overhead (DynaGuide requires iterative gradient steps at every denoising step).

Limitations (stated by authors)

The paper does not have a dedicated "Limitations" section, but the experiments reveal:

  • A handful of tasks (Bottle Adjust, Tool Adjust, Pick Apple Messy, Shoe Place) regress with ATE β€” guidance can occasionally fight the pre-training prior on contact-rich tasks the prior covers well.
  • Ξ± and Ξ» must be tuned (the sensitivity studies show a 1-order-of-magnitude window).
  • Step-1 VAE training itself takes ~12 h on the pre-training corpus per backbone, though this is amortised across all downstream adaptation tasks.
  • Currently requires access to the pre-training action distribution β€” not just the VLA checkpoint β€” to train V_Ο†. For closed checkpoints this would block direct application.

Significance & Positioning

ATE regularises adaptation in action-distribution space rather than weight space. Compared to:

  • Robust Parameter Merging β€” weight-space interpolation; complementary.
  • UniVLA / GR00T N1 β€” shared latent action space (VQ-VAE / discrete latents), used as an intermediate representation. ATE instead learns a probability prior over actions and uses it as a classifier-guidance signal.
  • Latent Action Diffusion (Bauer et al. 2025) β€” separate VAEs + contrastive learning; ATE replaces contrastive alignment with the more principled reverse-KL mode-seeking mechanism.
  • Ο€β‚€ FAST / OpenVLA fine-tuning β€” depend on data scale; ATE is most effective in the low-data, cross-embodiment regime where direct fine-tuning collapses.
  • Test-time steering (DynaGuide): ATE bakes guidance into training, paying zero inference cost (vs. DynaGuide's iterative gradient steps), echoing the inference-efficiency arguments in SP-VLA and Action-aware Dynamic Pruning.

Conceptually closest in spirit to classifier-guided diffusion (Dhariwal & Nichol 2021) and motion-planning diffusion (Carvalho et al. 2023), repurposed for the VLA adaptation problem. Plug-and-play across DP, RDT-1B and Ο€β‚€ is the central practical claim.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️