ICML 2026 Escaping the Diversity Trap in - Heungwoo/research GitHub Wiki

Escaping the Diversity Trap in Robotic Manipulation via Anchor-Centric Adaptation — Repeat before you expand

Venue: ICML 2026 (Poster) Category: Analysis-Insight Affiliations: Show Lab, National University of Singapore (Yanzhe Chen, Kevin Yuchen Ma, Qi Lv, Yiqi Lin, Zechen Bai, Chen Gao, Mike Zheng Shou) Traction (2026-06): 0 citations (arXiv)

Illustration of the diversity trap motivation: sparse single-shot sampling versus anchor-centric repetition, and the inverted-U trend of success vs. number of anchors (Figure 1 from Chen et al., 2026)

Problem

Vision-Language-Action (VLA) models offer broad generalist capabilities from pretraining, but deploying them on a specific physical platform requires real-world post-training to bridge the embodiment gap and subtle distribution shifts. Because robot demonstrations are expensive, this adaptation typically happens under a tight data budget — tens to hundreds of trajectories. The paper asks: under a fixed budget, what data-collection strategy yields the most robust policy?

The standard heuristic is to maximize coverage — collect single-shot demonstrations across as many distinct conditions as possible. The authors identify this as a diversity trap: spreading scarce samples too thin leaves each condition under-represented, inflating estimation variance and destabilizing the learned action vector field. Empirically, task success follows an inverted-U in the number of distinct conditions (anchors): past an optimum, more diversity erodes per-condition density and collapses policy stability.

Method

The authors formalize this as a Coverage–Density Trade-off. By analyzing conditional action vector fields (flow matching), they decompose worst-case policy error into an estimation (density) term and an extrapolation (coverage) term, proving the existence of an interior optimal allocation of unique conditions for a fixed budget — challenging the optimality of uniform diverse sampling under non-vanishing noise (σ > 0).

Guided by this analysis, Anchor-Centric Adaptation (ACA) is a two-stage framework:

  • Stage 1 — Anchor-Centric Stabilization: concentrate the budget on a minimal set of repeated "anchor" conditions to consolidate a low-variance "policy skeleton" (train the action expert for 20K steps).
  • Boundary Mining via Teacher-Forced Deviation: use the base policy's teacher-forced field error e(p) (L1 norm averaged over 7 dims: position, orientation, gripper) as a proxy for extrapolation risk to identify under-supported boundary regimes (k ∈ {2,3,5} mined conditions for budgets {50,100,150}).
  • Stage 2 — Constrained Residual Adaptation: integrate boundary-specific knowledge through a parameter-efficient residual pathway (LoRA, rank/α = 32, 10K steps), incorporating critical boundary data without degrading the consolidated core.

Overview of Anchor-Centric Adaptation (ACA) (Figure 2 from Chen et al., 2026)

Results

ACA is instantiated on flow-matching VLAs (π0 and π0.5) and evaluated on a 7-DoF Franka Panda over four tabletop tasks (Block Stacking, Cup Placement, Table Cleaning, Toy Tidying), using nested Region-Level Success Rates S@1 (core, 25% area), S@2 (50%), S@3 (near-boundary, 90%), averaged over 20 trials/task.

  • ACA outperforms the maximal-diversity baseline at every budget. At N=100, ACA reaches 72.5% mean success — +40.8% absolute over the baseline.
  • It alleviates boundary collapse: the diversity-first baseline frequently scores 0% in S@3, whereas in Block Stacking at N=150 ACA reaches 80% (16/20) vs. 20% for the baseline.
  • Ablations: even a spatially biased top-left anchor layout gives a 4× S@3 improvement (51.0% vs. 12.5%); centralized anchor layouts perform best; success follows a consistent inverted-U over the number of anchors K ∈ {0,…,12}, with the optimum shifting with budget.

Significance

The work reframes budget-constrained VLA adaptation from "what to cover" to "how to allocate," backing the principle of repetition before expansion with both a worst-case error bound and real-robot evidence. It offers a practical, backbone-agnostic recipe (stabilize on anchors → mine boundaries → residual-adapt) for the common low-data fine-tuning regime where naïve diversity sampling is actively harmful.

Links

← Back to ICML-2026