RSS 2026 DISC - Heungwoo/research GitHub Wiki

DISC: Decoupling Instruction from State-Conditioned Control via Policy Generation

Venue: RSS 2026 (Sydney, Jul 13โ€“17) ยท Session: Imitation learning 2 ยท paper #147 Authors: Hanxiang Ren, Pei Zhou, Xunzhe Zhou, Yanchao Yang arXiv: 2605.20856 ยท program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

DISC behavioral evidence of task-state entanglement (Figure 1 of arXiv 2605.20856, ยฉ the authors)

Figure 1. Two rollout comparisons illustrating observation leakage. Left ("put the white bowl to the right of the plate"): entangled Octo instead approaches the microwave, executing a behavior tied to a similar scene, while DISC follows the instruction. Right ("turn on the stove and put the frying pan on it"): pretrained ฯ€โ‚€.โ‚… skips the stove-activation subgoal and just places the pan, whereas DISC completes both steps.

Problem

Language-conditioned manipulation policies typically process instruction and observation through shared parameters โ€” "task-state entanglement." This lets networks learn scene-to-action shortcuts that bypass language grounding entirely (observation leakage), so policies succeed on familiar scenes but fail when instructions change or visual context is ambiguous.

Method

DISC removes the failure structurally: rather than conditioning one universal policy on language, a hypernetwork generates the entire parameter set of a task-specific visuomotor policy from the instruction alone, and that generated policy only ever sees observations โ€” never language โ€” so there is no pathway for leakage. Generating coherent high-dimensional weights is hard, so DISC uses a two-stage hypernetwork: a Weight Initialization Network maps the language embedding to a semantically informed point in parameter space, and an Iterative Refinement Module embeds the structure of gradient-based optimization (forward evaluation, error estimation, correction) as a feed-forward inductive bias, producing globally consistent parameters without actual gradient computation. It is trained entirely from scratch on standard data budgets with no external pretraining.

Results

DISC achieves 94.3% on LIBERO-90 and 92.2% on Meta-World, beating the strongest trained-from-scratch entangled baseline by 7.7% on LIBERO-90, with advantages widening on long-horizon tasks. It surpasses the large-scale pretrained ฯ€โ‚€ (91.6%) and remains competitive with ฯ€โ‚€.โ‚… (95.7%) despite using no pretraining data. On a real-world combinatorial benchmark where visual context is shared across all 9 tasks, DISC reaches 86.4% versus 78.5% for the best entangled baseline โ€” direct evidence that generated parameters, not visual shortcuts, drive behavior. The learned parameter manifold further supports few-shot adaptation and robustness to paraphrased instructions.

Significance

Offers an architectural (rather than data-scale) route to genuine language grounding, sharpening the instruction-following debates in Review-LBM-Cotraining and pretraining-centric VLA work.

โ† Back to RSS 2026 survey ยท RSS-2026-Papers ยท Home