RSS 2026 Legato - Heungwoo/research GitHub Wiki

Legato โ€” Learning Native Continuation for Action Chunking Flow Policies

Venue: RSS 2026 (Manipulation session) ยท Authors: Yufeng Liu, Hang Yu, Juntu Zhao, Bocheng Li, โ€ฆ Dequan Wang, Yang Gao โ€” SJTU ร— Spirit AI ร— Tsinghua ร— Tongji ร— USTC (work done during an internship at Spirit AI) ยท arXiv: 2602.12978 ยท project Category: Real-time execution for chunked flow policies Trend tag: RSS 2026 thread 7 โ€” action representation & inference mechanics

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

Legato vs RTC (Figure 1 of arXiv 2602.12978, ยฉ the authors)

Figure 1 of the paper. Top: smoothness (NSPARC) vs completion time across five real tasks (bowl, drawer, pickplace, towel, pour) โ€” the Legato points (blue) sit left of and below their RTC counterparts (grey) on both axes: shorter execution and smoother trajectories. Bottom: an execution trace on the pour task โ€” RTC's y-position trace shows multimodal switching with poor overlap alignment between the executed and discarded chunk continuations (photo inset: visible "hesitation"), while Legato's trace is smooth with well-aligned chunk overlaps ("moving"). Hesitation-induced slowdowns are the concrete cost that native continuation removes.

Problem

Action chunking gives VLAs real-time throughput but discontinuities at chunk boundaries. Real-Time Chunking patches this externally (inference-time guidance) โ€” which causes spurious multimodal switching and trajectories that are never intrinsically smooth.

Method

Legato makes continuation native to training:

  • Denoising initializes from a schedule-shaped mixture of known (already-committed) actions and noise, exposing the model to partial action information during training;
  • The learned flow dynamics are reshaped so denoising stays consistent between training and inference under per-step guidance;
  • Randomized schedule conditioning during training supports varying inference delays and yields controllable smoothness.

Results (as reported)

  • Smoother trajectories, less spurious multimodal switching, less hesitation.
  • Across five real-world manipulation tasks: โ‰ˆ10% improvements over RTC in both trajectory smoothness and task completion time.

Significance

The clearest instance of RSS 2026's "internalize the inference patch" pattern: RTC (NeurIPS 2025, test-time) โ†’ training-time RTC (ฮจโ‚€ adopts it after finding test-time guidance unstable) โ†’ Legato (fully native continuation with delay randomization). Together with AR-VLA's persistent-context expert (#85) and Action-to-Action flow (#209), chunk-boundary handling has become its own sub-field โ€” directly relevant to every flow-matching VLA in Review-VLA-Attention's streaming discussion.

โ† RSS 2026 survey ยท Home