RSS 2026 BagelVLA - Heungwoo/research GitHub Wiki
BagelVLA: Enhancing Long-Horizon Manipulation via Interleaved Vision-Language-Action Generation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #83 Authors: Yucheng Hu, Jianke Zhang, Yuanfei Luo, Yanjiang Guo, Xiaoyu Chen, Sun Xinshu, Kun Feng, Qingzhou Lu, Sheng Chen, Yangang Zhang, Wei Li, Jianyu Chen arXiv: 2602.09849 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

The hybrid training data wheel (center) mixes VQA data, self-collected robot data (Agibot, Open-galaxea, RoboTwin), open-resourced robot data, and egocentric human video. The bottom row shows the three interleaved capabilities: linguistic planning (chain-of-thought subtask selection, e.g. solving an equation before picking a block), visual forecasting (generating a subtask keyframe), and action generation (continuous action chunks).
Problem
VLA models built on foundation models typically do either linguistic planning or visual forecasting in isolation, and rarely integrate both to guide action generation — hurting complex, long-horizon manipulation that requires multi-stage reasoning (e.g., computing an arithmetic answer before placing symbol blocks).
Method
BagelVLA initializes from Bagel, a unified understanding-and-generation transformer, and trains it to interleave textual subtask reasoning, keyframe prediction, and action generation in one execution loop. Stage 1 builds a hybrid dataset (general multimodal data plus robotic datasets annotated with subtasks and keyframes); Stage 2 adds an action expert and fine-tunes the full model (7B understanding/generation expert + 2B action expert). Residual Flow Guidance (RFG) conditions on the current observation and uses a single denoising step to extract predictive visual features — foresight without full image synthesis — giving 1.2 s per 48-action chunk on one RTX 5090 (40 Hz; 72 Hz with asynchronous execution).
Results
On Calvin ABC-D, BagelVLA reaches 4.405 average completion length versus 4.329 for VPP, 4.078 for UP-VLA, and 3.648 for π0. On RoboTwin 2.0 (50 tasks) it scores 75.26% Clean / 20.87% Randomized versus 46.42/16.34 for π0; textual planning alone adds 21% on RoboTwin. On the real Aloha-AgileX bimanual platform, it averages 75.5% over 9 basic-task categories (π0 65.0, VPP 59.5) and on long-horizon planning tasks scores 73.3% (Stack Cubes in Requested Order) and 63.3% (Calculate and Place Symbol Blocks) versus 40.0/31.7 for π0, with ~90% planning accuracy.
Significance
A working demonstration that a single unified generative backbone can "think in text, imagine in pixels, and act" — with RFG addressing the latency objection that has kept visual forecasting out of real-time control loops. Related wiki threads: Review-VLA-Architecture · RL.
← Back to RSS 2026 survey · RSS-2026-Papers · Home