Latest Papers - Heungwoo/research GitHub Wiki

πŸ†• Latest Papers β€” preprint tracker

Recent preprints and notable model releases reviewed ahead of (or without) venue publication. These are fast-moving, often v1/early-revision works: numbers are quoted from the paper but unreplicated, and venue-of-record is pending. Once a paper is accepted, its page is cross-linked from the venue survey; until then it lives here.

Distinct from the conference surveys, which cover published/accepted work. Add newest first.

← Back to Home Β· full review catalog: Reviews


In-depth reviews

Posted Paper arXiv One-line
2026-08 [DYNA-2 (World-Action Model, human-video scaling law)](/Heungwoo/research/wiki/Review-Dyna2) (tech report β€” self-published, not peer-reviewed) Dyna Robotics' mixture-of-transformers WAM on ~1M h human egocentric video, no robot data in pre-training; joint next-frame+next-action; published 1kβ†’1M h power-law fits (RΒ²β‰ˆ0.88–0.93, ~50Γ— EgoScale) incl. zero-shot humanβ†’robot transfer, and 87% vs 46% over its DYNA-1 VLA β€” company self-reported, unreplicated
2026-08 [Ο‰-0 (humanoid loco-manipulation WAM)](/Heungwoo/research/wiki/Review-Omega0) 2608.06375 Whole-body humanoid WAM; reconstruction-free latent future prediction (not video) β†’ SONIC control; single model does 11 household tasks at 81.8% vs 44.5% best baseline; ships the 40 h Ο‰-HOME dataset
2025-11 β†’ 2026-05 [Stellar VLA (continual skill knowledge)](/Heungwoo/research/wiki/Review-Stellar-VLA) 2511.18085 Continual imitation learning for a fixed-size ~1B VLA with 1% replay; Dirichlet-Process self-evolving knowledge space (T-Stellar/TS-Stellar) + knowledge-routed diffusion MoE; SOTA CIL on LIBERO + real dual-arm
2026-04 [Cortex 2.0 (foresight-planning WAM+VLA)](/Heungwoo/research/wiki/Review-Cortex2) 2604.20246 Sereact's modular hybrid: world model generates k candidate futures in visual latent β†’ PRO scores progress/risk/termination β†’ flow action head commits, 30 Hz replan; deployment-reported 0.95–0.98 success / 0 interventions vs Ο€0.5 on industrial tasks (>10M-episode fleet) β€” company-reported
2026-04 [Being-H0.7 (latent world-action model)](/Heungwoo/research/wiki/Review-Being-H07) 2605.00078 BeingBeyond's unified + reactive hybrid: learnable latent queries as reasoning interface; future-informed posterior trains a deployable prior (no visual rollout at inference); InternVL3.5+Qwen3+V-JEPA2.1 on Being-H0.5; LIBERO 99.2%, RoboTwin2.0-Hard 89.6%
2025-12 [Motus (unified latent-action WAM)](/Heungwoo/research/wiki/Review-MOTUS) 2512.13030 THU-ML's unified hybrid: one 8B scheduled Mixture-of-Transformers is a world model / VLA / IDM / video generator on demand; optical-flow latent actions; open weights; RoboTwin 2.0 88.66% vs X-VLA 72.80% / Ο€0.5 42.98%

Where these connect

  • Ο‰-0 extends Humanoid VLA and sits in the World Models WAM taxonomy as the latent-predictive corner β€” a counterpoint to DreamZero's pixel-video WAM and to MotionWAM/DiT4DiT.
  • DYNA-2 is an industry WAM release backed by a self-published technical report (architecture + power-law fits, not peer-reviewed) pushing the EgoScale human-video scaling law and the DreamZero/Ο‰-0 joint video-action formulation to ~1M-hour, human-only scale β€” reviewed with an explicit not-peer-reviewed/unreplicated caveat.
  • Stellar VLA advances the continual-learning thread from RL for VLA and answers ICML 2026's Pretrained VLAs Resist Forgetting with a structured (not just replay-based) recipe.

← Back to Home