Changelog - Heungwoo/research GitHub Wiki
Changelog
Reverse-chronological log of page additions and major updates to this wiki. Dates are page-addition dates; month-level where the exact day is approximate. Back to Home.
2026-08
- 2026-08-11 — DYNA-2 (in-depth) added to Latest Papers — Dyna Robotics' World-Action Model launch (Aug 10, 2026): a WAM on ~1M h human egocentric video with no robot data in pre-training, joint next-frame+next-action, claiming the first human-to-robot scaling law smooth over 1k→1M h (~50× EgoScale) and 87% vs 46% zero-shot pass over its DYNA-1 VLA. Reviewed with an explicit vendor-claim caveat — company announcement, no technical paper/benchmark/weights. Cross-linked from World-Models, Human-Video-Transfer, Reviews.
- 2026-08-11 — New Latest Papers preprint tracker + two in-depth preprint reviews: ω-0 (arXiv 2608.06375 — whole-body humanoid World Action Model using reconstruction-free latent future prediction + SONIC control; single model does 11 household loco-manipulation tasks at 81.8% SR / 90.3% progress vs 44.5% / 59.6% for ψ-0; ships the 40 h six-modality ω-HOME dataset; Fig. 2 embedded) and Stellar VLA (arXiv 2511.18085 — continual imitation learning for a fixed ~1B VLA with 1% replay via a Dirichlet-Process self-evolving knowledge space + knowledge-routed diffusion MoE; SOTA CIL on LIBERO and 90.0% Final SR on real dual-arm; Fig. 2 embedded). Cross-linked from Home, Reviews, sidebar, Humanoid-VLA and World-Models reviews.
- 2026-08-10 — DreamZero (in-depth) — NVIDIA's World Action Models are Zero-shot Policies (arXiv 2602.15922): a 14B video-diffusion World Action Model that jointly predicts video+action, beats SOTA VLAs by >2× on unseen-env/unseen-object real-robot evals (62.2% vs 27.4% seen-task progress; VLAs ~0% at matched scale), with a 38× inference stack (DreamZero-Flash decoupled noise schedule) for 7 Hz closed-loop control and video-only cross-embodiment transfer (+42% from ≤20 min; 30-min new-robot adaptation). Fig. 4 architecture embedded. Cross-linked from World Models review, Home, Reviews, sidebar.
- 2026-08-05 — (fix) RSS 2026 survey hit GitHub's wiki render limit (262 wikilinks / 40 KB) — the full session tables (116 linked rows) moved to the new RSS-2026-Papers index page; the survey keeps a pointer and now renders at ~147 links / 21 KB.
- 2026-08-05 — RSS 2026 coverage completed: 101 new per-paper reference pages generated for every remaining in-scope paper (verbatim program abstract · session/authors/program-page metadata · related-topic-review links), bringing all 116 in-scope papers to full page coverage. RSS 2026 survey updated: all session tables fully linked (103 rows), the condensed Humanoids/WM/RL/Datasets block expanded into five per-session tables with abstract-first-line glosses (incl. a new hands/tactile picks table), 100 bolded mentions in the themed lists linked, and each of the 8 themed reading lists now opens with a 🧠 technology-level insight paragraph (reward-source differentiation in RL, the three-camp mechanism choice for human data, prediction-space selection for world models, touch-as-predicted-state, the hand canonicalization playbook, decomposition-beats-end-to-end for humanoids, ordered-tokens vs native-continuation levers, evaluation-as-systems-discipline). RSS hub updated.
- 2026-08-05 — Deep-dive surveys updated with ICML 2026 (previously RSS-only): 14 sections gain ICML evidence from the 99-paper index — architecture (VLANeXt recipe · From-Pixels-to-Tokens Oral · XR-1 Oral), wiring (MoT dual-systems HALO/LaST₀ · Move-Then-Operate · LangForce shortcut counter), attention/real-time (the 9-paper efficiency cluster: Reflex 50 Hz streaming · GridS −76% FLOPs · SpecPrune/EcoVLA · XPU two-phase profile · latent reasoning −90% latency), RL (VLAC/LAGEA/ReLAM rewards · VLA-MBPO/VLAW model-based +39.2% · test-time critics), memory (HiMe · SOMA · CAPS drift), world models (DreamDojo 44k h · LAC-WM +46.7% · dWorldEval action-token evaluator), dexterous (DexMachina · DECO · Tabero −70% grip force · CTSRL), cross-embodiment (OXE-AugE +24–45% · latent motion codes), evaluation (LIBERO-Gen tiers · VLA-Arena · FixBench · TRAP adversarial), human-video (video→world-model fourth use). Verdicts revised: forgetting is milder than assumed (VLA-Forgetting Oral) but untested under repeated RL; discrete-token verdict scoped to robot-action auxiliaries (latent-action tokens as VLM supervision are effective); the WM-evaluator action-input gap has its first crack (dWorldEval). Home fold-outs synced.
- 2026-08-05 — Decision-map restructure for scannability: every Home fold-out now leads with a prominent 📄 Deep dive → link (previously buried after the limitations text) and compresses 📈 Trend / ⚖️ Approaches / ⚠️ Open into a compact 3-row label table; the 12 detail-page survey sections rewritten to a uniform template —
## 🗓 State of the Field (updated Aug 2026)+ one-line Verdict quote +📈 Trend/⚖️ Approaches & trade-offs/ (optional✅ Established findings) /⚠️ Limitations & open problems— with content carried over and tightened; the three standalone surveys (Review-Human-Video-Transfer · Review-VLA-Evaluation · Review-Realtime-Execution) got matching (updated Aug 2026) headers with structure legends.
2026-07
- 2026-07-27 — Decision-map topics upgraded to survey-report depth on their detail pages: a dated "State of the Field — July 2026" section (trend arc · approach taxonomy with trade-offs · limitations) added to 12 topic reviews (Review-VLA-Architecture · Review-VLM-Action-Connection · Review-VLA-Attention · Review-Independent-Visual-Representation · Review-VLA-Training-Frameworks · RL · Review-VLA-Memory · Review-World-Models · Review-Dexterous-Manipulation · Review-Cross-Embodiment · Review-Humanoid-VLA · Review-LBM-Cotraining consolidated-verdict table), and three new topic surveys created for previously page-less questions: Review-Human-Video-Transfer (emergence vs decoupling vs synthesis, with a decision guide), Review-VLA-Evaluation (the indictment, approach taxonomy, emerging norms), Review-Realtime-Execution (RTC→Legato arc, approach comparison, latency-reporting gap). Home fold-outs now link the full surveys; Reviews catalog updated.
- 2026-07-27 — Home Research Decision Map rewritten from a link table into an insight map: all 15 topics are now collapsible entries, each carrying 📈 the trend arc through the latest venues (RSS 2026 / ICML / ICRA / ICLR 2026), ⚖️ the competing approaches with definitions and trade-offs, and ⚠️ current limitations — e.g. the flow-vs-AR arc with OAT's revival, the cross-attention-vs-concatenation ablation status, the RL-from-experience production milestone, the human-video emergence-vs-decoupling fork, the co-training verdicts, the evaluation-era shift, and the alignment-as-scaling-precondition reframe for cross-embodiment. All claims sourced from wiki-verified numbers.
- 2026-07-25 — Home redesign: stat strip added, Research Decision Map reorganized into four themed tables (🏗 Building · ⚡ Running & improving · 📊 Data & evaluation · 🦾 Embodiment) with a "current answer in one line" column, emoji section headers, compact 3-row reviews table; the Page Format / Maintenance Rule section moved off Home to the new Maintenance page (footer link).
- 2026-07-25 — Navigation restructure: new [Reviews]] page as the full in-depth-review catalog (topic reviews · lab programs · per-paper long-forms · RSS 2026 figure pages); [_Sidebar slimmed from ~250 to ~45 lines, keeping only top-level entries (reviews hub + 6 star topics, model lineages, ML hub, one link per venue year, foundational refs) — per-paper links now live on venue pages and in Reviews; Home §3/§5 merged into a compact reviews section pointing at the catalog.
- 2026-07-25 — Knowledge Graph updated with RSS 2026: venue node added to the overall map (+ previously-missing ICML), Tactile-VLA/Humanoid-VLA nodes and RSS-labeled edges in the reviews graph, lineage extensions (RECAP marked RSS-oral; RTC → Legato branch + Ψ₀ adoption; new TRI-LBM and LIBERO robustness lines), world-model cluster additions (mimic-video, LDA-1B under a new unified-WM branch, Qwen-RobotWorld), and a new fifth view: the RSS 2026 improvement-loop cluster mapping all six threads to their 15 pages.
- 2026-07-25 — Home Research Decision Map refreshed with RSS 2026 findings: four new question rows (policy improvement from experience · human-video-vs-robot-data fork · co-training data selection · credible evaluation) and five rows updated with RSS evidence (Legato/OAT for real-time inference, LDA-1B/mimic-video for world models, ViTacFormer/CGP for dexterity, cross-hand pages + camera-frame EEF for cross-embodiment, Ψ₀ for humanoids). Header stats and start-here pointer (→ RSS survey) updated.
- 2026-07-25 — Remaining six RSS 2026 per-paper pages upgraded to figure-illustrated reviews with confirmed arXiv IDs: RSS-2026-OAT (2602.04215, Harvard×Stanford — desiderata Venn + anytime-decoding chart), RSS-2026-Legato (2602.12978, SJTU×Spirit AI — smoothness-vs-time scatter + hesitation traces), RSS-2026-Contact-Grounded-Policy (2603.05687, Purdue×Meta — teleop channels + predicted-contact pipeline; CVPR-workshop Outstanding award noted), RSS-2026-DexGrasp-Zero (2603.16806 — paradigm comparison + unseen-hand deployment grid), RSS-2026-One-Hand (2602.16712, UNC — canonical-hand overview), RSS-2026-LIBERO-X (2602.06556, Meituan×Beihang — L1–L5 pyramid; 2,520 demos/600 tasks/100 scenes added). Sidebar: RSS added to Conferences + full RSS 2026 section + recent in-depth reviews (Qwen series, Ψ₀) listed.
- 2026-07-25 — RSS 2026 major papers upgraded to figure-illustrated 1-page reviews: key figures extracted from the arXiv originals (all verified) with explanatory captions added to PI-RECAP (RECAP loop overview), Review-Psi0 (G1 pantry teaser), RSS-2026-LDA-1B (data-tier/objective/results teaser), RSS-2026-Human2Robot-Emergence (emergence-vs-diversity chart), RSS-2026-mimic-video (VLA-vs-VAM + 10× chart), RSS-2026-LBM-Cotraining-Study (modality-matrix overview), RSS-2026-ViTacFormer (CVAE architecture + SharpaWave hardware details), RSS-2026-PolaRiS (4-part system + correlation scatter), RSS-2026-HoMMI (collection/gap/skills). arXiv IDs confirmed for all nine.
- 2026-07-25 — RSS 2026 — VLA & Manipulation Survey — Sydney, Jul 13–17; 210 accepted papers, ~116 manipulation/hand/humanoid in scope, all abstracts verified against the official program. Six threads: RL-from-experience for VLAs (π*0.6/RECAP flagship), human-video transfer (emergence vs decoupling), video/world models vs VLA backbones, contact-as-representation, cross-embodiment dexterous hands, evaluation infrastructure. + RSS venue hub, 14 new per-paper pages (LDA-1B · H2R-Emergence · mimic-video · LBM co-training study · ViTacFormer · DexGrasp-Zero · One-Hand · Contact-Grounded Policy · PolaRiS · LIBERO-X · OAT · Legato · HoMMI), and PI-RECAP updated with the RSS camera-ready results.
- 2026-07-25 — Ψ₀ (in-depth) — open humanoid loco-manipulation foundation model (USC PSI Lab × NVIDIA, RSS 2026, arXiv 2603.12263): decoupled human-video→VLM / robot-data→MM-DiT recipe; 800 h EgoDex + 30 h robot data beats >10× corpora (incl. GR00T N1.6) by >40 pp on 8 real Unitree-G1 tasks; full-paper verified.
- 2026-07-24 — Qwen-VLA (in-depth) (minor update) — re-verified against arXiv v2 (Jun 1, 2026): added the new no-T2A baseline (60.9% → T2A worth +10.2 pp) and the v2 clarification that T2A shares the downstream action representation (chunk-first-frame delta EEF); T2A ablation numbers aligned to v2's one-decimal precision. All benchmark tables unchanged between versions.
- 2026-07-24 — VLM4VLA (in-depth) (minor update) — re-verified against arXiv v2 (May 2026): corrected π0 Calvin Task-3 (0.786 → 0.686, the only substantive v1→v2 change); added version history and explicit Tsinghua × Qwen affiliation split to the header.
- 2026-07-24 — Qwen-RobotNav (in-depth) — navigation specialist on Qwen3-VL: parameterized observation interface (token budget · temporal decay · camera weights) with training-time randomization, 15.6M-sample corpus, agentic two-tier system (Qwen3.6-Plus planner + evidence notebook) with EQA SOTA; beats Qwen-VLA's nav numbers by ~15 pp.
- 2026-07-24 — Qwen-RobotWorld (in-depth) — language-actioned video world model: 20B double-stream MMDiT + frozen Qwen2.5-VL action encoder, 8.6M-pair EWK corpus with five-layer action-language annotation, Scene2Robot H2R editing; 1st on EWMBench/DreamGen Bench. Program page updated with both.
- 2026-07-24 — Qwen Team's VLA Program — cross-paper review of the Qwen team's VLA line (VLM4VLA → Qwen-VLA → Qwen-Robot Suite incl. RobotNav/RobotWorld): the shared doctrine (Qwen3.5-4B · flow matching · λ=0.1 VL co-training · synthetic-data scaling · language-as-interface), the diagnostic-to-flagship trace, and the internal contradictions between the two flagship VLAs.
- 2026-07-24 — Qwen-RobotManip (in-depth) — the Qwen team's second VLA (arXiv 2606.17846): alignment-first scaling thesis (80-dim canonical rep · camera-frame delta EEF + CaPE · in-context adaptation), ~38,100 h open-data-only corpus (24,808 h human-to-robot synthesis across 15 platforms), new RoboTwin-IF / RoboTwin-XE OOD benchmarks, RoboChallenge Table30-v1 generalist #1.
2026-06
- 2026-06-11 — Action Space: EEF vs Joint (in-depth) — taxonomy (joint vs EEF · absolute vs delta · chunking), the EEF-vs-joint verdict, and the O(k)-vs-O(1) chunk-wise-delta result.
- 2026-06-11 — VLA Training Frameworks — cross-paper review of StarVLA (Lego-modular) vs TRI VLA Foundry (LLM→VLM→VLA); structure, supported features, pros/cons.
- 2026-06-10 — RoboMME (in-depth) — memory-implementation breakdown (3 representations × 3 integration mechanisms on π0.5), full results tables, and the FrameSamp+Modulator code analysis.
- 2026-06-08 — ICML 2026 index — 99 manipulation papers in 10 categories (abstract-verified), the 🏅 Top-20 synthesized ranking, and 16 new per-paper pages.
- 2026-06 — Qwen-VLA — the Qwen team's first dedicated VLA.
- 2026-06 — DuoCore-FS — Astribot's parallel fast-slow whole-body stack.
- 2026-06 — Tactile VLA — touch/force-grounded VLAs: architecture × sensor hardware × 2026 trends.
- 2026-06 — World Models — model-centric robotics world-model taxonomy (incl. WM + inverse-dynamics).
- 2026-06 — WAM vs VLA Robustness — first controlled world-model-vs-VLA benchmark.
2026-05
- ICRA 2026 Survey — Vienna; record 5,088 submissions; 728 manip/VLA papers; deployment/sensor-centric (+12 topic analyses).
- CVPR 2026 Survey — Denver; VLA is now the modal manipulation contribution (+29 per-paper pages).
- Genesis GENE-26.5 · AsyncVLA · LBM Co-training (TRI) · Steerable Policies · GR00T N1→N1.7.
- OmniVTA (visuo-tactile WM) — visuo-tactile world model.
- Goal-Image Conditioning · System 0/1/2 · VLA Attention · ML foundations.
Earlier (2026-Q1 and before)
- Per-paper / per-series deep-dives: π0.7 · π0.6 · π Series Evolution · Fast-in-Slow · Discrete Diffusion VLA · VLM4VLA.
- Core topic reviews: VLA Architectures · VLM↔Action Connection · VLA Memory · Dexterous Manipulation · Cross-Embodiment · RL for VLA.
- Venue surveys: NeurIPS 2025 · CoRL 2025 · IROS 2025 · ICLR 2026 · CVPR 2025.
- Foundational references: OpenVLA · ReKep · AgiBot World Colosseo · RoboBrain 2.0.
How to maintain: when adding a page, prepend a dated bullet under the current month. Use exact YYYY-MM-DD when known, else YYYY-MM. Keep Home §2 to the ~6 most recent; everything else lives here.
← Back to Home