Review VLA Training Frameworks - Heungwoo/research GitHub Wiki
VLA Training Frameworks β StarVLA & TRI VLA Foundry
In-depth analysis of open-source VLA training frameworks: overall architecture, what they support, and their pros/cons. Snapshot 2026-06-10. Covered here: StarVLA (community, "Lego-like") and TRI VLA Foundry (Toyota Research Institute, "LLMβVLMβVLA in one stack"). Note: the request called the TRI project "VLA Factory" β its actual name is VLA Foundry (
TRI-ML/vla_foundry).
Why a "training framework" (vs a single model repo)
Most VLA papers ship a model repo β one architecture, one action head, one benchmark harness, often forked from OpenVLA/openpi and mutually incompatible. A framework instead factors VLA training into reusable axes (backbone Β· action head Β· data Β· training recipe Β· eval/deploy) so you can swap one axis without rewriting the rest. The two frameworks below attack this from opposite ends: StarVLA maximizes modular breadth at the action-modeling stage, while VLA Foundry owns the whole pretraining-to-action pipeline in one codebase.
TL;DR comparison
| Axis | StarVLA | TRI VLA Foundry |
|---|---|---|
| Origin | Community ("StarVLA Community") Β· arXiv 2604.05014 Β· β β2.8k | Toyota Research Institute Β· arXiv 2604.19728 Β· β β0.4k |
| Core idea | Lego: decoupled backbone β action-head | Foundry: end-to-end LLMβVLMβVLA in one stack |
| Scope | The action-modeling stage (assumes a pretrained VLM/WM backbone) | The whole pipeline, incl. language/VLM pretraining from scratch |
| Action paradigms | FAST (AR) Β· OFT (MLP) Β· Ο flow-matching Β· GR00T dual-system Β· discrete-diffusion | Primarily diffusion (diffusion-transformer / U-Net, LBM-style) |
| Backbones | Qwen-VL/Qwen3.5 (0.8β9B), Gemma, MiniCPM, Florence-2, world models (Cosmos-Predict2, Wan2.2) | From-scratch transformers (11Mβ3B), PaliGemma/SmolVLM ViTs, Qwen3-VL, HF weights |
| Infra / scale | DeepSpeed ZeRO-3, GPU + Ascend NPU; single-A100 friendly | FSDP2 + WebDataset, torchrun + AWS SageMaker, S3 streaming |
| Benchmarks | LIBERO(+plus), SimplerEnv, RoboTwin 2.0, RoboCasa-GR1/365, BEHAVIOR-1K, CALVIN, DOMINO, VLA-Arena, VLN-CE | LBM Eval (open-data sim) + STEP analysis tools |
| Deployment | Policy server + interface, sim and real robot (Franka) | gRPC policy server, LBM deployment guide |
| Best for | Prototyping/comparing many VLA designs across many benchmarks | Reproducible, scalable full-pipeline training of diffusion LBM-style policies |
StarVLA
Repo: starVLA/starVLA (branches: starVLA stable / starVLA_dev active) Β· Paper: arXiv 2604.05014 Β· β
β2,786 (2026-06)
"A Lego-like Codebase for VLA Model Developing" β every functional component (model, data, trainer, config, eval) follows high-cohesion/low-coupling separation, so building a new VLA reduces to swapping the backbone or the action head while reusing everything else.
Overall structure
The defining abstraction is a two-axis decomposition: backbone β action-head, both independently swappable under a shared data/training interface.
starVLA/
βββ model/
β βββ framework/
β β βββ VLM4A/ # "VLM-for-Action" backbones Γ heads:
β β β # QwenPI, QwenFast, QwenOFT, QwenGR00T, QwenDual,
β β β # QwenDiscreteDiffusion, Gemma4PI, MiniCPMPI/GR00T,
β β β # ABot_M0, LangForce, M1, QwenAdapter
β β βββ WM4A/ # "World-Model-for-Action": CosmoPredict2{PI,OFT,GR00T},
β β # Wan{PI,OFT,GR00T} (video-gen DiT as backbone)
β βββ modules/action_model/ # interchangeable ACTION HEADS:
β # MLP, DiT, GR00T, LayerwiseFM (flow-matching),
β # LayerwiseDiscreteDiffusion, VLA_Adapter, AML
βββ dataloader/ # LeRobot v3.0, gr00t_lerobot, qwenvl/llava-json
β # β returns a raw, MODEL-AGNOSTIC dict {image, lang, action}
βββ training/ # train_starvlm.py Β· train_starvla.py Β· train_starvla_cotrain.py
βββ config/ # YAML-driven
A single file model/framework/VLM4A/<YourFramework>.py is the only external API surface of a model and is meant to be structurally isomorphic to the paper's framework figure. Each submodule (model, dataloader) can be run standalone for smoke-testing (python starVLA/model/framework/VLM4A/QwenOFT.py --config_yaml ...).
What it supports
- Four headline variants (same data interface, only the head differs): StarVLA-FAST (AR discrete tokens, Οβ-fast-style), StarVLA-OFT (parallel continuous MLP, OpenVLA-OFT-style), StarVLA-PI (flow-matching expert, Οβ-style), StarVLA-GR00T (dual-system: VLM=System-2, flow-matching=System-1).
- Backbones: Qwen-VL & Qwen3.5 (0.8B/2B/4B/9B), Gemma, MiniCPM, Florence-2 (small enough for a single A100), and world-model backbones via WM4A (Cosmos-Predict2, Wan2.2 video-gen DiTs as action predictors).
- Training recipes (paradigm-agnostic): SFT Β· multimodal multi-objective co-training Β· cross-embodiment co-training Β· RL post-training (via the RLinf integration). It was the first OSS repo to offer train-your-VLM, train-your-VLA, and co-train-VLA-with-VLM entrypoints.
- Benchmarks (unified eval interface): LIBERO, LIBERO-plus, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, RoboCasa365, BEHAVIOR-1K, CALVIN, DOMINO, VLA-Arena, VLN-CE β plus real-robot deployment (Franka) via a policy server/interface.
- Infra niceties: DeepSpeed ZeRO-3, LeRobot v3.0, Ascend NPU support for Qwen backbones, AllenAI's
vla-evaluation-harnessfor fast eval, and agent-skills docs (StarVLA is explicitly engineered to be drivable by AI coding agents).
Pros / cons
Pros
- Widest design-space coverage of any OSS VLA framework: 4+ action paradigms Γ many backbones Γ 11 benchmarks β ideal for apples-to-apples method comparison.
- True modularity β backbone and head are genuinely independent; adding a new method is often one file.
- World-model-as-backbone (WM4A) is rare and forward-looking; small-VLM (Florence-2) path lowers the GPU bar to a single A100.
- Strong community momentum (β 2.8k, near-daily updates, RL/NPU/new-backbone PRs) and reproducible single-benchmark recipes that match/beat prior work.
Cons
- Research-grade velocity:
starVLA_devis explicitly "may be temporarily unstable"; pin the stablestarVLAbranch for results. - Action-stage-only scope: it consumes pretrained VLM/WM backbones β it does not pretrain an LLM/VLM from scratch in-framework.
- RL is integration-based (RLinf), not a first-class native recipe (the core checklist still marks RL adaptation as in-progress).
- Correctness/eval depends on many heavy external benchmark deps; breadth is a maintenance/setup cost.
TRI VLA Foundry
Repo: TRI-ML/vla_foundry Β· Paper: arXiv 2604.19728 Β· Weights: HF TRI-ML/vla-foundry Β· β
β395 (2026-06)
"A Unified Framework for Training Vision-Language-Action Models" β the differentiator is end-to-end ownership of the whole stack: train an LLM, use that checkpoint to train a VLM, use that to train a VLA β all in one codebase with no external pretraining dependencies.
Overall structure
vla_foundry/
βββ main.py / train.py # single entrypoint, torchrun / SageMaker
βββ models/ # pure-PyTorch transformer/ViT/diffusion blocks
βββ data/
β βββ pipelines/ # text Β· image_caption Β· robotics (unified)
β βββ preprocessing/ # β WebDataset .tar shards; LeRobot converter
β βββ dataloader.py # dataset MIXING (sources + ratios + balancing)
βββ eval/ + inference/ # LBM Eval sim, STEP tools, gRPC policy server
βββ config_presets/
β βββ models/ # transformer_{tiny,11m,100m,410m,1b}, vlm_{11m..3b},
β β # vit_paligemma / vit_smolvlm2, vla_diffusion_{11m..1b},
β β # diffusion_policy, unet, qwen_05b
β βββ data/ # lbm/* (4/6-camera, action-fields, language-annotations,
β β # 1past_14future / 5past_20future horizons)
β βββ training_jobs/ # lbm_multitask_4cams, diffusion_policy_lbm1, ...
βββ aws/ + sagemaker/ # cluster launch, S3 remote_sync
βββ distributed.py # FSDP2
The same main.py runs all three stages by switching --data.type between text, image_caption, and robotics β and chaining checkpoints (include directives compose model configs).
What it supports
- One codebase, three stages: LLM pretraining (text) β VLM training (image-caption) β VLA fine-tuning (robotics), each seeding the next via checkpoints. Also supports starting from HF pretrained backbones (PaliGemma, SmolVLM2, Qwen3-VL) instead of from scratch.
- Action head = diffusion (LBM-style diffusion transformer / U-Net / diffusion-policy presets) β the TRI Large-Behavior-Model lineage.
- Scales by config: model presets from 11M β 3B (transformer/vlm/vla-diffusion), multi-camera (4/6-cam) robot configs, configurable observation/action horizons.
- Production-grade infra: FSDP2 sharding, WebDataset streaming from S3,
torchrunfor local multi-GPU and AWS SageMaker for clusters,remote_syncof checkpoints. Dataset mixing (per-source ratios + batch balancing) at dataloading time. - Pure PyTorch, minimal deps β most modules avoid external libraries to keep the pipeline hackable; HF interop is opt-in.
- Eval & deploy: LBM Eval (open-data, open-source simulator) + STEP statistical analysis tools + gRPC policy server. Two released model families: a fully-open from-scratch LLMβVLMβVLA model (on par with TRI's prior closed-source LBM) and a Qwen3-VL-backbone model that beats the baseline on multi-task tabletop manipulation.
Pros / cons
Pros
- Only framework with true end-to-end LLMβVLMβVLA in one stack β no stitching of incompatible pretraining repos; full control from language pretraining to action expert.
- Built for scale & reproducibility: FSDP2 + WebDataset + SageMaker/S3 is genuine multi-node production infra, and LBM Eval uses open data so closed-loop results are reproducible.
- Pure-PyTorch, low-dependency design is easy to read, modify, and audit; clean dataset-mixing primitives.
- Backed by TRI's LBM research line (industrial rigor; the from-scratch model matches their prior closed-source work).
Cons
- Diffusion-centric: action modeling is essentially the LBM diffusion family β less paradigm variety than StarVLA (no FAST/OFT/dual-system menu out of the box).
- AWS/SageMaker-leaning workflow (S3 manifests,
secrets.env, SageMaker launcher) raises the setup bar for non-cloud users. - Narrower eval surface: centered on LBM Eval / tabletop manipulation, vs StarVLA's 10+ external benchmarks.
- Smaller community (β β0.4k) and newer, so fewer third-party backbones/recipes today.
Side-by-side: which to choose
| If you want to⦠| Use |
|---|---|
| Compare many action heads / backbones on many benchmarks | StarVLA |
| Plug in a world-model backbone (Cosmos/Wan) or run on a single A100 / Ascend NPU | StarVLA |
| Reproduce a published VLA method fast, or prototype a new head in one file | StarVLA |
| Pretrain LLMβVLMβVLA end-to-end in one controlled stack | VLA Foundry |
| Train diffusion (LBM-style) policies at scale on a cluster (FSDP2/SageMaker/S3) | VLA Foundry |
| Get reproducible closed-loop eval on open data (LBM Eval) | VLA Foundry |
Complementary, not competing. StarVLA optimizes the horizontal axis (breadth of VLA designs at the action stage); VLA Foundry optimizes the vertical axis (depth of one pipeline from pretraining to action, at scale). A plausible workflow: pretrain/scale a backbone in VLA Foundry, then prototype action-head variants and cross-benchmark comparisons in StarVLA.
Links
- StarVLA: repo Β· arXiv 2604.05014 Β· WM4A doc
- VLA Foundry: repo Β· arXiv 2604.19728 Β· docs
- Related TRI work in this wiki: LBM Co-training Study (TRI)
Related pages
- VLA Architectures Β· RL for VLA Β· Cross-Embodiment
- World Models (StarVLA's WM4A backbones)
π State of the Field (updated Aug 2026)
Verdict: modular (StarVLA) vs pipeline (TRI VLA Foundry) remains the base choice; RL-for-VLA infrastructure is the new third tier β and nothing yet covers the full lifecycle.
π Trend
RSS 2026 added the tier both incumbents lacked: RL-for-VLA frameworks (RLux-VLA β unified platform for fair algorithm Γ architecture comparison), mirroring the field's pivot from training policies to improving them ([RL]]). Evaluation is likewise becoming framework-shaped ([PolaRiS scene-builder + sharing hub).
βοΈ Approaches & trade-offs
| Framework class | Pros | Cons |
|---|---|---|
| Lego-modular (StarVLA) | Fast research iteration, component swap | Not a production pipeline |
| Full pipeline (TRI VLA Foundry) | LLMβVLMβVLA reproducibility | Heavier, opinionated |
| RL platforms (RLux-VLA) | Fair RL comparisons, rollout infra | New; narrow scope |
| Benchmarking frameworks ([VLA-Arena](/Heungwoo/research/wiki/ICML-2026-VLA-Arena), [CaP-X](/Heungwoo/research/wiki/ICML-2026-CaP-X) β ICML 2026) | Structured perturbation diagnostics (170 tasks, L0βL2); coding-agent evaluation | Evaluation-side only |
β οΈ Limitations & open problems
- No public framework covers pretrain β SFT β RL-from-deployment β OOD evaluation end-to-end.
- Flagship recipes (Ο, Qwen suite) are irreproducible in any open stack β weights withheld, hyperparameters partial.
- Fragmentation makes cross-paper numbers only loosely comparable.
β Back to Home