Review VLA Training Frameworks - Heungwoo/research GitHub Wiki

VLA Training Frameworks β€” StarVLA & TRI VLA Foundry

In-depth analysis of open-source VLA training frameworks: overall architecture, what they support, and their pros/cons. Snapshot 2026-06-10. Covered here: StarVLA (community, "Lego-like") and TRI VLA Foundry (Toyota Research Institute, "LLM→VLM→VLA in one stack"). Note: the request called the TRI project "VLA Factory" — its actual name is VLA Foundry (TRI-ML/vla_foundry).

Why a "training framework" (vs a single model repo)

Most VLA papers ship a model repo β€” one architecture, one action head, one benchmark harness, often forked from OpenVLA/openpi and mutually incompatible. A framework instead factors VLA training into reusable axes (backbone Β· action head Β· data Β· training recipe Β· eval/deploy) so you can swap one axis without rewriting the rest. The two frameworks below attack this from opposite ends: StarVLA maximizes modular breadth at the action-modeling stage, while VLA Foundry owns the whole pretraining-to-action pipeline in one codebase.

TL;DR comparison

Axis StarVLA TRI VLA Foundry
Origin Community ("StarVLA Community") Β· arXiv 2604.05014 Β· β˜…β‰ˆ2.8k Toyota Research Institute Β· arXiv 2604.19728 Β· β˜…β‰ˆ0.4k
Core idea Lego: decoupled backbone βŠ• action-head Foundry: end-to-end LLMβ†’VLMβ†’VLA in one stack
Scope The action-modeling stage (assumes a pretrained VLM/WM backbone) The whole pipeline, incl. language/VLM pretraining from scratch
Action paradigms FAST (AR) Β· OFT (MLP) Β· Ο€ flow-matching Β· GR00T dual-system Β· discrete-diffusion Primarily diffusion (diffusion-transformer / U-Net, LBM-style)
Backbones Qwen-VL/Qwen3.5 (0.8–9B), Gemma, MiniCPM, Florence-2, world models (Cosmos-Predict2, Wan2.2) From-scratch transformers (11Mβ†’3B), PaliGemma/SmolVLM ViTs, Qwen3-VL, HF weights
Infra / scale DeepSpeed ZeRO-3, GPU + Ascend NPU; single-A100 friendly FSDP2 + WebDataset, torchrun + AWS SageMaker, S3 streaming
Benchmarks LIBERO(+plus), SimplerEnv, RoboTwin 2.0, RoboCasa-GR1/365, BEHAVIOR-1K, CALVIN, DOMINO, VLA-Arena, VLN-CE LBM Eval (open-data sim) + STEP analysis tools
Deployment Policy server + interface, sim and real robot (Franka) gRPC policy server, LBM deployment guide
Best for Prototyping/comparing many VLA designs across many benchmarks Reproducible, scalable full-pipeline training of diffusion LBM-style policies

StarVLA

Repo: starVLA/starVLA (branches: starVLA stable / starVLA_dev active) Β· Paper: arXiv 2604.05014 Β· β˜…β‰ˆ2,786 (2026-06)

"A Lego-like Codebase for VLA Model Developing" β€” every functional component (model, data, trainer, config, eval) follows high-cohesion/low-coupling separation, so building a new VLA reduces to swapping the backbone or the action head while reusing everything else.

Overall structure

The defining abstraction is a two-axis decomposition: backbone βŠ• action-head, both independently swappable under a shared data/training interface.

starVLA/
β”œβ”€β”€ model/
β”‚   β”œβ”€β”€ framework/
β”‚   β”‚   β”œβ”€β”€ VLM4A/   # "VLM-for-Action" backbones Γ— heads:
β”‚   β”‚   β”‚            #   QwenPI, QwenFast, QwenOFT, QwenGR00T, QwenDual,
β”‚   β”‚   β”‚            #   QwenDiscreteDiffusion, Gemma4PI, MiniCPMPI/GR00T,
β”‚   β”‚   β”‚            #   ABot_M0, LangForce, M1, QwenAdapter
β”‚   β”‚   └── WM4A/    # "World-Model-for-Action": CosmoPredict2{PI,OFT,GR00T},
β”‚   β”‚                #   Wan{PI,OFT,GR00T}  (video-gen DiT as backbone)
β”‚   └── modules/action_model/   # interchangeable ACTION HEADS:
β”‚        # MLP, DiT, GR00T, LayerwiseFM (flow-matching),
β”‚        # LayerwiseDiscreteDiffusion, VLA_Adapter, AML
β”œβ”€β”€ dataloader/     # LeRobot v3.0, gr00t_lerobot, qwenvl/llava-json
β”‚                   # β†’ returns a raw, MODEL-AGNOSTIC dict {image, lang, action}
β”œβ”€β”€ training/       # train_starvlm.py Β· train_starvla.py Β· train_starvla_cotrain.py
└── config/         # YAML-driven

A single file model/framework/VLM4A/<YourFramework>.py is the only external API surface of a model and is meant to be structurally isomorphic to the paper's framework figure. Each submodule (model, dataloader) can be run standalone for smoke-testing (python starVLA/model/framework/VLM4A/QwenOFT.py --config_yaml ...).

What it supports

  • Four headline variants (same data interface, only the head differs): StarVLA-FAST (AR discrete tokens, Ο€β‚€-fast-style), StarVLA-OFT (parallel continuous MLP, OpenVLA-OFT-style), StarVLA-PI (flow-matching expert, Ο€β‚€-style), StarVLA-GR00T (dual-system: VLM=System-2, flow-matching=System-1).
  • Backbones: Qwen-VL & Qwen3.5 (0.8B/2B/4B/9B), Gemma, MiniCPM, Florence-2 (small enough for a single A100), and world-model backbones via WM4A (Cosmos-Predict2, Wan2.2 video-gen DiTs as action predictors).
  • Training recipes (paradigm-agnostic): SFT Β· multimodal multi-objective co-training Β· cross-embodiment co-training Β· RL post-training (via the RLinf integration). It was the first OSS repo to offer train-your-VLM, train-your-VLA, and co-train-VLA-with-VLM entrypoints.
  • Benchmarks (unified eval interface): LIBERO, LIBERO-plus, SimplerEnv, RoboTwin 2.0, RoboCasa-GR1, RoboCasa365, BEHAVIOR-1K, CALVIN, DOMINO, VLA-Arena, VLN-CE β€” plus real-robot deployment (Franka) via a policy server/interface.
  • Infra niceties: DeepSpeed ZeRO-3, LeRobot v3.0, Ascend NPU support for Qwen backbones, AllenAI's vla-evaluation-harness for fast eval, and agent-skills docs (StarVLA is explicitly engineered to be drivable by AI coding agents).

Pros / cons

Pros

  • Widest design-space coverage of any OSS VLA framework: 4+ action paradigms Γ— many backbones Γ— 11 benchmarks β€” ideal for apples-to-apples method comparison.
  • True modularity β€” backbone and head are genuinely independent; adding a new method is often one file.
  • World-model-as-backbone (WM4A) is rare and forward-looking; small-VLM (Florence-2) path lowers the GPU bar to a single A100.
  • Strong community momentum (β˜…2.8k, near-daily updates, RL/NPU/new-backbone PRs) and reproducible single-benchmark recipes that match/beat prior work.

Cons

  • Research-grade velocity: starVLA_dev is explicitly "may be temporarily unstable"; pin the stable starVLA branch for results.
  • Action-stage-only scope: it consumes pretrained VLM/WM backbones β€” it does not pretrain an LLM/VLM from scratch in-framework.
  • RL is integration-based (RLinf), not a first-class native recipe (the core checklist still marks RL adaptation as in-progress).
  • Correctness/eval depends on many heavy external benchmark deps; breadth is a maintenance/setup cost.

TRI VLA Foundry

Repo: TRI-ML/vla_foundry Β· Paper: arXiv 2604.19728 Β· Weights: HF TRI-ML/vla-foundry Β· β˜…β‰ˆ395 (2026-06)

"A Unified Framework for Training Vision-Language-Action Models" β€” the differentiator is end-to-end ownership of the whole stack: train an LLM, use that checkpoint to train a VLM, use that to train a VLA β€” all in one codebase with no external pretraining dependencies.

Overall structure

vla_foundry/
β”œβ”€β”€ main.py / train.py        # single entrypoint, torchrun / SageMaker
β”œβ”€β”€ models/                   # pure-PyTorch transformer/ViT/diffusion blocks
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ pipelines/            # text Β· image_caption Β· robotics  (unified)
β”‚   β”œβ”€β”€ preprocessing/        # β†’ WebDataset .tar shards; LeRobot converter
β”‚   └── dataloader.py         # dataset MIXING (sources + ratios + balancing)
β”œβ”€β”€ eval/ + inference/        # LBM Eval sim, STEP tools, gRPC policy server
β”œβ”€β”€ config_presets/
β”‚   β”œβ”€β”€ models/   # transformer_{tiny,11m,100m,410m,1b}, vlm_{11m..3b},
β”‚   β”‚             #   vit_paligemma / vit_smolvlm2, vla_diffusion_{11m..1b},
β”‚   β”‚             #   diffusion_policy, unet, qwen_05b
β”‚   β”œβ”€β”€ data/     # lbm/* (4/6-camera, action-fields, language-annotations,
β”‚   β”‚             #   1past_14future / 5past_20future horizons)
β”‚   └── training_jobs/        # lbm_multitask_4cams, diffusion_policy_lbm1, ...
β”œβ”€β”€ aws/ + sagemaker/         # cluster launch, S3 remote_sync
└── distributed.py            # FSDP2

The same main.py runs all three stages by switching --data.type between text, image_caption, and robotics β€” and chaining checkpoints (include directives compose model configs).

What it supports

  • One codebase, three stages: LLM pretraining (text) β†’ VLM training (image-caption) β†’ VLA fine-tuning (robotics), each seeding the next via checkpoints. Also supports starting from HF pretrained backbones (PaliGemma, SmolVLM2, Qwen3-VL) instead of from scratch.
  • Action head = diffusion (LBM-style diffusion transformer / U-Net / diffusion-policy presets) β€” the TRI Large-Behavior-Model lineage.
  • Scales by config: model presets from 11M β†’ 3B (transformer/vlm/vla-diffusion), multi-camera (4/6-cam) robot configs, configurable observation/action horizons.
  • Production-grade infra: FSDP2 sharding, WebDataset streaming from S3, torchrun for local multi-GPU and AWS SageMaker for clusters, remote_sync of checkpoints. Dataset mixing (per-source ratios + batch balancing) at dataloading time.
  • Pure PyTorch, minimal deps β€” most modules avoid external libraries to keep the pipeline hackable; HF interop is opt-in.
  • Eval & deploy: LBM Eval (open-data, open-source simulator) + STEP statistical analysis tools + gRPC policy server. Two released model families: a fully-open from-scratch LLMβ†’VLMβ†’VLA model (on par with TRI's prior closed-source LBM) and a Qwen3-VL-backbone model that beats the baseline on multi-task tabletop manipulation.

Pros / cons

Pros

  • Only framework with true end-to-end LLMβ†’VLMβ†’VLA in one stack β€” no stitching of incompatible pretraining repos; full control from language pretraining to action expert.
  • Built for scale & reproducibility: FSDP2 + WebDataset + SageMaker/S3 is genuine multi-node production infra, and LBM Eval uses open data so closed-loop results are reproducible.
  • Pure-PyTorch, low-dependency design is easy to read, modify, and audit; clean dataset-mixing primitives.
  • Backed by TRI's LBM research line (industrial rigor; the from-scratch model matches their prior closed-source work).

Cons

  • Diffusion-centric: action modeling is essentially the LBM diffusion family β€” less paradigm variety than StarVLA (no FAST/OFT/dual-system menu out of the box).
  • AWS/SageMaker-leaning workflow (S3 manifests, secrets.env, SageMaker launcher) raises the setup bar for non-cloud users.
  • Narrower eval surface: centered on LBM Eval / tabletop manipulation, vs StarVLA's 10+ external benchmarks.
  • Smaller community (β˜…β‰ˆ0.4k) and newer, so fewer third-party backbones/recipes today.

Side-by-side: which to choose

If you want to… Use
Compare many action heads / backbones on many benchmarks StarVLA
Plug in a world-model backbone (Cosmos/Wan) or run on a single A100 / Ascend NPU StarVLA
Reproduce a published VLA method fast, or prototype a new head in one file StarVLA
Pretrain LLM→VLM→VLA end-to-end in one controlled stack VLA Foundry
Train diffusion (LBM-style) policies at scale on a cluster (FSDP2/SageMaker/S3) VLA Foundry
Get reproducible closed-loop eval on open data (LBM Eval) VLA Foundry

Complementary, not competing. StarVLA optimizes the horizontal axis (breadth of VLA designs at the action stage); VLA Foundry optimizes the vertical axis (depth of one pipeline from pretraining to action, at scale). A plausible workflow: pretrain/scale a backbone in VLA Foundry, then prototype action-head variants and cross-benchmark comparisons in StarVLA.

Links

Related pages


πŸ—“ State of the Field (updated Aug 2026)

Verdict: modular (StarVLA) vs pipeline (TRI VLA Foundry) remains the base choice; RL-for-VLA infrastructure is the new third tier β€” and nothing yet covers the full lifecycle.

πŸ“ˆ Trend

RSS 2026 added the tier both incumbents lacked: RL-for-VLA frameworks (RLux-VLA β€” unified platform for fair algorithm Γ— architecture comparison), mirroring the field's pivot from training policies to improving them ([RL]]). Evaluation is likewise becoming framework-shaped ([PolaRiS scene-builder + sharing hub).

βš–οΈ Approaches & trade-offs

Framework class Pros Cons
Lego-modular (StarVLA) Fast research iteration, component swap Not a production pipeline
Full pipeline (TRI VLA Foundry) LLM→VLM→VLA reproducibility Heavier, opinionated
RL platforms (RLux-VLA) Fair RL comparisons, rollout infra New; narrow scope
Benchmarking frameworks ([VLA-Arena](/Heungwoo/research/wiki/ICML-2026-VLA-Arena), [CaP-X](/Heungwoo/research/wiki/ICML-2026-CaP-X) β€” ICML 2026) Structured perturbation diagnostics (170 tasks, L0–L2); coding-agent evaluation Evaluation-side only

⚠️ Limitations & open problems

  • No public framework covers pretrain β†’ SFT β†’ RL-from-deployment β†’ OOD evaluation end-to-end.
  • Flagship recipes (Ο€, Qwen suite) are irreproducible in any open stack β€” weights withheld, hyperparameters partial.
  • Fragmentation makes cross-paper numbers only loosely comparable.

← Back to Home