Review GR00T Series - Heungwoo/research GitHub Wiki
Author: NVIDIA Β· Papers / reports: GR00T N1 arXiv 2503.14734 (N1 only β N1.5/N1.6/N1.7 have no arXiv paper, only research blogs + HF model cards + GitHub code) Code: https://github.com/NVIDIA/Isaac-GR00T (tags:
n1-release,n1.5-release,n1.6-release,n1.7-release) N1.7 early access announcement: https://forums.developer.nvidia.com/t/early-access-isaac-gr00t-n1-7-open-reasoning-vla-model-for-humanoid-robotics/366916 Built on N1.7: RoboTTT β adds Test-Time-Training fast-weight layers to the N1.7 DiT action head for 8K-timestep context (constant latency).
Architectural details below are read directly from the released code at each git tag β not inferred from blog posts. The key architectural questions answered here are which specific VLM layer feeds the DiT, what preprocessing sits between VLM and DiT, and how those choices evolved across four versions.
GR00T is the most-iterated open-humanoid VLA series of 2025β2026. Each release follows the same high-level structure (dual-system S2 VLM + S1 DiT action head with cross-attention) but varies the VLM backbone, the VLMβDiT connection details, and the training recipe. As of N1.7 (Apr 17, 2026), GR00T became the first Apache-2.0-weights generalist humanoid VLA β production-deployable, unlike prior releases under NVIDIA One-Way Noncommercial.
Companion reviews: VLA Architectures Β· VLMβAction Connection Β· Ο series evolution (the closed-weights counterpart).
Across all four versions, GR00T keeps the same Category C connection (cross-attention from a separate DiT into VLM hidden states) but varies which VLM layer feeds the DiT and what preprocessing sits between them. Four major shifts:
-
N1 β N1.5: Eagle-2 β frozen Eagle-2.5 + FLARE auxiliary loss; mid-layer tap (layer 12) replaced by last-layer tap (
select_layer=-1) after physical truncation. -
N1.5 β N1.6: Eagle-2.5 β Cosmos-Reason-2B (reasoning VLM); DiT depth 16 β 32 layers; adds a
vllnLayerNorm on VLM features before DiT cross-attention; introducestune_top_llm_layersoption (unfreeze top-N LLM layers); introducesAlternateVLDiTβ a VL-version of RDT-1B-style alternating image/text cross-attention. -
N1.6 β N1.7: Cosmos-Reason-derived β Qwen3-VL-2B (Cosmos-Reason2-2B); adds
vl_self_attentionβ a new SelfAttentionTransformer between LayerNorm and DiT cross-attention; ONNX/TensorRT export; Apache-2.0 weights (first in the series). -
Across all four: separate-stack DiT with flow matching; 16-step action chunks; Kβ{4,5} denoising; per-embodiment MLPs with
MultiEmbodimentActionEncoder+CategorySpecificMLP.
License flip at N1.7 is arguably more significant than any single architectural delta.
Training-data thread (detailed in Β§6): N1 established a three-tier data pyramid (web/ego-video β synthetic sim + neural β real teleop). Each version then scaled a different tier hardest: N1/N1.5 synthetic-scaled (DexMimicGen 780k + DreamGen), N1.6 real-teleop-scaled (thousands of hours of YAM/AGIBot/Galaxea/G1), N1.7 human-video-scaled (EgoScale 20,854h action-labeled ego-video, with a log-linear scaling law fit L_val = 0.024 β 0.003Β·ln(D), RΒ²=0.9983).
flowchart LR
subgraph S2[System 2 β VLM backbone]
direction TB
I[Multi-view images 224x224]
L[Language instruction]
VE[Vision encoder<br/>SigLip2 class]
VT[Text tokenizer]
LLM[LLM layers<br/>truncated at select_layer<br/>last-layer output used]
I --> VE --> LLM
L --> VT --> LLM
end
LLM --> BF[backbone features]
BF --> VLLN[vlln LayerNorm]
VLLN --> VLSA[vl_self_attention]
VLSA --> VLE[vl_embeds]
subgraph INP[State and action inputs]
direction TB
S[State history] --> SE[state_encoder<br/>CategorySpecificMLP]
AN[Noised action chunk] --> AE[action_encoder<br/>MultiEmbodimentActionEncoder]
EID[Embodiment ID] --> SE
EID --> AE
end
SE --> CONCAT[concat state + action tokens]
AE --> CONCAT
subgraph S1[System 1 β DiT action head]
direction TB
DIT[DiT blocks<br/>self-attn on state+action<br/>cross-attn on vl_embeds]
AD[action_decoder<br/>CategorySpecificMLP]
DIT --> AD
end
CONCAT --> DIT
VLE -. cross-attention source .-> DIT
AD --> OUT[Action chunk H=16<br/>flow matching K=4 to 5 steps]
classDef new fill:#ffe8c2,stroke:#b47820,color:#000
class VLLN,VLSA new
Reading the diagram:
- S2 (VLM backbone) produces
backbone_featuresfrom the last layer of a physically truncated LLM (theselect_layercodelayers.pop()s all layers past that index β there's no separate "tap" layer). - The orange path between S2 and S1 is where GR00T evolved version-to-version:
vllnLayerNorm added in N1.6,vl_self_attentionSelfAttentionTransformer added in N1.7. N1 and N1.5 feed rawbackbone_featuresdirectly to DiT asencoder_hidden_states. - S1 (DiT action head) does self-attention over concatenated state + action tokens and cross-attention into
vl_embedsinside eachBasicTransformerBlock.N1.6+'sAlternateVLDiToptionally alternates image-only vs. image+text cross-attention across blocks. - The
state_encoder/action_encoder/action_decoderare all per-embodiment (CategorySpecificMLPselects which head byembodiment_id), which is what gives GR00T cross-embodiment support.
The core pattern (separate DiT cross-attending to truncated-VLM features with per-embodiment MLPs on state/action) is unchanged N1 β N1.7. Evolution is in the orange pre-DiT path, the VLM identity, and the training recipe β Β§3 below details each.
This is the most frequently asked architectural question about GR00T. The answer, read from each release tag's code:
All four releases use the same unusual pattern: physically pop VLM LLM layers past select_layer rather than tap an intermediate layer:
N1 (gr00t/model/backbone/eagle_backbone.py, tag n1-release):
class EagleBackbone(nn.Module):
def __init__(self, ..., select_layer: int = 12, ...):
...
while len(self.model.language_model.model.layers) > select_layer:
self.model.language_model.model.layers.pop(-1)-
select_layer=12is the default β this is the famous "mid-layer tap" from the paper - After popping,
hidden_states[-1]is the layer-12 output (there's no layer 13+ anymore)
N1.5 (same file, tag n1.5-release):
select_layer: int = -1, # default changed: last layer after truncation
while len(self.eagle_model.language_model.model.layers) > select_layer:
self.eagle_model.language_model.model.layers.pop(-1)
# ...
eagle_features = eagle_output.hidden_states[self.select_layer]- Default becomes
-1(use the last layer of the truncated model) - Adds explicit
hidden_states[self.select_layer]β can pick any layer of the truncated stack
N1.6 (gr00t/model/modules/eagle_backbone.py, tag n1.6-release):
class EagleBackbone(torch.nn.Module):
def __init__(self, ..., tune_llm=False, select_layer=-1,
tune_top_llm_layers: int = 0, ...):
while len(self.model.language_model.model.layers) > select_layer:
self.model.language_model.model.layers.pop(-1)
# ...
if tune_top_llm_layers > 0:
for layer in self.model.language_model.model.layers[-tune_top_llm_layers:]:
layer.requires_grad_(True)-
New:
tune_top_llm_layersβ unfreeze the top-N LLM layers for joint training (the research blog's "unfroze top 4 VLM layers" detail)
N1.7 (gr00t/model/modules/qwen3_backbone.py):
class Qwen3Backbone(torch.nn.Module):
def __init__(self, ..., select_layer: int = -1, ...):
while len(self.model.language_model.layers) > select_layer:
self.model.language_model.layers.pop(-1)
...
def forward(self, vl_input):
outputs = self.model(**vl_input, output_hidden_states=True)
return outputs.hidden_states[-1]- Same truncation pattern, now on
Qwen3-VL(Cosmos-Reason2-2B) - Same
tune_top_llm_layersknob - Hierarchy
language_model.layers(not.model.layers) reflects Qwen3 structure
Between backbone_features and DiT cross-attention, each version added a new module:
| Version | Preprocessing pipeline |
|---|---|
| N1 |
vl_embeds = backbone_features (raw) |
| N1.5 |
vl_embeds = backbone_features (raw, but with FLARE auxiliary loss on separate path) |
| N1.6 |
vl_embeds = vlln(backbone_features) β LayerNorm added
|
| N1.7 |
vl_embeds = vl_self_attention(vlln(backbone_features)) β LayerNorm + SelfAttentionTransformer
|
The N1.6 Gr00tN1d6ActionHead code:
self.vlln = (nn.LayerNorm(config.backbone_embedding_dim)
if config.use_vlln else nn.Identity())
def process_backbone_output(self, backbone_output):
backbone_features = backbone_output["backbone_features"]
backbone_features = self.vlln(backbone_features)
backbone_output["backbone_features"] = backbone_features
return backbone_outputThe N1.7 Gr00tN1d7ActionHead adds:
self.vlln = nn.LayerNorm(config.backbone_embedding_dim) if config.use_vlln else nn.Identity()
vl_self_attention_cfg = getattr(config, "vl_self_attention_cfg", None)
if vl_self_attention_cfg and vl_self_attention_cfg.get("num_layers", 0) > 0:
self.vl_self_attention = SelfAttentionTransformer(**vl_self_attention_cfg)
else:
self.vl_self_attention = nn.Identity()
def process_backbone_output(self, backbone_output):
backbone_features = backbone_output["backbone_features"]
backbone_features = self.vlln(backbone_features)
backbone_features = self.vl_self_attention(backbone_features)
...SelfAttentionTransformer is a configurable transformer block (num_layers, hidden_size, etc.) that lets the VL features self-mix before being passed to DiT. This is a genuine architectural addition β the equivalent of a Q-former-lite or VL post-processor. It can be turned off by setting num_layers=0 β becomes nn.Identity().
All four versions instantiate the same DiT (or AlternateVLDiT) module with VLM features as encoder_hidden_states:
model_output, _ = self.model(
hidden_states=sa_embs, # state + action tokens
encoder_hidden_states=vl_embeds, # VLM features β CROSS-ATTENTION source
encoder_attention_mask=vl_attn_mask,
timestep=t_discretized,
...
)So across all GR00T versions the interface is:
- Cross-attention from a separate action transformer into VLM hidden states (Category C in Review-VLM-Action-Connection)
- Self-attention over
[state_features, action_features]tokens inside the DiT - Interleaved cross-attention β self-attention in
BasicTransformerBlock
Introduced in N1.6, inherited in N1.7. From gr00t/model/modules/dit.py:
class AlternateVLDiT(DiT):
def __init__(self, *args, attend_text_every_n_blocks: int = 2, **kwargs):
...
def forward(self, ..., image_mask, backbone_attention_mask, ...):
image_attention_mask = image_mask & backbone_attention_mask
non_image_attention_mask = (~image_mask) & backbone_attention_mask
for idx, block in enumerate(self.transformer_blocks):
if idx % (2 * self.attend_text_every_n_blocks) == 0:
# cross-attend to TEXT (+ image)
else:
# cross-attend to IMAGE only
...Translation: most DiT blocks cross-attend to image tokens only; every 2N-th block attends to text (+ image). The rationale is the same as RDT-1B's Alternating Condition Injection: image tokens would drown out text tokens if both were cross-attended every layer. This is a direct architectural cousin of RDT-1B's trick, adapted for VL inputs (rather than just image + language).
-
arXiv: 2503.14734 Β· HF:
nvidia/GR00T-N1-2B -
VLM: Eagle-2 (1.34B); layer 12 of LLM tapped (default
select_layer=12) - Action head: DiT with AdaLN, cross-attention to VLM features, flow matching, H=16, K=4 steps
- Total params: ~2.2B
- Training mix: 88h GR-1 teleop + OpenX + AgiBot-Alpha (140k traj) + 780k DexMimicGen sim traj + 827h "neural trajectories" (video-generation rollouts) + 7 ego-video datasets labeled with VQ-VAE latent actions
- Embodiments: Fourier GR-1, Franka, bimanual Panda, RoboCasa mobile
- Benchmarks: RoboCasa 32.1% (vs. Diffusion Policy 25.6), DexMimicGen 66.5% (56.1), GR-1 Tabletop 50.0%; Real GR-1 10% data: 42.6% avg (DP 10.2); full: 76.8% (46.4)
- Latency: 63.9 ms / 16-action chunk on L40 (bf16)
- License: NVIDIA One-Way Noncommercial (weights), Apache 2.0 (code)
- Signature novelty: dual-system VLA with mid-layer cross-attention + 780k synthetic trajectories
-
HF:
nvidia/GR00T-N1.5-3BΒ· Research blog: https://research.nvidia.com/labs/gear/gr00t-n1_5/ - VLM: Eagle-2.5 (2.1B), frozen during pretrain + finetune
-
Interface: default
select_layer=-1(last layer after truncation); adapter MLP + LayerNorm on VL outputs - New auxiliary: FLARE (Future LAtent Representation Alignment, arXiv 2505.15659) β world-model-style latent alignment loss alongside flow matching
- Training scale: 250k steps on 1k H100s, batch 16,384
- New data: AgiBot-Beta, DreamGen neural trajectories, DexMG
- New embodiments: Unitree G1, SO-100/SO-101
-
Benchmarks vs N1:
- Language Table: 93.2% vs 52.8%
- Real GR-1 language following: 93.3% vs 46.6%
- RoboCasa-30: 47.5 vs 17.4
- Grounding IoU on GR-1: 40.4 (Eagle-2.5) vs 35.5 (Qwen2.5-VL)
- Ecosystem: released with GR00T-Dreams blueprint (36h sim data = 3 months teleop)
- Signature novelty: frozen Eagle-2.5 + FLARE loss (world-model auxiliary); adapter MLP interface replaces mid-layer plumbing
-
HF:
nvidia/GR00T-N1.6-3B,nvidia/GR00T-N1.6-BEHAVIOR1kΒ· Research blog: https://research.nvidia.com/labs/gear/gr00t-n1_6/ - VLM: Cosmos-Reason-2B variant β a reasoning VLM with chain-of-thought on physical-AI scenes
-
Interface: dropped N1.5's post-VLM adapter; unfroze top 4 VLM layers (
tune_top_llm_layers=4); addedvllnLayerNorm on backbone features before DiT cross-attention - DiT: 2Γ deeper β 32 layers (N1.5 had 16)
-
AlternateVLDiTintroduced: alternating image/text cross-attention every2Β·Nblocks (N defaults to 2 β text every 4 blocks) - Action: state-relative action chunks for most embodiments
- Total params: ~3B
- Training: 300k pretrain steps, global batch 16,384; several-thousand hours of teleop from bimanual YAM, AGIBot Genie1, Galaxea R1 Pro, Unitree G1 whole-body, BEHAVIOR sim, DROID
- Ecosystem: Newton physics engine (with DeepMind + Disney) in Isaac Lab 2.3; Cosmos Reason as VLM; DreamGen video-world-model training; COMPASS sim data for navigation
- Benchmarks: no published head-to-head; research blog claims "outperforms N1.5 on sim + YAM/AGIBot/G1"
- Partners evaluating: AeiROBOT, Franka, LG, Lightwheel, Mentee, Neura, Solomon, Techman, UCR; Boston Dynamics Atlas demoed Isaac Lab 2.3
- Signature novelty: reasoning VLA (Cosmos-Reason CoT) + 32-layer DiT + Newton/Isaac/DreamGen co-launch
-
HF:
nvidia/GR00T-N1.7-3B+ task finetunes (DROID, LIBERO, SimplerEnv) Β· Announcement: https://forums.developer.nvidia.com/t/early-access-isaac-gr00t-n1-7-open-reasoning-vla-model-for-humanoid-robotics/366916 -
VLM: Cosmos-Reason2-2B = Qwen3-VL-2B-Instruct (2.44B)
- Native aspect-ratio vision
- Up to 256k-token context
- 2D/3D point localization
- Chain-of-thought reasoning
-
Interface: new
vl_self_attentionβ a SelfAttentionTransformer betweenvllnand DiT cross-attention (configurable byvl_self_attention_cfg.num_layers) -
DiT: class inherited from N1.6 (
AlternateVLDiTavailable) - Data: +20,000 hours of action-labeled egocentric human video from NVIDIA's new EgoScale release (arXiv 2602.16710) β >20Γ prior ego-video efforts; established log-linear scaling law linking human-data volume to validation loss
- Reasoning: "structured reasoning at both task and subtask levels" per NVIDIA's announcement
-
Deployment: full pipeline export to ONNX and TensorRT with higher action frequency;
uv-based installs across dGPU, Jetson (Thor/Orin), DGX Spark - Finetuned checkpoints published: DROID, LIBERO, SimplerEnv variants
- License: Apache 2.0 for BOTH code AND weights β the first GR00T release under a commercial license
- Partners adopting at launch: LG Electronics, NEURA, Noble Machines
- Benchmarks: "comparable performance to N1.6 with improved generalization and language-following" (no head-to-head tables published yet)
-
Signature novelty: commercial-grade open weights + EgoScale human-video scaling +
vl_self_attentionpre-DiT transformer + ONNX/TRT deployment
| Axis | N1 (Mar 2025) | N1.5 (Jun 2025) | N1.6 (Sep 2025) | N1.7 (Apr 2026) |
|---|---|---|---|---|
| VLM backbone | Eagle-2 (1.34B) | Eagle-2.5 (2.1B, frozen) | Cosmos-Reason-2B | Cosmos-Reason2-2B (Qwen3-VL-2B, 2.44B) |
select_layer default |
12 (mid-layer) |
-1 (last after truncation) |
-1 |
-1 |
| VL preprocessing | none (raw) | adapter MLP + LayerNorm | vlln LayerNorm |
vlln + vl_self_attention |
tune_top_llm_layers |
N/A | N/A (frozen) | new knob (default 0, used 4) | inherited |
| Auxiliary loss | flow matching only | + FLARE | flow matching only | flow matching only |
| DiT layers | 12 (default) | 16 | 32 | 32 (inherited) |
AlternateVLDiT |
not available | not available | introduced | inherited |
| Attention scheme | cross-attn to VLM (all blocks) | cross-attn to VLM | cross-attn, optional alternating image/text every 2N blocks | cross-attn, optional alternating |
| Action horizon H | 16 | 16 | 16 | 16 |
| Denoising steps K | 4 | 4β5 | 4β5 | 4β5 |
| Total params | ~2.2B (N1-2B) | ~3B (N1.5-3B) | ~3B (N1.6-3B) | ~3B (N1.7-3B) |
| Embodiments | Fourier GR-1, Franka, bimanual Panda, RoboCasa | + Unitree G1, SO-100/101 | + YAM, AGIBot Genie1, Galaxea R1 Pro, Atlas | + DROID/LIBERO/SimplerEnv finetunes; LG, NEURA, Noble |
| Human-video data | 7 ego-video datasets + VQ-VAE latent actions | inherited | inherited | + 20kh EgoScale (the big data delta) |
| NVIDIA stack | Isaac Sim + DexMimicGen | + GR00T-Dreams | + Newton, Cosmos Reason, Isaac Lab 2.3, COMPASS | + Cosmos-Reason2, EgoScale, ONNX/TRT export |
| Benchmarks | Real GR-1 10%: 42.6; full: 76.8 | LangTable 93.2, RealGR-1 93.3 | "outperforms N1.5" (no table) | "comparable to N1.6" (no table) |
| Weights license | NVIDIA One-Way Noncommercial | One-Way Noncommercial | One-Way Noncommercial | Apache 2.0 β |
| Code license | Apache 2.0 | Apache 2.0 | Apache 2.0 | Apache 2.0 |
| arXiv paper | 2503.14734 | β | β | β |
| Signature novelty | mid-layer cross-attn + 780k synthetic trajs | FLARE + frozen VLM + G1 | Reasoning VLA + 32-layer DiT + Newton/Isaac/DreamGen | Commercial-grade open + EgoScale + vl_self_attention |
This is the second most-asked question about GR00T after VLMβDiT connection: which data source is "the main" training set for each version, and how is it processed? The answer shifts significantly release-to-release β each version bets on a different tier of the same underlying pyramid.
The N1 paper (Β§4) introduced a three-tier blueprint that every subsequent release kept:
flowchart TB
WV["Base: web + human ego-video<br/>7 ego-video datasets<br/>VQ-VAE latent actions<br/>unlabeled, pseudo-actions"]
SIM["Middle: synthetic<br/>780k DexMimicGen sim traj<br/>~6,500h equivalent<br/>+ 827h neural trajectories<br/>action-labeled, cheap to scale"]
ROBOT["Peak: real robot teleop<br/>88h GR-1 teleop<br/>+ OpenX embodiments<br/>+ 140k AgiBot-Alpha<br/>action-labeled, grounded, scarce"]
WV --> SIM --> ROBOT
classDef base fill:#e8f5e9,stroke:#2e7d32,color:#000
classDef mid fill:#fff3e0,stroke:#ef6c00,color:#000
classDef peak fill:#e3f2fd,stroke:#1565c0,color:#000
class WV base
class SIM mid
class ROBOT peak
- Peak (real robot). Smallest, most grounded layer. 88h Fourier GR-1 teleop + Open-X-Embodiment (RT-1, Bridge-v2, Language Table, DROID, MUTEX, RoboSet, Plex) + 140k AgiBot-Alpha trajectories across 100 robots.
-
Middle (synthetic). Two generators:
- DexMimicGen: segments dozens of human teleop demos into object-centric subtasks β aligns with new object positions β interpolates β executes β retains only successful trajectories. Yields 780k demos (~6,500h) across 54 sourceΓtarget receptacle combinations.
-
Neural trajectories: fine-tunes an image-to-video model on 3k real-robot samples with novel language prompts (commercial LLM generates feasible
pick up {object} from {A} to {B}task combos) β 827h of generated video (β10Γ the 88h seed) β commercial LLM filters 8-frame samples for instruction adherence β re-captioning.
-
Base (web + ego-video). Seven public ego-video datasets β Ego4D, Ego-Exo4D, Assembly-101, EPIC-KITCHENS, HOI4D, HoloAssist, RH20T-Human. No action labels, so N1 trains a VQ-VAE-style latent-action encoder on
(x_t, x_{t+H})future-frame pairs and uses the pre-quantized embeddings as pseudo-actions for a dedicated "LAPA" embodiment head.
Co-training recipe: flow-matching loss jointly on all three layers; action-less sources use latent / IDM-predicted pseudo-actions; batches sample heterogeneously across the pyramid. ~50,000 H100-hours for GR00T-N1-2B pretrain on up to 1,024 H100s via NVIDIA OSMO orchestration.
The answer changes every release. Volumes below are volume-dominant, not impact-dominant β the small real-robot peak remains the "golden" set that grounds everything.
| Version | Volume-dominant source | What shifted |
|---|---|---|
| N1 (Mar 2025) | 780k DexMimicGen sim trajectories (~6,500h) dominate by volume; 88h real GR-1 is the grounded peak | Establishes the pyramid; synthetic-scaled |
| N1.5 (Jun 2025) | DreamGen-generated neural trajectories + sim GR-1 (DexMG) β still synthetic-scaled; adds AgiBot-Beta on the real side; new embodiment data (Unitree G1, SO-100/101) | DreamGen (arXiv 2505.12705) replaces N1's ad-hoc video generator; DreamGen blueprint claims 36h sim β 3 months of teleop |
| N1.6 (Sep 2025) | Several thousand hours of real teleop becomes dominant β bimanual YAM + AGIBot Genie1 + Galaxea R1 Pro + Unitree G1 whole-body + BEHAVIOR sim + DROID | Real-teleop-scaled β first release where teleop corpus size is measured in "thousands of hours" |
| N1.7 (Apr 2026) | EgoScale β 20,854h action-labeled egocentric human video becomes the single largest source | Human-video-scaled β ~20Γ every prior human-video effort; robot teleop inherited from N1.6 |
Reading the shift: N1 / N1.5 were synthetic-scaled; N1.6 was real-teleop-scaled; N1.7 is human-video-scaled. Each phase bets on a different tier of the pyramid β the pyramid itself is invariant, only which layer is being scaled hardest changes.
N1.7 is the first GR00T where the data story is the headline, not the architecture. The EgoScale pipeline (arXiv 2602.16710) produces action-labeled human video at unprecedented scale.
- ~20,000h of in-the-wild egocentric recordings β household, industrial, retail, educational (the majority of the 20,854h total)
- 829h of EgoDex β Apple-Vision-Pro-collected ego video with higher-precision hand/wrist tracking; the precise-pose subset used for anchoring
- Total: 20,854 hours β dwarfs Ego4D-scale efforts and makes the base layer of the pyramid the single largest source
flowchart LR
V["Raw ego-video frame"] --> SLAM["Monocular SLAM<br/>T_wβc β SE(3)"]
V --> HPE["Hand-pose estimator<br/>21 keypoints per frame<br/>H_c,i β SE(3)"]
SLAM --> RT["Optimization-based retargeting<br/>CasADi + IPOPT<br/>URDF FK, joint limits,<br/>keypoint-pose error min"]
HPE --> RT
RT --> ACT["Action label<br/>wrist: ΞW_t β SE(3)<br/>hand: 22-DoF joint angles"]
ACT --> SMOOTH["Exponential smoothing"]
SMOOTH --> OUT["Training-ready<br/>action-labeled frames"]
classDef step fill:#f3e5f5,stroke:#6a1b9a,color:#000
class SLAM,HPE,RT,ACT,SMOOTH step
-
Camera motion. Off-the-shelf monocular SLAM recovers per-frame camera extrinsics
T_wβc^t β SE(3). -
Hand pose. Hand-pose estimator outputs 21 keypoints per frame as rigid transforms
H_c,i^tin camera frame. - Retargeting. Nonlinear optimization (CasADi + IPOPT) maps 21 human keypoints β dexterous-robot-hand joint angles, subject to URDF-based FK, joint limits, and keypoint-pose error minimization.
-
Action representation.
-
Wrist: relative end-effector SE(3) motion
ΞW_tbetween consecutive frames - Hand: 22-DoF joint angles on the default dexterous hand
-
Wrist: relative end-effector SE(3) motion
- Smoothing. Exponential smoothing over the optimized trajectory. No explicit filtering thresholds for noisy SLAM / pose estimates are reported β the scaling-law fit suggests the quantity-over-quality bet pays off.
| Stage | Data | Steps | Batch | LR | Unfrozen |
|---|---|---|---|---|---|
| I β Human pretrain | 20,854h EgoScale | 100k | 8,192 | 5e-5 | All params |
| II β Aligned mid-training | 50h human play + 4h robot play (matched scenes, calibrated cameras, Vive wrist + Manus glove GT) | 50k | 2,048 | 3e-5 | Vision encoder + DiT; VLM backbone frozen |
| III β Post-training | Task-specific robot demos (~100 trajectories per task) | 10k | 512 | 3e-5 | Task head |
Why the three stages? The key novelty is Stage II's aligned mid-training β a tiny dataset (50h human + 4h robot) where humans and the robot perform the same tabletop task in the same scene under matched camera intrinsics. Ground-truth wrist pose from Vive trackers (3D position + orientation); ground-truth hand pose from Manus gloves (25 joint transforms). This bridges the human β robot distribution gap before the tiny post-training set (Stage III) sees the task. Without Stage II, the massive Stage-I human prior would need to cross the embodiment gap during Stage III's ~100-trajectory post-train β too little signal.
EgoScale fits a log-linear relationship between validation loss and human-pretrain hours across the 1kβ20k-hour range:
with RΒ² = 0.9983. The paper claims the fit "strongly correlates with downstream real-robot completion scores" β the quantitative basis for NVIDIA's bet on human-video scaling. Independent replication has not happened yet (see Β§9).
Continuity threads:
- DexMimicGen sim (from N1+) and DreamGen neural trajectories (from N1.5+) remain in every subsequent mix
- Seven ego-video datasets from N1 stay in the pretrain β EgoScale augments, not replaces
- The data-pyramid blueprint itself is invariant; only the layer being scaled hardest shifts
Rupture points:
- N1.5's DreamGen introduction displaces N1's ad-hoc video model with a first-class world-model-based trajectory generator
- N1.6's shift to thousand-hour teleop when partner-collected data finally reached VLA-scale
- N1.7's EgoScale redefines what the "base layer" means: from passive internet ego-video (weak context via VQ-VAE latents) β 20kh of action-labeled ego-video as a first-class training signal with its own scaling law
The N1.7 shift is the central narrative: NVIDIA frames it as the first robotics-scaling law where human-video hours is the x-axis β analogous to LLM compute/data scaling, but with an embodied validation-loss y-axis. Whether this scales to 200kh remains the open research question for post-N1.7.
From Review-VLM-Action-Connection: GR00T is the canonical Category C (cross-attention) VLA. Specifically:
-
Cross-attention layer choice: GR00T physically truncates the VLM at
select_layer, so "cross-attending to layer N" = "cross-attending to the last layer of a truncated stack." This is a different pattern from the paper description "cross-attention at layer 12 of Eagle-2" β the paper's wording is post-hoc; the code actually pops layers 13+. -
N1's
select_layer=12default aligns the code with the paper. N1.5+ moved to full-length VLM with-1default. -
N1.6 added
vllnLayerNorm β normalizing VLM features before cross-attention is a small but empirically-motivated hygiene step (analog of T5's pre-norm). -
N1.7 added
vl_self_attentionβ a configurable extra transformer that lets VL features self-mix before feeding DiT. Think of it as a "VL post-processor" or a lightweight Q-former-style module that can be enabled/disabled at config time. -
AlternateVLDiT(N1.6+) directly mirrors RDT-1B's Alternating Condition Injection β image tokens would drown out text without alternation.
Comparison to other Category C VLAs:
| Paper | Cross-attn target | Special handling |
|---|---|---|
| GR00T N1 | Last layer of truncated (at 12) Eagle-2 | None β raw vl_embeds
|
| GR00T N1.6/N1.7 | Last layer of truncated VLM | + LayerNorm (N1.6), + self-attention block (N1.7), + alternating image/text (optional) |
| RDT-1B | Separate SigLIP + T5-XXL | Alternating Condition Injection (image and text cross-attended on alternate layers) |
| ST4VLA (ICLR 2026) | k intermediate VLM layers | Query transformer over multiple layers |
GR00T N1's original mid-layer tap (layer 12) was empirically motivated in the paper for "higher policy success and lower inference latency" β a data point for the "which VLM layer?" open question flagged in Review-VLM-Action-Connection Β§8. The N1.5+ shift to last-layer tap (with physical truncation) suggests NVIDIA didn't find mid-layer clearly better in subsequent experiments.
GR00T changed VLM backbone three times in 13 months: Eagle-2 β Eagle-2.5 β Cosmos-Reason β Cosmos-Reason2. The rationale per release:
- Eagle-2 β Eagle-2.5: better multimodal pretraining; enables freezing (N1 had to unfreeze LLM; N1.5's frozen backbone is only possible with stronger pretrain)
- Eagle-2.5 β Cosmos-Reason: chain-of-thought capability β the headline N1.6 shift from "VLM that reads scenes" to "VLM that reasons about physics"
- Cosmos-Reason β Cosmos-Reason2 (Qwen3-VL): 256k context, native aspect-ratio, 2D/3D point localization β the N1.7 reasoning VLA gets stronger grounding
This is very different from the Ο-series, which did one backbone swap (PaliGemma β Gemma3-4B at Ο0.6) and froze afterward. GR00T is the "NVIDIA VLM-of-the-month" series in a good sense: every open VLM NVIDIA releases becomes a GR00T backbone candidate.
- No published head-to-head benchmark tables for N1.6 or N1.7 against predecessors. NVIDIA's claims are qualitative ("outperforms N1.5", "comparable to N1.6 with improved generalization"). Rigorous comparison is impossible without this.
- No arXiv paper for N1.5, N1.6, N1.7 β information is scattered across HF model cards, research blogs, forum posts, and GitHub READMEs.
- HF model cards for N1.5/N1.6/N1.7 are template-cloned and all incorrectly list "SigLip2 + T5" β research blogs are the authoritative source for backbone identity.
- Category C (cross-attention) locks in a latency overhead vs. the Ο-series' same-stack MoE. GR00T's 63.9 ms / chunk (N1 on L40) is competitive but the architectural cost of a separate DiT is real.
- Mid-layer vs. last-layer choice is unresolved β N1 used layer 12, N1.5+ abandoned it. No ablation published explaining the shift.
- EgoScale scaling claim needs independent validation. 20kh β validation-loss log-linear fit is a strong claim; third-party replication has not happened yet.
- "Apache 2.0 weights" at N1.7 is early-access only. The actual license-file check in HF model cards should be verified before commercial deployment.
- No head-to-head on contact-rich or dexterous tasks with the latest Ο-series (Ο0.6, Ο0.7) or DDVLA on LIBERO. Public benchmarking is limited to NVIDIA's own suites.
-
vl_self_attentionin N1.7 has no public ablation β NVIDIA adds architectural components release-by-release but doesn't publish which ones matter. A matched-compute ablation of vlln + vl_self_attention + AlternateVLDiT across the three versions is the ablation the community needs.
- Most-iterated open humanoid VLA of 2025β2026. Four releases in 13 months with public code; the Ο-series had five but with closed weights.
- Reference dual-system architecture. "GR00T N1's System-2/System-1 split has become reference architecture" β cited in Review-VLA-Architecture, WholeBodyVLA, HiMoE-VLA, ReinFlow (which runs RL on GR00T-N1.5), and VLA-0 (NVIDIA, arXiv 2510.13054).
- Ecosystem integration. Newton (physics), Cosmos (world model), DreamGen (video rollouts), Isaac Lab, COMPASS, Jetson Thor β no competitor has this stack breadth.
- License flip at N1.7 is the most commercially significant 2026 VLA event. Until Apr 17, 2026, every production-grade generalist VLA was closed-weights (Ο0.6/Ο0.7, Gemini Robotics). GR00T N1.7 breaks that.
- VLM backbone churn as a feature. Every ~5 months NVIDIA swaps in a stronger VLM. This means GR00T benefits from general VLM progress faster than single-backbone-frozen systems like the Ο-series.
| Axis | GR00T | Ο-series |
|---|---|---|
| Developer | NVIDIA (40+ authors) | Physical Intelligence |
| Weights | Apache 2.0 (N1.7) / Noncommercial (N1βN1.6) | Closed (model cards only) |
| VLMβAction interface | Category C cross-attn into truncated VLM | Category B same-stack MoE with prefix-KV |
| Action head | Flow-matching DiT, H=16, K=4 | Flow-matching expert, H=50, K=5β10 |
| Backbone path | Eagle-2 β Eagle-2.5 β Cosmos-Reason β Cosmos-Reason2 | PaliGemma-3B β Gemma3-4B (stable since Ο0.6) |
| Gradient insulation | (tune_top_llm_layers + freeze) | Knowledge Insulation (2505.23705) |
| Humanoid focus | Primary (Fourier GR-1, Unitree G1, Atlas) | Manipulation-centric; UR5e zero-shot |
| RL integration | via DreamGen / Isaac Lab recipes | RECAP (Ο*0.6) + distillation into Ο0.7 |
| Ecosystem | Newton, Cosmos, DreamGen, Isaac Lab, Jetson | OpenPi (open code for Ο0/Ο0.5-KI) |
| Signature novelty per release | Backbone swaps + stack integration | Algorithmic depth (KI, MEM, RTC, RECAP, metadata CFG) |
As of April 2026: GR00T N1.7 is the first credible generalist humanoid VLA with Apache 2.0 weights AND the NVIDIA inference stack. The Ο-series remains closed but has a stronger per-release algorithmic story. Different strengths: GR00T for humanoid breadth + ecosystem + commercial deployment; Ο-series for production-tested single-model depth.
flowchart LR
A[Read arXiv 2503.14734<br/>N1 paper] --> B[Then research blogs<br/>N1.5 + N1.6]
B --> C[Then this review Β§3<br/>for code-level connection evolution]
C --> D[Then inspect git tags<br/>n1.5-release / n1.6-release / n1.7-release]
D --> E[Finally pair with EgoScale<br/>arXiv 2602.16710 for N1.7 data]
- arXiv N1 paper: https://arxiv.org/abs/2503.14734 Β· HTML v1: https://arxiv.org/html/2503.14734v1
- NVIDIA Research β N1 publication: https://research.nvidia.com/publication/2025-03_nvidia-isaac-gr00t-n1-open-foundation-model-humanoid-robots
- Research blog β N1.5: https://research.nvidia.com/labs/gear/gr00t-n1_5/
- Research blog β N1.6: https://research.nvidia.com/labs/gear/gr00t-n1_6/
- Research blog β EgoScale: https://research.nvidia.com/labs/gear/egoscale/
- N1.7 Early Access forum post (Apr 17, 2026): https://forums.developer.nvidia.com/t/early-access-isaac-gr00t-n1-7-open-reasoning-vla-model-for-humanoid-robotics/366916
-
GitHub: https://github.com/NVIDIA/Isaac-GR00T (tags
n1-release,n1.5-release,n1.6-release,n1.7-release) -
HF checkpoints:
nvidia/GR00T-N1-2BΒ·nvidia/GR00T-N1.5-3BΒ·nvidia/GR00T-N1.6-3BΒ·nvidia/GR00T-N1.6-BEHAVIOR1kΒ·nvidia/GR00T-N1.7-3B(+ DROID/LIBERO/SimplerEnv finetunes) - Newsroom: N1 launch Β· Computex 2025 humanoid Β· CoRL 2025 N1.6 + Newton
- FLARE (used in N1.5): https://arxiv.org/abs/2505.15659
- DreamGen (used in N1.5+): https://arxiv.org/abs/2505.12705 Β· DreamGen
- EgoScale (used in N1.7): https://arxiv.org/abs/2602.16710
- Cosmos-Reason2-2B: https://huggingface.co/nvidia/Cosmos-Reason2-2B
- The Robot Report on Newton + N1.6: https://www.therobotreport.com/nvidia-launches-newton-physics-engine-gr00t-ai-corl-2025/
- Ο series evolution β the closed-weights counterpart
- Review-pi06 Β· Review-pi07 β per-paper Ο-series reviews
- Review-VLM-Action-Connection β Β§4.3 (Cross-attention into VLM hidden states) β GR00T is the canonical exemplar
- Review-VLA-Architecture β Β§5.F (Hierarchical / dual-system / MoE) β GR00T's S2/S1 split as reference
- DreamGen β used in N1.5/N1.6 training pipelines
- WholeBodyVLA Β· HiMoE-VLA β hierarchical humanoid VLAs inheriting GR00T's decomposition
- ReinFlow β online RL fine-tune runs on GR00T-N1.5