Review GR00T Series - Heungwoo/research GitHub Wiki

In-Depth Review β€” NVIDIA Isaac GR00T N1 β†’ N1.5 β†’ N1.6 β†’ N1.7

Author: NVIDIA Β· Papers / reports: GR00T N1 arXiv 2503.14734 (N1 only β€” N1.5/N1.6/N1.7 have no arXiv paper, only research blogs + HF model cards + GitHub code) Code: https://github.com/NVIDIA/Isaac-GR00T (tags: n1-release, n1.5-release, n1.6-release, n1.7-release) N1.7 early access announcement: https://forums.developer.nvidia.com/t/early-access-isaac-gr00t-n1-7-open-reasoning-vla-model-for-humanoid-robotics/366916 Built on N1.7: RoboTTT β€” adds Test-Time-Training fast-weight layers to the N1.7 DiT action head for 8K-timestep context (constant latency).

Architectural details below are read directly from the released code at each git tag β€” not inferred from blog posts. The key architectural questions answered here are which specific VLM layer feeds the DiT, what preprocessing sits between VLM and DiT, and how those choices evolved across four versions.

πŸ“Ž Context β€” GR00T in the VLA landscape

GR00T is the most-iterated open-humanoid VLA series of 2025–2026. Each release follows the same high-level structure (dual-system S2 VLM + S1 DiT action head with cross-attention) but varies the VLM backbone, the VLMβ†’DiT connection details, and the training recipe. As of N1.7 (Apr 17, 2026), GR00T became the first Apache-2.0-weights generalist humanoid VLA β€” production-deployable, unlike prior releases under NVIDIA One-Way Noncommercial.

Companion reviews: VLA Architectures Β· VLM↔Action Connection Β· Ο€ series evolution (the closed-weights counterpart).


1. TL;DR

Across all four versions, GR00T keeps the same Category C connection (cross-attention from a separate DiT into VLM hidden states) but varies which VLM layer feeds the DiT and what preprocessing sits between them. Four major shifts:

  1. N1 β†’ N1.5: Eagle-2 β†’ frozen Eagle-2.5 + FLARE auxiliary loss; mid-layer tap (layer 12) replaced by last-layer tap (select_layer=-1) after physical truncation.
  2. N1.5 β†’ N1.6: Eagle-2.5 β†’ Cosmos-Reason-2B (reasoning VLM); DiT depth 16 β†’ 32 layers; adds a vlln LayerNorm on VLM features before DiT cross-attention; introduces tune_top_llm_layers option (unfreeze top-N LLM layers); introduces AlternateVLDiT β€” a VL-version of RDT-1B-style alternating image/text cross-attention.
  3. N1.6 β†’ N1.7: Cosmos-Reason-derived β†’ Qwen3-VL-2B (Cosmos-Reason2-2B); adds vl_self_attention β€” a new SelfAttentionTransformer between LayerNorm and DiT cross-attention; ONNX/TensorRT export; Apache-2.0 weights (first in the series).
  4. Across all four: separate-stack DiT with flow matching; 16-step action chunks; K∈{4,5} denoising; per-embodiment MLPs with MultiEmbodimentActionEncoder + CategorySpecificMLP.

License flip at N1.7 is arguably more significant than any single architectural delta.

Training-data thread (detailed in Β§6): N1 established a three-tier data pyramid (web/ego-video β†’ synthetic sim + neural β†’ real teleop). Each version then scaled a different tier hardest: N1/N1.5 synthetic-scaled (DexMimicGen 780k + DreamGen), N1.6 real-teleop-scaled (thousands of hours of YAM/AGIBot/Galaxea/G1), N1.7 human-video-scaled (EgoScale 20,854h action-labeled ego-video, with a log-linear scaling law fit L_val = 0.024 βˆ’ 0.003Β·ln(D), RΒ²=0.9983).


2. The unifying architecture β€” dual-system VLA

flowchart LR
  subgraph S2[System 2 β€” VLM backbone]
    direction TB
    I[Multi-view images 224x224]
    L[Language instruction]
    VE[Vision encoder<br/>SigLip2 class]
    VT[Text tokenizer]
    LLM[LLM layers<br/>truncated at select_layer<br/>last-layer output used]
    I --> VE --> LLM
    L --> VT --> LLM
  end

  LLM --> BF[backbone features]
  BF --> VLLN[vlln LayerNorm]
  VLLN --> VLSA[vl_self_attention]
  VLSA --> VLE[vl_embeds]

  subgraph INP[State and action inputs]
    direction TB
    S[State history] --> SE[state_encoder<br/>CategorySpecificMLP]
    AN[Noised action chunk] --> AE[action_encoder<br/>MultiEmbodimentActionEncoder]
    EID[Embodiment ID] --> SE
    EID --> AE
  end

  SE --> CONCAT[concat state + action tokens]
  AE --> CONCAT

  subgraph S1[System 1 β€” DiT action head]
    direction TB
    DIT[DiT blocks<br/>self-attn on state+action<br/>cross-attn on vl_embeds]
    AD[action_decoder<br/>CategorySpecificMLP]
    DIT --> AD
  end

  CONCAT --> DIT
  VLE -. cross-attention source .-> DIT
  AD --> OUT[Action chunk H=16<br/>flow matching K=4 to 5 steps]

  classDef new fill:#ffe8c2,stroke:#b47820,color:#000
  class VLLN,VLSA new
Loading

Reading the diagram:

  • S2 (VLM backbone) produces backbone_features from the last layer of a physically truncated LLM (the select_layer code layers.pop()s all layers past that index β€” there's no separate "tap" layer).
  • The orange path between S2 and S1 is where GR00T evolved version-to-version: vlln LayerNorm added in N1.6, vl_self_attention SelfAttentionTransformer added in N1.7. N1 and N1.5 feed raw backbone_features directly to DiT as encoder_hidden_states.
  • S1 (DiT action head) does self-attention over concatenated state + action tokens and cross-attention into vl_embeds inside each BasicTransformerBlock. N1.6+'s AlternateVLDiT optionally alternates image-only vs. image+text cross-attention across blocks.
  • The state_encoder / action_encoder / action_decoder are all per-embodiment (CategorySpecificMLP selects which head by embodiment_id), which is what gives GR00T cross-embodiment support.

The core pattern (separate DiT cross-attending to truncated-VLM features with per-embodiment MLPs on state/action) is unchanged N1 β†’ N1.7. Evolution is in the orange pre-DiT path, the VLM identity, and the training recipe β€” Β§3 below details each.


3. VLM→Action connection — code-level evolution

This is the most frequently asked architectural question about GR00T. The answer, read from each release tag's code:

3.1 Layer selection β€” truncate, don't tap

All four releases use the same unusual pattern: physically pop VLM LLM layers past select_layer rather than tap an intermediate layer:

N1 (gr00t/model/backbone/eagle_backbone.py, tag n1-release):

class EagleBackbone(nn.Module):
    def __init__(self, ..., select_layer: int = 12, ...):
        ...
        while len(self.model.language_model.model.layers) > select_layer:
            self.model.language_model.model.layers.pop(-1)
  • select_layer=12 is the default β€” this is the famous "mid-layer tap" from the paper
  • After popping, hidden_states[-1] is the layer-12 output (there's no layer 13+ anymore)

N1.5 (same file, tag n1.5-release):

select_layer: int = -1,       # default changed: last layer after truncation
while len(self.eagle_model.language_model.model.layers) > select_layer:
    self.eagle_model.language_model.model.layers.pop(-1)
# ...
eagle_features = eagle_output.hidden_states[self.select_layer]
  • Default becomes -1 (use the last layer of the truncated model)
  • Adds explicit hidden_states[self.select_layer] β€” can pick any layer of the truncated stack

N1.6 (gr00t/model/modules/eagle_backbone.py, tag n1.6-release):

class EagleBackbone(torch.nn.Module):
    def __init__(self, ..., tune_llm=False, select_layer=-1,
                 tune_top_llm_layers: int = 0, ...):
        while len(self.model.language_model.model.layers) > select_layer:
            self.model.language_model.model.layers.pop(-1)
        # ...
        if tune_top_llm_layers > 0:
            for layer in self.model.language_model.model.layers[-tune_top_llm_layers:]:
                layer.requires_grad_(True)
  • New: tune_top_llm_layers β€” unfreeze the top-N LLM layers for joint training (the research blog's "unfroze top 4 VLM layers" detail)

N1.7 (gr00t/model/modules/qwen3_backbone.py):

class Qwen3Backbone(torch.nn.Module):
    def __init__(self, ..., select_layer: int = -1, ...):
        while len(self.model.language_model.layers) > select_layer:
            self.model.language_model.layers.pop(-1)
        ...
    def forward(self, vl_input):
        outputs = self.model(**vl_input, output_hidden_states=True)
        return outputs.hidden_states[-1]
  • Same truncation pattern, now on Qwen3-VL (Cosmos-Reason2-2B)
  • Same tune_top_llm_layers knob
  • Hierarchy language_model.layers (not .model.layers) reflects Qwen3 structure

3.2 Preprocessing between VLM and DiT

Between backbone_features and DiT cross-attention, each version added a new module:

Version Preprocessing pipeline
N1 vl_embeds = backbone_features (raw)
N1.5 vl_embeds = backbone_features (raw, but with FLARE auxiliary loss on separate path)
N1.6 vl_embeds = vlln(backbone_features) β€” LayerNorm added
N1.7 vl_embeds = vl_self_attention(vlln(backbone_features)) β€” LayerNorm + SelfAttentionTransformer

The N1.6 Gr00tN1d6ActionHead code:

self.vlln = (nn.LayerNorm(config.backbone_embedding_dim)
             if config.use_vlln else nn.Identity())

def process_backbone_output(self, backbone_output):
    backbone_features = backbone_output["backbone_features"]
    backbone_features = self.vlln(backbone_features)
    backbone_output["backbone_features"] = backbone_features
    return backbone_output

The N1.7 Gr00tN1d7ActionHead adds:

self.vlln = nn.LayerNorm(config.backbone_embedding_dim) if config.use_vlln else nn.Identity()

vl_self_attention_cfg = getattr(config, "vl_self_attention_cfg", None)
if vl_self_attention_cfg and vl_self_attention_cfg.get("num_layers", 0) > 0:
    self.vl_self_attention = SelfAttentionTransformer(**vl_self_attention_cfg)
else:
    self.vl_self_attention = nn.Identity()

def process_backbone_output(self, backbone_output):
    backbone_features = backbone_output["backbone_features"]
    backbone_features = self.vlln(backbone_features)
    backbone_features = self.vl_self_attention(backbone_features)
    ...

SelfAttentionTransformer is a configurable transformer block (num_layers, hidden_size, etc.) that lets the VL features self-mix before being passed to DiT. This is a genuine architectural addition β€” the equivalent of a Q-former-lite or VL post-processor. It can be turned off by setting num_layers=0 β†’ becomes nn.Identity().

3.3 DiT cross-attention mechanism

All four versions instantiate the same DiT (or AlternateVLDiT) module with VLM features as encoder_hidden_states:

model_output, _ = self.model(
    hidden_states=sa_embs,              # state + action tokens
    encoder_hidden_states=vl_embeds,    # VLM features β†’ CROSS-ATTENTION source
    encoder_attention_mask=vl_attn_mask,
    timestep=t_discretized,
    ...
)

So across all GR00T versions the interface is:

  • Cross-attention from a separate action transformer into VLM hidden states (Category C in Review-VLM-Action-Connection)
  • Self-attention over [state_features, action_features] tokens inside the DiT
  • Interleaved cross-attention ↔ self-attention in BasicTransformerBlock

3.4 AlternateVLDiT β€” GR00T's answer to RDT-1B's "Alternating Condition Injection"

Introduced in N1.6, inherited in N1.7. From gr00t/model/modules/dit.py:

class AlternateVLDiT(DiT):
    def __init__(self, *args, attend_text_every_n_blocks: int = 2, **kwargs):
        ...

    def forward(self, ..., image_mask, backbone_attention_mask, ...):
        image_attention_mask = image_mask & backbone_attention_mask
        non_image_attention_mask = (~image_mask) & backbone_attention_mask
        for idx, block in enumerate(self.transformer_blocks):
            if idx % (2 * self.attend_text_every_n_blocks) == 0:
                # cross-attend to TEXT (+ image)
            else:
                # cross-attend to IMAGE only
            ...

Translation: most DiT blocks cross-attend to image tokens only; every 2N-th block attends to text (+ image). The rationale is the same as RDT-1B's Alternating Condition Injection: image tokens would drown out text tokens if both were cross-attended every layer. This is a direct architectural cousin of RDT-1B's trick, adapted for VL inputs (rather than just image + language).


4. Version-by-version details

4.1 GR00T N1 (March 18, 2025)

  • arXiv: 2503.14734 Β· HF: nvidia/GR00T-N1-2B
  • VLM: Eagle-2 (1.34B); layer 12 of LLM tapped (default select_layer=12)
  • Action head: DiT with AdaLN, cross-attention to VLM features, flow matching, H=16, K=4 steps
  • Total params: ~2.2B
  • Training mix: 88h GR-1 teleop + OpenX + AgiBot-Alpha (140k traj) + 780k DexMimicGen sim traj + 827h "neural trajectories" (video-generation rollouts) + 7 ego-video datasets labeled with VQ-VAE latent actions
  • Embodiments: Fourier GR-1, Franka, bimanual Panda, RoboCasa mobile
  • Benchmarks: RoboCasa 32.1% (vs. Diffusion Policy 25.6), DexMimicGen 66.5% (56.1), GR-1 Tabletop 50.0%; Real GR-1 10% data: 42.6% avg (DP 10.2); full: 76.8% (46.4)
  • Latency: 63.9 ms / 16-action chunk on L40 (bf16)
  • License: NVIDIA One-Way Noncommercial (weights), Apache 2.0 (code)
  • Signature novelty: dual-system VLA with mid-layer cross-attention + 780k synthetic trajectories

4.2 GR00T N1.5 (June 11, 2025)

  • HF: nvidia/GR00T-N1.5-3B Β· Research blog: https://research.nvidia.com/labs/gear/gr00t-n1_5/
  • VLM: Eagle-2.5 (2.1B), frozen during pretrain + finetune
  • Interface: default select_layer=-1 (last layer after truncation); adapter MLP + LayerNorm on VL outputs
  • New auxiliary: FLARE (Future LAtent Representation Alignment, arXiv 2505.15659) β€” world-model-style latent alignment loss alongside flow matching
  • Training scale: 250k steps on 1k H100s, batch 16,384
  • New data: AgiBot-Beta, DreamGen neural trajectories, DexMG
  • New embodiments: Unitree G1, SO-100/SO-101
  • Benchmarks vs N1:
    • Language Table: 93.2% vs 52.8%
    • Real GR-1 language following: 93.3% vs 46.6%
    • RoboCasa-30: 47.5 vs 17.4
    • Grounding IoU on GR-1: 40.4 (Eagle-2.5) vs 35.5 (Qwen2.5-VL)
  • Ecosystem: released with GR00T-Dreams blueprint (36h sim data = 3 months teleop)
  • Signature novelty: frozen Eagle-2.5 + FLARE loss (world-model auxiliary); adapter MLP interface replaces mid-layer plumbing

4.3 GR00T N1.6 (CoRL 2025, Sept 29, 2025)

  • HF: nvidia/GR00T-N1.6-3B, nvidia/GR00T-N1.6-BEHAVIOR1k Β· Research blog: https://research.nvidia.com/labs/gear/gr00t-n1_6/
  • VLM: Cosmos-Reason-2B variant β€” a reasoning VLM with chain-of-thought on physical-AI scenes
  • Interface: dropped N1.5's post-VLM adapter; unfroze top 4 VLM layers (tune_top_llm_layers=4); added vlln LayerNorm on backbone features before DiT cross-attention
  • DiT: 2Γ— deeper β†’ 32 layers (N1.5 had 16)
  • AlternateVLDiT introduced: alternating image/text cross-attention every 2Β·N blocks (N defaults to 2 β†’ text every 4 blocks)
  • Action: state-relative action chunks for most embodiments
  • Total params: ~3B
  • Training: 300k pretrain steps, global batch 16,384; several-thousand hours of teleop from bimanual YAM, AGIBot Genie1, Galaxea R1 Pro, Unitree G1 whole-body, BEHAVIOR sim, DROID
  • Ecosystem: Newton physics engine (with DeepMind + Disney) in Isaac Lab 2.3; Cosmos Reason as VLM; DreamGen video-world-model training; COMPASS sim data for navigation
  • Benchmarks: no published head-to-head; research blog claims "outperforms N1.5 on sim + YAM/AGIBot/G1"
  • Partners evaluating: AeiROBOT, Franka, LG, Lightwheel, Mentee, Neura, Solomon, Techman, UCR; Boston Dynamics Atlas demoed Isaac Lab 2.3
  • Signature novelty: reasoning VLA (Cosmos-Reason CoT) + 32-layer DiT + Newton/Isaac/DreamGen co-launch

4.4 GR00T N1.7 (April 17, 2026, early access)

  • HF: nvidia/GR00T-N1.7-3B + task finetunes (DROID, LIBERO, SimplerEnv) Β· Announcement: https://forums.developer.nvidia.com/t/early-access-isaac-gr00t-n1-7-open-reasoning-vla-model-for-humanoid-robotics/366916
  • VLM: Cosmos-Reason2-2B = Qwen3-VL-2B-Instruct (2.44B)
    • Native aspect-ratio vision
    • Up to 256k-token context
    • 2D/3D point localization
    • Chain-of-thought reasoning
  • Interface: new vl_self_attention β€” a SelfAttentionTransformer between vlln and DiT cross-attention (configurable by vl_self_attention_cfg.num_layers)
  • DiT: class inherited from N1.6 (AlternateVLDiT available)
  • Data: +20,000 hours of action-labeled egocentric human video from NVIDIA's new EgoScale release (arXiv 2602.16710) β€” >20Γ— prior ego-video efforts; established log-linear scaling law linking human-data volume to validation loss
  • Reasoning: "structured reasoning at both task and subtask levels" per NVIDIA's announcement
  • Deployment: full pipeline export to ONNX and TensorRT with higher action frequency; uv-based installs across dGPU, Jetson (Thor/Orin), DGX Spark
  • Finetuned checkpoints published: DROID, LIBERO, SimplerEnv variants
  • License: Apache 2.0 for BOTH code AND weights β€” the first GR00T release under a commercial license
  • Partners adopting at launch: LG Electronics, NEURA, Noble Machines
  • Benchmarks: "comparable performance to N1.6 with improved generalization and language-following" (no head-to-head tables published yet)
  • Signature novelty: commercial-grade open weights + EgoScale human-video scaling + vl_self_attention pre-DiT transformer + ONNX/TRT deployment

5. Comparison table

Axis N1 (Mar 2025) N1.5 (Jun 2025) N1.6 (Sep 2025) N1.7 (Apr 2026)
VLM backbone Eagle-2 (1.34B) Eagle-2.5 (2.1B, frozen) Cosmos-Reason-2B Cosmos-Reason2-2B (Qwen3-VL-2B, 2.44B)
select_layer default 12 (mid-layer) -1 (last after truncation) -1 -1
VL preprocessing none (raw) adapter MLP + LayerNorm vlln LayerNorm vlln + vl_self_attention
tune_top_llm_layers N/A N/A (frozen) new knob (default 0, used 4) inherited
Auxiliary loss flow matching only + FLARE flow matching only flow matching only
DiT layers 12 (default) 16 32 32 (inherited)
AlternateVLDiT not available not available introduced inherited
Attention scheme cross-attn to VLM (all blocks) cross-attn to VLM cross-attn, optional alternating image/text every 2N blocks cross-attn, optional alternating
Action horizon H 16 16 16 16
Denoising steps K 4 4–5 4–5 4–5
Total params ~2.2B (N1-2B) ~3B (N1.5-3B) ~3B (N1.6-3B) ~3B (N1.7-3B)
Embodiments Fourier GR-1, Franka, bimanual Panda, RoboCasa + Unitree G1, SO-100/101 + YAM, AGIBot Genie1, Galaxea R1 Pro, Atlas + DROID/LIBERO/SimplerEnv finetunes; LG, NEURA, Noble
Human-video data 7 ego-video datasets + VQ-VAE latent actions inherited inherited + 20kh EgoScale (the big data delta)
NVIDIA stack Isaac Sim + DexMimicGen + GR00T-Dreams + Newton, Cosmos Reason, Isaac Lab 2.3, COMPASS + Cosmos-Reason2, EgoScale, ONNX/TRT export
Benchmarks Real GR-1 10%: 42.6; full: 76.8 LangTable 93.2, RealGR-1 93.3 "outperforms N1.5" (no table) "comparable to N1.6" (no table)
Weights license NVIDIA One-Way Noncommercial One-Way Noncommercial One-Way Noncommercial Apache 2.0 βœ…
Code license Apache 2.0 Apache 2.0 Apache 2.0 Apache 2.0
arXiv paper 2503.14734 β€” β€” β€”
Signature novelty mid-layer cross-attn + 780k synthetic trajs FLARE + frozen VLM + G1 Reasoning VLA + 32-layer DiT + Newton/Isaac/DreamGen Commercial-grade open + EgoScale + vl_self_attention

6. Training data β€” version-by-version evolution

This is the second most-asked question about GR00T after VLM→DiT connection: which data source is "the main" training set for each version, and how is it processed? The answer shifts significantly release-to-release — each version bets on a different tier of the same underlying pyramid.

6.1 The N1 "data pyramid" blueprint (inherited by all versions)

The N1 paper (Β§4) introduced a three-tier blueprint that every subsequent release kept:

flowchart TB
  WV["Base: web + human ego-video<br/>7 ego-video datasets<br/>VQ-VAE latent actions<br/>unlabeled, pseudo-actions"]
  SIM["Middle: synthetic<br/>780k DexMimicGen sim traj<br/>~6,500h equivalent<br/>+ 827h neural trajectories<br/>action-labeled, cheap to scale"]
  ROBOT["Peak: real robot teleop<br/>88h GR-1 teleop<br/>+ OpenX embodiments<br/>+ 140k AgiBot-Alpha<br/>action-labeled, grounded, scarce"]
  WV --> SIM --> ROBOT
  classDef base fill:#e8f5e9,stroke:#2e7d32,color:#000
  classDef mid fill:#fff3e0,stroke:#ef6c00,color:#000
  classDef peak fill:#e3f2fd,stroke:#1565c0,color:#000
  class WV base
  class SIM mid
  class ROBOT peak
Loading
  • Peak (real robot). Smallest, most grounded layer. 88h Fourier GR-1 teleop + Open-X-Embodiment (RT-1, Bridge-v2, Language Table, DROID, MUTEX, RoboSet, Plex) + 140k AgiBot-Alpha trajectories across 100 robots.
  • Middle (synthetic). Two generators:
    • DexMimicGen: segments dozens of human teleop demos into object-centric subtasks β†’ aligns with new object positions β†’ interpolates β†’ executes β†’ retains only successful trajectories. Yields 780k demos (~6,500h) across 54 sourceΓ—target receptacle combinations.
    • Neural trajectories: fine-tunes an image-to-video model on 3k real-robot samples with novel language prompts (commercial LLM generates feasible pick up {object} from {A} to {B} task combos) β†’ 827h of generated video (β‰ˆ10Γ— the 88h seed) β†’ commercial LLM filters 8-frame samples for instruction adherence β†’ re-captioning.
  • Base (web + ego-video). Seven public ego-video datasets β€” Ego4D, Ego-Exo4D, Assembly-101, EPIC-KITCHENS, HOI4D, HoloAssist, RH20T-Human. No action labels, so N1 trains a VQ-VAE-style latent-action encoder on (x_t, x_{t+H}) future-frame pairs and uses the pre-quantized embeddings as pseudo-actions for a dedicated "LAPA" embodiment head.

Co-training recipe: flow-matching loss jointly on all three layers; action-less sources use latent / IDM-predicted pseudo-actions; batches sample heterogeneously across the pyramid. ~50,000 H100-hours for GR00T-N1-2B pretrain on up to 1,024 H100s via NVIDIA OSMO orchestration.

6.2 Which source is "the main data" per version?

The answer changes every release. Volumes below are volume-dominant, not impact-dominant β€” the small real-robot peak remains the "golden" set that grounds everything.

Version Volume-dominant source What shifted
N1 (Mar 2025) 780k DexMimicGen sim trajectories (~6,500h) dominate by volume; 88h real GR-1 is the grounded peak Establishes the pyramid; synthetic-scaled
N1.5 (Jun 2025) DreamGen-generated neural trajectories + sim GR-1 (DexMG) β€” still synthetic-scaled; adds AgiBot-Beta on the real side; new embodiment data (Unitree G1, SO-100/101) DreamGen (arXiv 2505.12705) replaces N1's ad-hoc video generator; DreamGen blueprint claims 36h sim β‰ˆ 3 months of teleop
N1.6 (Sep 2025) Several thousand hours of real teleop becomes dominant β€” bimanual YAM + AGIBot Genie1 + Galaxea R1 Pro + Unitree G1 whole-body + BEHAVIOR sim + DROID Real-teleop-scaled β€” first release where teleop corpus size is measured in "thousands of hours"
N1.7 (Apr 2026) EgoScale β€” 20,854h action-labeled egocentric human video becomes the single largest source Human-video-scaled β€” ~20Γ— every prior human-video effort; robot teleop inherited from N1.6

Reading the shift: N1 / N1.5 were synthetic-scaled; N1.6 was real-teleop-scaled; N1.7 is human-video-scaled. Each phase bets on a different tier of the pyramid β€” the pyramid itself is invariant, only which layer is being scaled hardest changes.

6.3 N1.7 training-data pipeline in detail (EgoScale)

N1.7 is the first GR00T where the data story is the headline, not the architecture. The EgoScale pipeline (arXiv 2602.16710) produces action-labeled human video at unprecedented scale.

6.3.1 Source composition

  • ~20,000h of in-the-wild egocentric recordings β€” household, industrial, retail, educational (the majority of the 20,854h total)
  • 829h of EgoDex β€” Apple-Vision-Pro-collected ego video with higher-precision hand/wrist tracking; the precise-pose subset used for anchoring
  • Total: 20,854 hours β€” dwarfs Ego4D-scale efforts and makes the base layer of the pyramid the single largest source

6.3.2 Action-labeling pipeline

flowchart LR
  V["Raw ego-video frame"] --> SLAM["Monocular SLAM<br/>T_w←c ∈ SE(3)"]
  V --> HPE["Hand-pose estimator<br/>21 keypoints per frame<br/>H_c,i ∈ SE(3)"]
  SLAM --> RT["Optimization-based retargeting<br/>CasADi + IPOPT<br/>URDF FK, joint limits,<br/>keypoint-pose error min"]
  HPE --> RT
  RT --> ACT["Action label<br/>wrist: Ξ”W_t ∈ SE(3)<br/>hand: 22-DoF joint angles"]
  ACT --> SMOOTH["Exponential smoothing"]
  SMOOTH --> OUT["Training-ready<br/>action-labeled frames"]
  classDef step fill:#f3e5f5,stroke:#6a1b9a,color:#000
  class SLAM,HPE,RT,ACT,SMOOTH step
Loading
  1. Camera motion. Off-the-shelf monocular SLAM recovers per-frame camera extrinsics T_w←c^t ∈ SE(3).
  2. Hand pose. Hand-pose estimator outputs 21 keypoints per frame as rigid transforms H_c,i^t in camera frame.
  3. Retargeting. Nonlinear optimization (CasADi + IPOPT) maps 21 human keypoints β†’ dexterous-robot-hand joint angles, subject to URDF-based FK, joint limits, and keypoint-pose error minimization.
  4. Action representation.
    • Wrist: relative end-effector SE(3) motion Ξ”W_t between consecutive frames
    • Hand: 22-DoF joint angles on the default dexterous hand
  5. Smoothing. Exponential smoothing over the optimized trajectory. No explicit filtering thresholds for noisy SLAM / pose estimates are reported β€” the scaling-law fit suggests the quantity-over-quality bet pays off.

6.3.3 Three-stage training recipe

Stage Data Steps Batch LR Unfrozen
I β€” Human pretrain 20,854h EgoScale 100k 8,192 5e-5 All params
II β€” Aligned mid-training 50h human play + 4h robot play (matched scenes, calibrated cameras, Vive wrist + Manus glove GT) 50k 2,048 3e-5 Vision encoder + DiT; VLM backbone frozen
III β€” Post-training Task-specific robot demos (~100 trajectories per task) 10k 512 3e-5 Task head

Why the three stages? The key novelty is Stage II's aligned mid-training β€” a tiny dataset (50h human + 4h robot) where humans and the robot perform the same tabletop task in the same scene under matched camera intrinsics. Ground-truth wrist pose from Vive trackers (3D position + orientation); ground-truth hand pose from Manus gloves (25 joint transforms). This bridges the human β†’ robot distribution gap before the tiny post-training set (Stage III) sees the task. Without Stage II, the massive Stage-I human prior would need to cross the embodiment gap during Stage III's ~100-trajectory post-train β€” too little signal.

6.3.4 The scaling law

EgoScale fits a log-linear relationship between validation loss and human-pretrain hours across the 1k–20k-hour range:

$$\mathcal{L}_{\text{val}} = 0.024 - 0.003 \cdot \ln(D)$$

with RΒ² = 0.9983. The paper claims the fit "strongly correlates with downstream real-robot completion scores" β€” the quantitative basis for NVIDIA's bet on human-video scaling. Independent replication has not happened yet (see Β§9).

6.4 Continuity vs. rupture across versions

Continuity threads:

  • DexMimicGen sim (from N1+) and DreamGen neural trajectories (from N1.5+) remain in every subsequent mix
  • Seven ego-video datasets from N1 stay in the pretrain β€” EgoScale augments, not replaces
  • The data-pyramid blueprint itself is invariant; only the layer being scaled hardest shifts

Rupture points:

  • N1.5's DreamGen introduction displaces N1's ad-hoc video model with a first-class world-model-based trajectory generator
  • N1.6's shift to thousand-hour teleop when partner-collected data finally reached VLA-scale
  • N1.7's EgoScale redefines what the "base layer" means: from passive internet ego-video (weak context via VQ-VAE latents) β†’ 20kh of action-labeled ego-video as a first-class training signal with its own scaling law

The N1.7 shift is the central narrative: NVIDIA frames it as the first robotics-scaling law where human-video hours is the x-axis β€” analogous to LLM compute/data scaling, but with an embodied validation-loss y-axis. Whether this scales to 200kh remains the open research question for post-N1.7.


7. Where GR00T's interface sits in the taxonomy

From Review-VLM-Action-Connection: GR00T is the canonical Category C (cross-attention) VLA. Specifically:

  • Cross-attention layer choice: GR00T physically truncates the VLM at select_layer, so "cross-attending to layer N" = "cross-attending to the last layer of a truncated stack." This is a different pattern from the paper description "cross-attention at layer 12 of Eagle-2" β€” the paper's wording is post-hoc; the code actually pops layers 13+.
  • N1's select_layer=12 default aligns the code with the paper. N1.5+ moved to full-length VLM with -1 default.
  • N1.6 added vlln LayerNorm β€” normalizing VLM features before cross-attention is a small but empirically-motivated hygiene step (analog of T5's pre-norm).
  • N1.7 added vl_self_attention β€” a configurable extra transformer that lets VL features self-mix before feeding DiT. Think of it as a "VL post-processor" or a lightweight Q-former-style module that can be enabled/disabled at config time.
  • AlternateVLDiT (N1.6+) directly mirrors RDT-1B's Alternating Condition Injection β€” image tokens would drown out text without alternation.

Comparison to other Category C VLAs:

Paper Cross-attn target Special handling
GR00T N1 Last layer of truncated (at 12) Eagle-2 None β€” raw vl_embeds
GR00T N1.6/N1.7 Last layer of truncated VLM + LayerNorm (N1.6), + self-attention block (N1.7), + alternating image/text (optional)
RDT-1B Separate SigLIP + T5-XXL Alternating Condition Injection (image and text cross-attended on alternate layers)
ST4VLA (ICLR 2026) k intermediate VLM layers Query transformer over multiple layers

GR00T N1's original mid-layer tap (layer 12) was empirically motivated in the paper for "higher policy success and lower inference latency" β€” a data point for the "which VLM layer?" open question flagged in Review-VLM-Action-Connection Β§8. The N1.5+ shift to last-layer tap (with physical truncation) suggests NVIDIA didn't find mid-layer clearly better in subsequent experiments.


8. Why iterate the VLM every release?

GR00T changed VLM backbone three times in 13 months: Eagle-2 β†’ Eagle-2.5 β†’ Cosmos-Reason β†’ Cosmos-Reason2. The rationale per release:

  • Eagle-2 β†’ Eagle-2.5: better multimodal pretraining; enables freezing (N1 had to unfreeze LLM; N1.5's frozen backbone is only possible with stronger pretrain)
  • Eagle-2.5 β†’ Cosmos-Reason: chain-of-thought capability β€” the headline N1.6 shift from "VLM that reads scenes" to "VLM that reasons about physics"
  • Cosmos-Reason β†’ Cosmos-Reason2 (Qwen3-VL): 256k context, native aspect-ratio, 2D/3D point localization β€” the N1.7 reasoning VLA gets stronger grounding

This is very different from the Ο€-series, which did one backbone swap (PaliGemma β†’ Gemma3-4B at Ο€0.6) and froze afterward. GR00T is the "NVIDIA VLM-of-the-month" series in a good sense: every open VLM NVIDIA releases becomes a GR00T backbone candidate.


9. Limitations & gaps

Authors' (implied / stated)

  • No published head-to-head benchmark tables for N1.6 or N1.7 against predecessors. NVIDIA's claims are qualitative ("outperforms N1.5", "comparable to N1.6 with improved generalization"). Rigorous comparison is impossible without this.
  • No arXiv paper for N1.5, N1.6, N1.7 β€” information is scattered across HF model cards, research blogs, forum posts, and GitHub READMEs.
  • HF model cards for N1.5/N1.6/N1.7 are template-cloned and all incorrectly list "SigLip2 + T5" β€” research blogs are the authoritative source for backbone identity.

Reviewer's concerns

  • Category C (cross-attention) locks in a latency overhead vs. the Ο€-series' same-stack MoE. GR00T's 63.9 ms / chunk (N1 on L40) is competitive but the architectural cost of a separate DiT is real.
  • Mid-layer vs. last-layer choice is unresolved β€” N1 used layer 12, N1.5+ abandoned it. No ablation published explaining the shift.
  • EgoScale scaling claim needs independent validation. 20kh β†’ validation-loss log-linear fit is a strong claim; third-party replication has not happened yet.
  • "Apache 2.0 weights" at N1.7 is early-access only. The actual license-file check in HF model cards should be verified before commercial deployment.
  • No head-to-head on contact-rich or dexterous tasks with the latest Ο€-series (Ο€0.6, Ο€0.7) or DDVLA on LIBERO. Public benchmarking is limited to NVIDIA's own suites.
  • vl_self_attention in N1.7 has no public ablation β€” NVIDIA adds architectural components release-by-release but doesn't publish which ones matter. A matched-compute ablation of vlln + vl_self_attention + AlternateVLDiT across the three versions is the ablation the community needs.

10. Significance β€” why the GR00T series matters

  1. Most-iterated open humanoid VLA of 2025–2026. Four releases in 13 months with public code; the Ο€-series had five but with closed weights.
  2. Reference dual-system architecture. "GR00T N1's System-2/System-1 split has become reference architecture" β€” cited in Review-VLA-Architecture, WholeBodyVLA, HiMoE-VLA, ReinFlow (which runs RL on GR00T-N1.5), and VLA-0 (NVIDIA, arXiv 2510.13054).
  3. Ecosystem integration. Newton (physics), Cosmos (world model), DreamGen (video rollouts), Isaac Lab, COMPASS, Jetson Thor β€” no competitor has this stack breadth.
  4. License flip at N1.7 is the most commercially significant 2026 VLA event. Until Apr 17, 2026, every production-grade generalist VLA was closed-weights (Ο€0.6/Ο€0.7, Gemini Robotics). GR00T N1.7 breaks that.
  5. VLM backbone churn as a feature. Every ~5 months NVIDIA swaps in a stronger VLM. This means GR00T benefits from general VLM progress faster than single-backbone-frozen systems like the Ο€-series.

11. Positioning vs Ο€-series

Axis GR00T Ο€-series
Developer NVIDIA (40+ authors) Physical Intelligence
Weights Apache 2.0 (N1.7) / Noncommercial (N1–N1.6) Closed (model cards only)
VLM→Action interface Category C cross-attn into truncated VLM Category B same-stack MoE with prefix-KV
Action head Flow-matching DiT, H=16, K=4 Flow-matching expert, H=50, K=5–10
Backbone path Eagle-2 β†’ Eagle-2.5 β†’ Cosmos-Reason β†’ Cosmos-Reason2 PaliGemma-3B β†’ Gemma3-4B (stable since Ο€0.6)
Gradient insulation (tune_top_llm_layers + freeze) Knowledge Insulation (2505.23705)
Humanoid focus Primary (Fourier GR-1, Unitree G1, Atlas) Manipulation-centric; UR5e zero-shot
RL integration via DreamGen / Isaac Lab recipes RECAP (Ο€*0.6) + distillation into Ο€0.7
Ecosystem Newton, Cosmos, DreamGen, Isaac Lab, Jetson OpenPi (open code for Ο€0/Ο€0.5-KI)
Signature novelty per release Backbone swaps + stack integration Algorithmic depth (KI, MEM, RTC, RECAP, metadata CFG)

As of April 2026: GR00T N1.7 is the first credible generalist humanoid VLA with Apache 2.0 weights AND the NVIDIA inference stack. The Ο€-series remains closed but has a stronger per-release algorithmic story. Different strengths: GR00T for humanoid breadth + ecosystem + commercial deployment; Ο€-series for production-tested single-model depth.


12. Reading path

flowchart LR
  A[Read arXiv 2503.14734<br/>N1 paper] --> B[Then research blogs<br/>N1.5 + N1.6]
  B --> C[Then this review Β§3<br/>for code-level connection evolution]
  C --> D[Then inspect git tags<br/>n1.5-release / n1.6-release / n1.7-release]
  D --> E[Finally pair with EgoScale<br/>arXiv 2602.16710 for N1.7 data]
Loading

13. Sources

14. Related wiki pages

← Back to ICLR-2026 Β· Home

⚠️ **GitHub.com Fallback** ⚠️