Review System 0 1 2 - Heungwoo/research GitHub Wiki
Compiled April 2026 Β· Focus: how the Kahneman-style cognitive trichotomy (extended with "System 0") is being mapped onto humanoid manipulation stacks β what each tier is, who actually builds it that way, and where "System 0" ends up living when you bolt a foundation model onto a real robot.
This is the canonical "what does System 0 / 1 / 2 mean for a humanoid?" cross-paper / cross-company review. It complements three existing pages:
- Review-VLA-Architecture Β§5.F β the broader "hierarchical / dual-system / MoE" category
- Review-Fast-in-Slow β the deepest single-paper treatment of an embedded dual-system
- Review-VLM-Action-Connection Β§4 β the seven mechanisms by which S2 hands off to S1
Where those pages dissect architecture, this page dissects cognitive framing: how five robotics companies and ~20 manipulation papers all reach for the same Kahneman vocabulary, and what each one actually means by it.
Three independent pressures pushed humanoid VLAs into a layered cognitive vocabulary in 2025β2026:
- Frequency mismatch. A 7B VLM runs at 5β10 Hz on Jetson-class hardware. A balance controller for a 35-DoF biped needs β₯500 Hz to stay upright. Whatever sits between them needs an explicit name.
- Failure-mode decomposition. "The robot fell over" and "the robot misunderstood the instruction" are not the same bug. A single end-to-end policy can't be debugged, tested, or insured separately. A cognitive partition gives engineers (and regulators) something to point at.
- Industry messaging. Once Figure put "System 1 / System 2" in their Helix announcement (Feb 2025) and NVIDIA GR00T N1 echoed it in arXiv 2503.14734 (Mar 2025), every humanoid pitch deck started reaching for the same vocabulary. Whether the underlying engineering is genuinely tiered or just marketed that way varies considerably β this page calls out which is which.
The key 2026 wrinkle: by January 2026 two companies β Figure (Helix-02) and Sharpa Robotics (CraftNet) β went further and announced an explicit System 0 layer. They define System 0 very differently. That divergence is the most interesting open question on this page.
- Origin. Kahneman (2011) named S1/S2 in psychology. Bengio (NeurIPS 2019 Posner Lecture) imported them into deep learning. Chiriatti et al. (Nature Human Behaviour, 2024) coined "System 0" β but in their original framing System 0 sits above a human's S1/S2 (AI as cognitive infrastructure that pre-processes the world for you). The robotics community has inverted this: in humanoid stacks, "System 0" is repurposed below the visuomotor policy as a reflex / safety / sub-policy controller. No single paper has canonised the inverted usage yet.
- Industry. Only two companies publicly insist on a 3-tier S0/S1/S2 stack in April 2026: Figure (Helix-02, Jan 2026) and Sharpa Robotics (CraftNet, CES 2026). They disagree on what System 0 is: Figure's S0 is a 1 kHz, ~10M-parameter whole-body neural prior for balance/coordination/actuator-level control trained on >1,000 hours of human motion data; Sharpa's S0 is a ~100 Hz tactile fingertip "Interaction Brain" for fine-motor contact adjustment.
- Dual-system (S1+S2 only) is the dominant pattern. Figure Helix v1, NVIDIA GR00T N1/N1.5/N1.6/N1.7, Physical Intelligence Ο0.5/Ο0.7 (without the labels), Hi Robot, RoboDual, Fast-in-Slow, ChatVLA-2, ThinkAct, OpenHelix, CogACT, VITA-VLA, VLA-OS β all instantiate variations of the dual-system pattern. By ICRA 2026 this is the systems-community default β still label-free: Galaxea G0 (Qwen2.5-VL planner + PaliGemma-3B flow actor) and DualVLN (Ground Slow, Move Fast) ship explicit S2/S1 splits driven by the frequency mismatch, neither using the "System" vocabulary β confirming the pattern's dominance and Β§5.5's point that the labels are a framing, not a requirement.
- Triple-system papers exist but are rare. TriVLA (arXiv 2507.01424), Critic-in-the-Loop tri-system, VLSA safety-projection layer (arXiv 2512.11891), MinD, PhysiFlow, plus the LeVERB / WholeBodyVLA / SkillBlender family β but none yet uses the literal "System 0" term in a published paper.
- Where System 0 actually lives in 2026 humanoid stacks. Three concrete realisations: (a) whole-body MPC / RL controller beneath a latent-verb interface (LeVERB, WholeBodyVLA, SkillBlender), (b) safety / projection layer that overrides the VLA when constraints are violated (VLSA, SafeVLA, Latent Policy Barrier), (c) balance / actuator prior trained from teleop and human-motion-capture (Figure S0). The first two come from academia; the third is currently industry-only.
The canonical psychological dichotomy:
- System 1 β fast, automatic, intuitive, unconscious pattern-matching. Cheap and parallel; prone to bias and heuristic shortcuts.
- System 2 β slow, deliberate, effortful, rule-based reasoning. Expensive and serial; gets invoked when S1 lacks confidence or the problem is novel.
Kahneman never claimed these are literal brain modules β they are behavioural categories. Robotics papers tend to elide this caveat.
Delivered 11 Dec 2019, Vancouver (slides, virtual page). Bengio's thesis: deep learning so far had mastered System 1 (perception, intuition learned from static datasets); the next frontier was System 2 (reasoning, planning, causal inference, systematic generalisation), enabled by soft-attention bottlenecks ("the attention spotlight as conscious thought") and agent-centric rather than dataset-centric objectives.
This is the talk every dual-system VLA paper cites in its introduction. It legitimised the import of Kahneman's vocabulary into ML.
Chiriatti M., Ganapini M.B., Panai E., Ubiali M., Riva G. "The case for humanβAI interaction as system 0 thinking." Nature Human Behaviour 8(10):1829β1830 (2024). DOI 10.1038/s41562-024-01995-5.
Verified β this is where "System 0" was coined. Crucially, the authors define it as:
"A new psychological substrate distinct from Kahneman's tiers: a data-driven, externalised pre-reflective cognitive scaffold that precedes System 1/2 by pre-processing the world before the human even forms intuitions or reasoning."
In Chiriatti et al.'s framing, System 0 sits above a human's S1/S2 β generative AI as cognitive infrastructure, a "cognitive extension" that arranges the world for human consumption. The follow-up by Riva et al. ("System 0: Transforming AI into a Cognitive Extension," Cyberpsychology, Behavior, and Social Networking, 2025) and the Stanford "Toward a New Science of AI as Cognitive Infrastructure" piece (arXiv 2507.22893) extend the same direction: AI as a layer outside and prior to the human cognitive stack.
The robotics community has taken the term and flipped its position. In humanoid VLA stacks, "System 0" is repurposed as a tier beneath the System 1 visuomotor policy:
| Chiriatti et al. (humans) | Robotics inversion (humanoids) | |
|---|---|---|
| System 0 | AI cognitive extension above the human | Reflex / safety / sub-policy below the visuomotor policy |
| System 1 | Human fast intuition (untouched) | Fast neural visuomotor policy |
| System 2 | Human slow reasoning (untouched) | Slow neural VLM reasoning |
Why the inversion? Because the robot is the cognitive system being decomposed. The "thing pre-processing the world before higher-level cognition" β for a humanoid β is proprioception, balance, contact reflexes, and actuator-level control, not the cloud LLM. So System 0 in robotics has come to mean the kHz-rate, model-based or small-neural reflex layer that handles the parts of the world where the VLM and the visuomotor policy can't run fast enough. No published paper has formalised this inverted usage. Figure's Helix-02 announcement (Jan 2026) and Sharpa's CraftNet release (Jan 2026) are the two industry instantiations that named it.
Working definitions used throughout this page:
| Tier | Function | Typical realisation | Operating frequency |
|---|---|---|---|
| System 2 | Slow semantic reasoning Β· scene understanding Β· task decomposition Β· language grounding Β· goal selection | Pretrained VLM (PaliGemma, Eagle-2, Cosmos-Reason, Qwen3-VL, Gemma3-4B), 2Bβ14B parameters | 1β10 Hz |
| System 1 | Fast intuitive visuomotor policy Β· maps current observation + S2 latent β action chunk Β· "the cerebellum" | Diffusion / flow-matching / discrete-diffusion / AR action head, 80Mβ1B parameters; sometimes the last few VLM blocks (Fast-in-Slow) | 50β200 Hz |
| System 0 | Reflex / safety / sub-policy controller Β· enforces feasibility, balance, contact dynamics, joint/torque limits Β· "the spinal cord + brainstem" | Whole-body MPC, RL controller, balance prior, tactile servo, projection-based safety filter | 100 Hz β 1 kHz |
The inter-tier interfaces are where the design space diverges most:
- S2 β S1. Latent vector (Figure Helix, ThinkAct, RoboDual, WholeBodyVLA) Β· subtask text (Ο0.5 / Hi Robot) Β· cross-attention into VLM hidden states (GR00T N1.x) Β· shared parameters (Fast-in-Slow) Β· MoE-routed tokens (ChatVLA-2, HiMoE-VLA, AdaMoE).
- S1 β S0. Joint targets Β· end-effector waypoints Β· contact forces Β· or latent verbs (LeVERB) that an RL whole-body controller decodes. The Figure Helix-02 path is the cleanest published instance: S1 emits whole-body joint targets at 200 Hz, S0 turns them into actuator commands at 1 kHz with stability guarantees baked in.
What System 0 is NOT (in the robotics sense). The Chiriatti/Riva "AI as cognitive extension" sense doesn't apply to robots β the robot has no human cognition for AI to sit above. Don't conflate the two senses.
flowchart TB
subgraph S2[System 2 β slow deliberative Β· 1-10 Hz]
direction LR
Lang[Language instruction] --> VLM
Img[Multi-view RGB] --> VLM[Pretrained VLM<br/>2B-14B params<br/>scene understanding<br/>task decomposition]
Mem[Episodic memory] --> VLM
VLM --> Latent[Latent / subtask<br/>communication vector]
end
subgraph S1[System 1 β fast intuitive Β· 50-200 Hz]
direction LR
Latent --> AH[Action head<br/>diffusion / flow / AR<br/>80M-1B params]
ImgHF[High-freq RGB] --> AH
Tact[Tactile / 3D] --> AH
State[Robot state] --> AH
AH --> Cmd[Whole-body joint targets<br/>or end-effector waypoints]
end
subgraph S0[System 0 β reflex / safety Β· 100 Hz - 1 kHz]
direction LR
Cmd --> WBC[Whole-body controller<br/>MPC / RL / neural prior<br/>~10M params if neural]
Prop[Proprioception<br/>IMU + joint encoders] --> WBC
Fsafe[Safety constraints<br/>contact / torque / joint limits] --> WBC
WBC --> Act[Actuator torques /<br/>position references]
end
Act --> Rob[35-DoF humanoid<br/>servo loop]
Rob -. proprioception .-> S0
Rob -. cameras .-> S1
Rob -. cameras .-> S2
classDef s2 fill:#e3f2fd,stroke:#1565c0,color:#000
classDef s1 fill:#fff3e0,stroke:#ef6c00,color:#000
classDef s0 fill:#e8f5e9,stroke:#2e7d32,color:#000
class S2,VLM,Latent s2
class S1,AH,Cmd s1
class S0,WBC,Act s0
Reading the diagram. Each tier has its own sensors and its own clock. S2 sees everything but at 5 Hz; S0 sees only proprioception but at 1 kHz. Latency budget per tier roughly inverts: S2 has hundreds of milliseconds, S1 has tens, S0 has single-digit. The interfaces between tiers are the architectural design space β see Review-VLM-Action-Connection for the seven distinct S2βS1 mechanisms catalogued so far.
Figure is the company that has most aggressively promoted the S0/S1/S2 vocabulary in public marketing. Two checkpoints:
Helix v1 (Feb 2025) β first humanoid VLA to publicly adopt the "System 1 / System 2" labels.
- S2: Internet-pretrained VLM (~7B params), runs at 7β9 Hz asynchronously. Handles scene + language understanding; emits a latent communication vector into shared memory between the two systems.
- S1: 80M-parameter reactive visuomotor transformer at 200 Hz. Consumes the S2 latent (projected into S1's token space and concatenated with vision-backbone features) plus current observation. Outputs continuous upper-body actions for a 35-DoF humanoid. Gradients flow back from S1 into S2 through the latent during joint training β the two networks are trained end-to-end despite running asynchronously.
- No System 0 in v1; the low-level servo / balance loop is hand-engineered C/C++.
- Training: ~500 hours of teleoperated demonstrations total, in a single training stage with a single set of weights (no separate action heads or per-task fine-tuning).
- Source: https://www.figure.ai/news/helix
Helix-02 (Jan 2026) β adds an explicit System 0 tier; this is the first public industry stack with all three.
- S2: Slow semantic reasoning over goals, scenes, language. Carried over.
- S1: "Pixels-to-whole-body" visuomotor policy at 200 Hz; outputs joint targets for the entire frame (not just upper body). Adds palm cameras + 3-gram fingertip tactile sensors.
- S0: ~10M-param neural prior at 1 kHz trained on >1,000 hours of human motion-capture data. Takes full-body joint state + base motion, outputs joint-level actuator commands. Replaces ~109,504 lines of hand-engineered C++ balance/coordination code with a single learned model.
- Source: https://www.figure.ai/news/helix-02 Β· Humanoids Daily deep dive
Why Helix-02 matters for the framing. Figure is the first public humanoid stack to instantiate "System 0" as an actual neural prior trained from human motion data β rather than as a model-based MPC fallback. The 1 kHz / 10M-param / 1,000-hour spec sets a concrete reference point the field can argue with. The framing-level claim β "the brain has S1+S2; the spinal cord+brainstem are S0; the robot needs all three" β is rhetorically strong but only partially supported by neuroscience. (Real cerebellum and brainstem don't separate cleanly along these lines.)
Sharpa Robotics (Singapore, founded 2024). Hierarchical VTLA (Vision-Tactile-Language-Action) model deployed on the North full-body humanoid. Announced Jan 2026; live autonomous demos at CES 2026.
- S2 β "Reasoning Brain": VLM at ~1 Hz. Task decomposition, scene perception, planning.
- S1 β "Motion Brain": Foundation model at ~10 Hz. Coarse upper-body motions and pre-contact hand poses.
- S0 β "Interaction Brain": ~100 Hz tactile-feedback model. Continuous fine-motor adjustment of hand/finger positions during contact.
- CraftNet itself is explicitly defined as "the combination of System 0 and System 1" β so for Sharpa the headline product is the (S0+S1) tactile + motion fast-loop, with the VLM treated as a swappable upstream component.
- Source: https://www.sharpa.com/blogs/news/sharpa-announces-craftnet-a-hierarchical-vtla-model-for-fine-manipulation Β· AI Journal coverage Β· CES 2026 demo recap
This is the most interesting open disagreement in the 2026 humanoid landscape:
| Figure (Helix-02) | Sharpa (CraftNet) | |
|---|---|---|
| What is "System 0"? | Whole-body balance + actuator-level neural prior | Tactile-driven fingertip micro-adjustment loop |
| Frequency | 1 kHz | ~100 Hz |
| Inputs | Full-body joint state + base motion (proprioception) | Fingertip tactile + local hand state |
| Outputs | Joint-level actuator commands for all joints | Hand/finger micro-adjustments only |
| Replaces | ~109,504 lines of hand-engineered C/C++ balance code | Hand-tuned impedance / force controllers on the gripper |
| Why "S0"? | "The fastest reflex layer beneath the policy" | "The closest-to-physical-contact layer" |
Both definitions are defensible. Figure's S0 is Kahneman-pure: pre-reflective, reflexive, automatic. Sharpa's S0 is more about closing the contact-physics loop β closer in spirit to a tactile servo than to a balance reflex. The field has not decided yet whether a single "System 0" name should cover both, or whether they're really different tiers (call them S0_balance and S0_tactile) that coexist under S1. A future humanoid stack that needs both β e.g., a bimanual humanoid doing fine assembly while ambulating β would need both.
NVIDIA uses the S1/S2 vocabulary openly but does not claim a System 0 in any of N1 β N1.7. See Review-GR00T-Series for the code-level architectural treatment.
- S2: Pretrained VLM. Eagle-2 (N1) β Eagle-2.5 frozen (N1.5) β Cosmos-Reason 2B (N1.6) β Cosmos-Reason2 / Qwen3-VL (N1.7). Runs at ~10 Hz on L40.
- S1: Diffusion Transformer with flow matching. 16 layers (N1) β 32 layers (N1.6+). Cross-attends to the truncated VLM's last layer. Runs at ~120 Hz with H=16 chunks, K=4β5 denoising steps.
- No System 0. GR00T relies on whatever low-level controller comes with the host robot (Fourier GR-1, Unitree G1, Atlas β¦). NVIDIA explicitly leaves that tier to the embodiment vendor and to its own Isaac Lab / Newton physics tools.
- Sources: N1 launch press release Β· arXiv 2503.14734
The Ο-series (Ο0.5 / Ο0.6 / Ο0.7) and Hi Robot (arXiv 2502.19417) instantiate a clean hierarchical pattern:
- High-level semantic subtask predictor (the "inner voice" β emits text like "pick up the cup")
- Low-level flow-matching action expert (decodes a 50-step continuous chunk per call)
But β and this is unusual β Physical Intelligence has consistently avoided the literal "System 1 / System 2" labels in its public blog posts. The Ο0.5 blog (physicalintelligence.company/blog/pi05) describes the architecture as "hierarchical" and references the "inner voice"; nowhere does it call the two tiers System 1 and System 2. Whether this is an intentional refusal of the framing or just stylistic is unclear from public materials.
| Company | Stack name | Public framing | Has S0? |
|---|---|---|---|
| 1X Technologies | Redwood AI on NEO | Two-stage: high-level kinematic planner + low-level RL controller, plus separate World Model. Does not use S1/S2 labels. | Implicit (the RL controller plays an S0 role) |
| AgiBot | GO-1 ViLLA | VLM + MoE Latent Planner + Action Expert. Three-tier but framed as "ViLLA," not S0/S1/S2. | No (planner is S2-ish, not S0) |
| Skild AI | Skild Brain | Two-tier: low-frequency high-level policy β high-frequency low-level joint policy. No S0/S1/S2 wording. | Implicit |
| Sanctuary AI | Carbon / Phoenix | Hybrid symbolic + LLM + RL "cognitive architecture" mimicking memory/sight/sound/touch subsystems. Different vocabulary entirely. | No public S0 |
| Apptronik | Apollo | Public stack: NVIDIA GR00T on Jetson Orin + (separately) Google DeepMind Gemini Robotics. No proprietary cognitive framing. | Inherits from GR00T |
| Boston Dynamics + RAI | Atlas | "Cognitive AI" + "Athletic AI" pillars. No public S0/S1/S2 terminology. | Implicit (Atlas's MPC is decades-old; RL controllers are recent) |
| Tesla Optimus | (FSD-derived) | Single end-to-end neural net + xAI Grok for language. Explicitly anti-decomposition. | No |
| Unitree | UnifoLM-VLA-0 / WMA-0 | VLA + world-model-action. No S0/S1/S2. | No public S0 |
| Galbot, Galaxea, Fourier, UBTech | (various) | Mostly use NVIDIA GR00T or in-house single-stack policies. | No public S0 |
flowchart LR
subgraph EXP[Explicit S0/S1/S2 in public materials]
F["Figure<br/>Helix-02<br/>(Jan 2026)"]
SH["Sharpa<br/>CraftNet<br/>(CES 2026)"]
end
subgraph S1S2[Explicit S1/S2 only]
FH["Figure<br/>Helix v1<br/>(Feb 2025)"]
GR["NVIDIA GR00T<br/>N1 β N1.7"]
end
subgraph IMP[Hierarchical but doesn't use S labels]
PI["Physical Intelligence<br/>Ο0.5 / Ο0.7 / Hi Robot"]
OnX["1X Redwood"]
AGI["AgiBot GO-1"]
SK["Skild Brain"]
SAN["Sanctuary Carbon"]
end
subgraph NONE[Single-stack / anti-decomposition]
TE["Tesla Optimus"]
end
classDef exp fill:#e8f5e9,stroke:#2e7d32,color:#000
classDef s12 fill:#fff3e0,stroke:#ef6c00,color:#000
classDef imp fill:#e3f2fd,stroke:#1565c0,color:#000
classDef none fill:#fce4ec,stroke:#ad1457,color:#000
class F,SH exp
class FH,GR s12
class PI,OnX,AGI,SK,SAN imp
class TE none
Ordered roughly by influence and publication date. Bold = has a dedicated wiki page.
| Paper | arXiv / Venue | S2 (slow) | S1 (fast) | Headline mechanism |
|---|---|---|---|---|
| SayCan | 2204.01691 / Apr 2022 | LLM proposes high-level action | Learned affordance value-function disposes | First "LLM picks the verb, robot grounds it" pipeline β pre-VLA precursor |
| Inner Monologue | 2207.05608 / Jul 2022 | LLM with environment feedback in a closed loop | Primitive controller | Closed-loop LLM β adds reactivity to SayCan |
| Code-as-Policies | 2209.07753 / Sep 2022 | LLM emits Python | Perception/control primitives invoked by code | Programs as the inter-tier interface |
| RT-H | 2403.01823 / Mar 2024 | Predicts language motions ("move arm forward") | Conditions actions on the language motion | Language as the action hierarchy |
| RoboDual | 2410.08001 / Oct 2024 | OpenVLA generalist | 20M-param diffusion-transformer specialist | +26.7% real over OpenVLA with ~1 hr specialist training; 3.8Γ higher control rate |
| CogACT | 2411.19650 / Nov 2024 | DINO/SigLIP + Llama VLM | Diffusion-Transformer action module | +35% sim, +55% real over OpenVLA-7B; cleanest "cognition vs action" separation |
| Hi Robot | 2502.19417 / Feb 2025 | Ο0-class VLM emits sub-instructions | low-level Ο0 VLA | Open-ended instruction following with mid-task feedback ("that's not trash") |
| GR00T N1 | 2503.14734 / Mar 2025 | Eagle-2 VLM | Diffusion Transformer with flow matching | Open-foundation cross-embodiment humanoid VLA β see Review-GR00T-Series |
| OpenHelix | 2505.03912 / May 2025 | MLLM | 3D-aware policy (RVT-2-style) | Open-source replication + ablation of dual-system design space; SOTA on CALVIN ABC-D among open dual-system VLAs |
| ChatVLA-2 | 2505.21906 / NeurIPS 2025 | MoE VLM preserving open-world knowledge | Action expert | Two-stage training preserves OCR / math / spatial reasoning; MoE-routed dual-system |
| Fast-in-Slow | 2506.01953 / NeurIPS 2025 | Full VLM | Last 2 transformer blocks of the same VLM, re-run at high frequency | Embedded dual-system via parameter sharing; 117.7 Hz @ chunk=8 |
| ThinkAct | 2507.16815 / NeurIPS 2025 (NVIDIA) | RL-trained MLLM emitting visual latent plans | Downstream action decoder | Reinforced visual latent planning β RL-rewarded plans, planβlatentβaction |
| VLA-OS | NeurIPS 2025 | Configurable | Configurable | Controlled head-to-head ablation: Hierarchical > Integrated > Action-Only across 2D/3D, sim/real |
| VITA-VLA | 2510.09607 / Oct 2025 (Tencent) | Pretrained VLM, untouched | Distilled small action expert (reverse direction) | Cheap recipe to add action capability without expensive VLA pretraining |
| HiMoE-VLA | 2512.05693 / ICLR 2026 | Shared MLLM trunk | Hierarchical MoE action experts | Layer-wise heterogeneous routing; 97.8% LIBERO |
| WholeBodyVLA | 2512.11047 / ICLR 2026 | Latent-action VLA | LMO RL whole-body controller | Unified latent β RL WBC; +21.3% over baseline on AgiBot X2 humanoid for loco-manipulation β the controller plays an S0 role under the latent-verb interface |
| LeVERB | 2506.13751 / Jun 2025 (Berkeley/CMU/SFU) | Vision-language policy emits latent verbs | RL whole-body controller | First sim-to-real benchmark for vision-language humanoid WBC; 7.8Γ over naive hierarchical |
| Galaxea G0 | 2509.00576 / Sep 2025 | VLM | Action expert | Open-world dataset + dual-system VLA |
| DuoCore-FS | 2512.20188 / Dec 2025 (Astribot) | Ο0-FAST on PaliGemma-3B @ 1β3 Hz, autoregressive RVQ-VAE action tokens + CoT + bbox + fusion-query embeddings (Jacobi-decoded) | Pi0-small-style flow-matching transformer @ 25β30 Hz, 32-step whole-body chunks over 29-dim action | Truly parallel slow-fast via a written-and-read bridge buffer (instruction + fusion-query embeddings; fusion-query params trained through fast-side loss); whole-body 3-stream RVQ-VAE tokenizer (pos / 6D-SO(3) / gripper Γ codebook 1024 + geodesic SO(3) loss); 32.3 Hz on Astribot S1 25-DoF mobile dual-arm; differentiates from FiS-VLA / Helix / Hume |
Reading the table. Note that several "two-tier" papers β WholeBodyVLA, LeVERB, SkillBlender β actually have a third tier underneath: an RL whole-body controller that decodes the VLA's latent-verb output into 1 kHz joint torques. That third tier is functionally a System 0 even when the paper doesn't label it as such. This is the academic mirror of Figure's S0 design choice.
ICRA 2026 (survey) is the deployment-and-sensor-centric corner of the 2026 VLA wave, and even there the dual-system split shows up as a load-bearing design choice rather than a slogan β none of these papers reaches for the literal "System 1/2" branding, yet each instantiates the slow-VLM-plans / fast-policy-acts partition cleanly. The entries below place the relevant ICRA (and one closely-paired ICLR navigation sibling) papers in the S2/S1 framing.
| Paper | Venue | S2 (slow) | S1 (fast) | S2βS1 interface | Notes for this page |
|---|---|---|---|---|---|
| Galaxea / G0 | ICRA 2026 | G0-VLM β Qwen2.5-VL planner (7B/32B/72B evaluated), subtask reasoning | G0-VLA β PaliGemma-3B + SigLIP, FAST tokenizer + flow-matching action expert | Subtask goal (text) | Cleanest ICRA instance of the VLM-planner + separate flow action-expert pattern; S2 fine-tuned G0-VLM hits 83.3% on Table Bussing vs 32.0% for Gemini-2.5-pro. The empirical headline is data-side (single-embodiment pre-training), but the architecture is textbook dual-system. No S0. |
| CollabVLA | ICRA 2026 | Self-reflective reasoning loop that "dreams together" with a human, soliciting help | Underlying visuomotor policy | Reflection / uncertainty-triggered human hand-off | An inference-time S2 bolted onto a frozen S1 β the slow tier is a reflection-and-collaboration layer rather than a pretrained planner. Belongs to ICRA's reasoning/introspection cluster (Topic-VLA Β§3). No S0. |
| DualVLN β Ground Slow, Move Fast | ICLR 2026 (navigation) | 7B VLM global planner, image-grounded reasoning β mid-term pixel-goal waypoint + latent, ~2 Hz | Lightweight multi-modal-conditioning Diffusion Transformer local policy, ~30 Hz | Pixel goal + continuous latent (dual interface) | The VLN sibling of the manipulation dual-systems; ports the split into navigation as a foundation-model recipe. Its pixel-goal-plus-latent interface is a distinct S2βS1 mechanism worth filing alongside the seven in Review-VLM-Action-Connection Β§4. No S0. |
What ICRA adds to the framing. Three signals reinforce the dominant dual-system thesis of this page. (1) The split has crossed from ML venues into the systems community without the marketing vocabulary β ICRA authors adopt slow-plan / fast-act because it solves the frequency-mismatch problem on real robots, not because Figure popularised the labels. (2) The S2βS1 interface keeps diversifying: Galaxea uses subtask text (the Ο0.5/Hi Robot lineage), DualVLN pairs a discrete pixel goal with a continuous latent, and CollabVLA makes S2 an inference-time reflection loop over a frozen S1 β three more points in the design space Β§6 already maps. (3) Crucially for the System-0 thread: none of the ICRA dual-system entries has a System 0. Galaxea's R1 Lite, the CollabVLA platforms, and DualVLN's continuous-control base all inherit their low-level controller from the embodiment, exactly like GR00T. ICRA β the venue closest to real hardware β thus confirms the Β§7 observation that S0 remains an industry-and-whole-body-RL concern, not yet a manipulation-VLA one.
True three-tier architectures are rare. Most published "S0-like" candidates either repurpose the third tier for world modelling rather than for reflex/safety, or live in supplementary papers that other VLAs build on top of:
| Paper | arXiv / Year | Third tier | Role |
|---|---|---|---|
| TriVLA | 2507.01424 / Jul 2025 | Video-diffusion world model | S2 = perception VLM, S3 = world-model dynamics, S1 = policy. Third tier is world model, not reflex |
| Critic-in-the-Loop / Tri-System VLA | 2603.05185 / 2026 | Asynchronous visual Critic | VLM brain + VLA cerebellum + critic that monitors execution + triggers anomaly recovery |
| VLSA β VLA with Plug-and-Play Safety Constraint Layer | 2512.11891 / Dec 2025 | Safety-projection layer | Inserts a projection that overrides VLA actions to satisfy constraints β closest published analogue to a System-0 reflex |
| PhysiFlow | 2603.05410 / 2026 | Robust tracking controller for humanoid WBC | Multi-brain latent flow matching with a tracking controller that plays the System-0 role |
| MinD | (referenced in 2510.17111) | Low-frequency video-prediction planner + high-frequency diffusion policy + ... | Listed in the efficient-VLA survey as a triple-system architecture |
| SkillBlender | 2506.09366 / Jun 2025 | Skill library + VLA | RL whole-body skill library that the VLA composes β the skills are functionally an S0 |
| Reachability-Constrained SLS for humanoids | 2604.07644 / Apr 2026 | MPC with reachability constraints | Explicit S0 in the safety / control-theory tradition |
| SafeVLA | NeurIPS 2025 Spotlight | Constrained-learning safety alignment | Training-time safety pillar β bakes constraints into the VLA itself rather than adding a runtime S0 |
| Latent Policy Barrier | NeurIPS 2025 Spotlight | Inference-time barrier for OOD recovery | Latent-space safety projection β closest NeurIPS 2025 analogue to a runtime S0 |
Key observation. No 2024β2026 paper yet uses the exact term "System 0" for a robot reflex tier. The closest explicit naming sits in the conceptual literature (Chiriatti et al., Riva extensions) and survey papers (arXiv 2510.17111 discusses dual-/triple-system layering including reactive safety controllers). So when a paper builds something that functions as System 0 β VLSA's safety-projection, LeVERB's RL whole-body controller, SkillBlender's skill library, PhysiFlow's tracking controller β they call it almost anything but "System 0." The vocabulary lag matters because it makes the cluster harder to find by literature search.
flowchart TB
S0root[Where does the System 0 tier actually sit?]
S0root --> A[A. Whole-body MPC / RL controller<br/>under a latent-verb interface]
S0root --> B[B. Safety / projection layer<br/>that overrides the VLA when<br/>constraints are violated]
S0root --> C[C. Balance / actuator neural prior<br/>trained from teleop +<br/>human motion-capture]
A --> A1[LeVERB Β· WholeBodyVLA Β· SkillBlender<br/>academic, RL-trained, latent-verb interface]
B --> B1[VLSA Β· SafeVLA Β· Latent Policy Barrier<br/>academic, projection-based, training-time or inference-time]
C --> C1[Figure Helix-02<br/>industry, neural prior, 10M params at 1 kHz<br/>currently the only public instance]
classDef cat fill:#fff3e0,stroke:#ef6c00,color:#000
class A,B,C cat
Three live realisations, three independent communities. The unification β what should "System 0" actually mean for a humanoid? β has not happened.
Pulling everything together, the operational definitions that are actually load-bearing for a 2026 humanoid stack:
Purpose: Bridge from natural language and unstructured scene to a structured execution plan. Inputs: Multi-view RGB Β· language instruction Β· long-horizon memory. Outputs: Subtask sequence / latent plan vector / cross-attendable hidden states. Time budget: 100β500 ms per inference, asynchronous with respect to S1. Why it has to be a VLM: Web-pretrained semantic priors are the only way to handle novel objects, novel verbs, and novel scene layouts without per-task fine-tuning. Removing this tier reverts the policy to the 2022 OpenVLA-style "trained on the union of all robot tasks" recipe β which doesn't generalise to instructions like "load the air fryer" without a demo of that air fryer. Open question: Whether reasoning-trained VLMs (Cosmos-Reason, Qwen3-VL with CoT) genuinely improve task success or just look like they should. VLM4VLA is the canonical study showing VLM-benchmark scores correlate weakly with VLA performance.
Purpose: Map the current observation + S2 latent to smooth, high-rate, continuous robot actions. Inputs: High-frequency RGB + tactile + 3D + robot state + latest S2 latent. Outputs: Action chunk (typically 16β50 steps; flow-matching, diffusion, AR, or discrete-diffusion decoded). Time budget: 5β20 ms per chunk, must keep up with the servo loop. Architectural design space: see Review-VLA-Architecture Β§5 β this is the entire VLA categorical taxonomy. Flow-matching expert (Ο-series), continuous diffusion (RDT-1B / DexVLA), discrete diffusion (DDVLA / dVLA), AR action tokens (OpenVLA / VLA-0), embedded sharing (Fast-in-Slow). Open question: Whether a small, distilled S1 (FLOWER 950M, SmolVLA <0.5B) can match a tightly-coupled S1+S2 β this is the VITA-VLA distillation thesis.
Purpose: Guarantee feasibility (joint/torque/contact limits), maintain balance, close the contact-physics loop, and recover from S1 failures fast enough that they never escalate to a fall or a damaged object. Inputs: Proprioception (IMU + joint encoders) Β· contact / tactile Β· safety constraints. Outputs: Actuator torques or position references at 100 Hz β 1 kHz. Realisations in 2026:
- (a) Whole-body MPC / RL controller β academic standard. LeVERB, WholeBodyVLA, SkillBlender, Reachability-Constrained SLS. Strong stability guarantees; less sample-efficient to retarget.
- (b) Safety / projection layer β VLSA, SafeVLA, Latent Policy Barrier. Doesn't replace the controller, sits in front of the action stream and projects onto a safe subset.
- (c) Neural balance / actuator prior β Figure Helix-02. 10M-param, trained from human motion capture. Strong generalisation to novel poses; weaker formal stability guarantees than (a).
- (d) Tactile-driven micro-adjustment β Sharpa CraftNet's "Interaction Brain." Strictly local (fingertip / hand); doesn't address whole-body balance.
The big open question in 2026: does System 0 need to be one tier, or several? Figure says one (a 1 kHz whole-body neural prior). Sharpa says one but defines it differently (a 100 Hz tactile loop). LeVERB / WholeBodyVLA say one (an RL whole-body controller). VLSA says one (a safety projection). A future humanoid that needs all of: balance recovery + tactile servo + safety projection + actuator-level neural prior would have a fragmented S0. Whether they can be unified β or whether "System 0" really means "everything below S1, under whatever name" β is unresolved.
-
Should "System 0" be a single tier or a stack? Figure's S0 (1 kHz balance) and Sharpa's S0 (100 Hz tactile) both claim the name but address different physics. A unified definition would have to either (a) cover both, in which case S0 is itself layered, or (b) admit that "System 0" is a marketing term for "everything below the visuomotor policy."
-
Does System 0 have to be neural? Figure's neural prior is the boldest published claim. A model-based whole-body MPC (Atlas-class) plays the same role with formal stability guarantees and zero training data. The Figure claim β "we replaced 109,504 lines of C++ with one neural net" β is rhetorically powerful but doesn't yet have a published ablation showing the neural version is better on real tasks, only that it's learnable.
-
Does the inverted-System-0 framing survive contact with neuroscience? The Kahneman / Bengio / Chiriatti tradition is rooted in observable cognitive behaviour. Figure's "spinal cord + brainstem = S0" is a metaphor at best β real cerebellum and brainstem are heavily entangled with what robotics calls S1. If the framing fails, the field will fall back to "high-level / mid-level / low-level controller," which is what control theory called it for 40 years before VLAs renamed it.
-
Where does memory live? Review-VLA-Memory argues memory is its own architectural axis. In the S0/S1/S2 framing, memory could plausibly sit at S2 (long-term episodic), at S1 (short-term context window), or as its own tier. No published S0/S1/S2 paper has a clean answer.
-
Where does the world model live? TriVLA puts a world model at the third tier ("S3"). Cosmos Policy / DreamGen / Ctrl-World put it as an auxiliary or replacement for S2. mimic-video / VAM replaces the VLM with a video-generation backbone. None of these is comfortably "System 0/1/2."
-
Is the 3-tier framing inevitable, or is it an artefact of current hardware? Tesla's anti-decomposition position (single end-to-end network derived from FSD) is a falsifier. If a 2027 humanoid runs a 100B-param network at 1 kHz on next-gen silicon, the three tiers might collapse into one. Conversely, if dexterity demands push S0 into two tiers (balance + tactile), 2027 might have S0a/S0b/S1/S2.
-
Does "System 0" actually need a separate paper? Of the academic candidates above, all either inherit the safety / control / RL literature wholesale (LeVERB, VLSA, SkillBlender) or repurpose the third tier for something else (TriVLA's world model). A 2026 paper that explicitly claims and benchmarks a neural System 0 for a humanoid VLA stack β comparable in rigour to FiS-VLA's S1+S2 study β has not been published. Figure's Helix-02 is industry-only and lacks an ablation table.
- Kahneman, Thinking, Fast and Slow (Farrar, Straus and Giroux, 2011)
- Bengio NeurIPS 2019 Posner Lecture: virtual page Β· slides PDF Β· TechTalks summary
- Chiriatti et al., "The case for humanβAI interaction as system 0 thinking," Nature Human Behaviour 8(10):1829β1830 (2024) Β· DOI 10.1038/s41562-024-01995-5
- Riva et al., "System 0: Transforming AI into a Cognitive Extension," Cyberpsychology, Behavior, and Social Networking (2025) Β· https://www.liebertpub.com/doi/10.1089/cyber.2025.0201
- Stanford, "Toward a New Science of AI as Cognitive Infrastructure" β arXiv 2507.22893
- Figure Helix: https://www.figure.ai/news/helix
- Figure Helix-02: https://www.figure.ai/news/helix-02 Β· Humanoids Daily deep dive
- Sharpa CraftNet: https://www.sharpa.com/blogs/news/sharpa-announces-craftnet-a-hierarchical-vtla-model-for-fine-manipulation Β· AI Journal Β· CES 2026 demo recap
- NVIDIA GR00T N1: press release Β· arXiv 2503.14734
- Physical Intelligence Ο0.5: https://www.physicalintelligence.company/blog/pi05
- AgiBot GO-1 / ViLLA: https://www.globenewswire.com/news-release/2025/03/11/3040608/0/en/AgiBot-GO-1-The-Evolution-of-Generalist-Embodied-Foundation-Model-from-VLA-to-ViLLA.html
- 1X Redwood AI: https://www.1x.tech/discover/redwood-ai
- Skild AI: https://www.skild.ai/blogs/building-the-general-purpose-robotic-brain
- Sanctuary AI: https://www.sanctuary.ai/technology
- Apptronik Apollo: https://apptronik.com/apollo
- Boston Dynamics Γ RAI: https://bostondynamics.com/news/boston-dynamics-and-the-robotics-ai-institute-partner/
- Tesla Optimus brain architecture: https://news.accelerationrobotics.com/tesla-optimus-robot-brain-computer-architecture-hardware-software/
- SayCan (2204.01691) Β· Inner Monologue (2207.05608) Β· Code-as-Policies (2209.07753) Β· RT-H (2403.01823)
- RoboDual (2410.08001) Β· CogACT (2411.19650)
- Hi Robot (2502.19417) Β· GR00T N1 (2503.14734)
- OpenHelix (2505.03912) Β· ChatVLA-2 (2505.21906) Β· Fast-in-Slow (2506.01953)
- LeVERB (2506.13751) Β· ThinkAct (2507.16815)
- Galaxea G0 (2509.00576) Β· VITA-VLA (2510.09607)
- HiMoE-VLA (2512.05693) Β· WholeBodyVLA (2512.11047)
- VLA-OS β NeurIPS 2025 poster
- TriVLA (2507.01424) Β· Critic-in-the-Loop (2603.05185)
- VLSA (2512.11891) Β· PhysiFlow (2603.05410)
- SkillBlender (2506.09366) Β· Reachability-Constrained SLS (2604.07644)
- SafeVLA (NeurIPS-2025-SafeVLA) Β· Latent Policy Barrier (NeurIPS-2025-Latent-Policy-Barrier)
- Pure VLA Models β Comprehensive Survey (2509.19012)
- Efficient VLA Models for Embodied Manipulation (2510.17111)
- VLA Models: An Action-Tokenization Perspective (2507.01925)
- Large VLM-based VLA for Robotic Manipulation (2508.13073)
- A Survey on Efficient VLA Models (2510.24795)
- Behavior Foundation Model survey (2506.20487)
- Review: VLA Architectures Β§5.F β the broader hierarchical / dual-system / MoE category
- Review: VLMβAction Connection Β§4 β the seven distinct mechanisms by which S2 hands off to S1
- Review: Fast-in-Slow β deepest single-paper treatment of an embedded S1-in-S2
- Review: GR00T Series β the four-release NVIDIA dual-system humanoid VLA lineage
- Review: Humanoid VLA β applies this S0/S1/S2 stack to whole-body & bipedal loco-manipulation (balance Γ high-DoF action space Γ data collection)
- Review: VLA Memory β where memory sits in the cognitive stack
- ChatVLA-2 Β· ThinkAct Β· Fast-in-Slow Β· VLA-OS β the four NeurIPS 2025 papers that diversified the dual-system interface design space
- HiMoE-VLA Β· WholeBodyVLA β ICLR 2026 dual-system descendants
- Galaxea / G0 Β· DualVLN β Ground Slow, Move Fast β ICRA/ICLR 2026 dual-system entries (see Β§6.1); plus the broader ICRA 2026 VLA survey
- SafeVLA Β· Latent Policy Barrier β the safety pillar that doubles as a System-0 candidate
β Back to Home