Review RLDX 1 - Heungwoo/research GitHub Wiki
In-Depth Review โ RLDX-1: A Dexterity-First Foundation Model for Robot Hands (RLWRLD)
Model: RLDX-1 โ dexterity-first robot-hand foundation model ยท RLWRLD (Seoul, South Korea) Announced: May 2026 (SF Exploratorium launch) ยท built on NVIDIA's physical-AI stack. Status: industry release โ all figures below are company/vendor claims, not peer-reviewed. Filed alongside DYNA-2 as an industrial datapoint. Sources: RLWRLD ยท The Robot Report
โ ๏ธ Sourcing caveat. RLDX-1 is a company launch; there is no peer-reviewed paper, no released weights, and the benchmarks are RLWRLD's own vs its chosen baselines. Treat every number as a vendor claim.
Companions: Dexterous-Hand Data Pyramid ยท VLA Hybrid Architectures ยท DYNA-2 ยท Dexterous Manipulation.
1. TL;DR
- Four modality streams in one transformer. RLDX-1 uses a Multi-Stream Action Transformer (MSAT) โ Vision (fine-tuned Qwen3-VL 8B, spatial reasoning), Motion (spatio-temporal video features + object velocity), Memory (64 learnable "cognition tokens" for long-horizon), and Torque/Tactile (a Physics Module treating tactile + torque as native modalities for weight estimation and contact detection). "Each modality gets its own processing stream, and joint self-attention lets them interact."
- Human-hand-first data. Rather than teleop-heavy collection, RLDX records from the bare human hand and closes the gap in software with a retargeting framework built for five-finger dexterity โ reportedly >200 demonstrations/hour.
- Synthetic augmentation. Video-generation models vary lighting/surface/position; inverse dynamics annotate actions; quality-filtered. A claimed ~5ร data-scale increase โ +9.2% average success.
- Runs across embodiments. WIRobotics ALLEX humanoid (five-finger hands), Franka Research 3 (AnySkin tactile + joint torque), OpenArm + Inspire 6-DoF hand.
2. Why it matters (and the caveats)
- An industrial bet that mirrors the research pyramid. RLDX-1 explicitly composes bare-human-hand capture โ five-finger retargeting โ synthetic augmentation โ small teleop set โ the exact L1/L2 โ L4 โ L5 โ L6 stack the data pyramid describes, with torque/tactile as a native stream (the cross-cut). It's the commercial instantiation of "human-hand data first, retargeting is the bridge."
- MSAT is a multi-stream MoT. Per-modality streams + joint self-attention is the same family as the three-expert Mixture-of-Transformers in Review-VLA-Hybrid-Architectures โ here specialized to add motion, memory, and physics streams for contact-rich dexterity.
- But it's vendor-reported. Unlike the arXiv works on the pyramid, RLDX-1 has no paper/weights/independent eval; the ฯ0.5 / GR00T comparisons are RLWRLD's own.
3. Architecture & data (as described)
MSAT streams:
| Stream | What it does |
|---|---|
| Vision | robot-specialized VLM (fine-tuned Qwen3-VL 8B), spatial reasoning |
| Motion | spatio-temporal features from video; tracks object velocity |
| Memory | 64 learnable cognition tokens compress the scene for long-horizon tasks |
| Torque / Tactile | Physics Module โ tactile + torque as native modalities โ weight estimation, contact detection |
Data pipeline: bare-human-hand recording + five-finger kinematic retargeting (>200 demos/hr) โ synthetic augmentation (video-gen trajectories + inverse-dynamics action labels + quality filtering) โ small real teleop set. Built on NVIDIA's physical-AI stack.
Embodiments: ALLEX humanoid (five-finger) ยท Franka Research 3 (AnySkin + torque) ยท OpenArm + Inspire 6-DoF; supports single-arm / dual-arm / humanoid.
4. Reported results (โ ๏ธ vendor claims vs ฯ0.5 & GR00T N1.6)
- ALLEX humanoid: baselines report "success rates below 30%"; RLDX-1 "reaches nearly 90%" on motion / history / physical-signal tasks.
- OpenArm: RLDX-1 keeps "balanced performance across task types," while GR00T N1.6 "completely fails on the object identification task."
- Conveyor pick-and-place: +37.5 percentage points over GR00T N1.6.
- Data scaling: ~5ร data โ +9.2% average success.
(No absolute protocol, held-out set, or independent replication is disclosed.)
5. Significance & limitations
Significance. RLDX-1 is a notable industrial vote for the human-hand-first + retargeting + synthetic data pyramid, and for torque/tactile as a first-class stream in a multi-stream transformer โ aimed squarely at five-finger, contact-rich humanoid dexterity.
Limitations.
- Vendor-reported, not peer-reviewed / not replicated; no weights, no shared benchmark, self-chosen baselines.
- Bare-human-hand retargeting fidelity is the load-bearing assumption (the pyramid's L4 bottleneck) โ asserted, not independently measured.
- Synthetic-augmentation quality (video-gen + inverse dynamics) inherits pixel-world-model limits on contact physics.
- Absolute success protocols undisclosed โ "nearly 90%" / "below 30%" lack task-level detail.
6. Links
- Sources: RLWRLD RLDX-1 ยท The Robot Report
- Pyramid placement: L1/L2 human-hand โ L4 five-finger retarget โ L5 synthetic โ L6 teleop, + torque/tactile stream โ Dexterous-Hand Data Pyramid
- Architecture kin: VLA Hybrid Architectures (multi-stream MoT) ยท industrial sibling: DYNA-2
- Dexterous Manipulation ยท Tactile VLA