ICLR 2026 AutoBio - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Zhiqian Lan, Yuxuan Jiang, Ruiqi Wang, Xuanbing Xie, Rongkui Zhang, Yicheng Zhu, Peihang Li, Tianshuo Yang, Tianxing Chen, Haoyu Gao, Xiaokang Yang, Xuelong Li, Hongyuan Zhang, Yao Mu, Ping Luo · Paper: arXiv 2505.14030 (May 2025) · Category: Simulation framework + VLA benchmark · Trend tag: Professional/scientific-domain robot manipulation
flowchart LR
Real[Real lab instruments<br/>centrifuge, pipette, thermal cycler] --> Dig[Digitization pipeline]
Dig --> Sim[MuJoCo + custom physics plugins<br/>thread / detent / eccentric / quasi-static liquid]
Sim --> Render[PBR rendering<br/>dynamic panels, transparent materials]
Render --> Tasks[Biology-grounded tasks<br/>3 difficulty levels: Easy / Medium / Hard]
Tasks --> Demo[Demonstration generation]
Demo --> VLA[VLA integration: pi0, RDT]
VLA --> Eval[Standardized eval:<br/>precision, visual reasoning, instruction following]
VLA models are advancing on domestic tasks, but professional, science-oriented domains remain underexplored. Biology-lab automation combines structured experimental protocols with demanding precision, transparent/specular materials, dynamic digital instrument interfaces, and specialized mechanisms (threads, detents, eccentric drives, liquid handling) that existing manipulation simulators do not model.
AutoBio extends simulation along three axes:
- Instrument digitization pipeline to turn real-world lab apparatus into simulated assets.
- Specialized MuJoCo physics plugins for mechanisms ubiquitous in lab workflows — thread mechanisms, detent mechanisms, eccentric mechanisms, and quasi-static liquid computation — rarely addressed by prior simulators.
- Rendering stack supporting dynamic instrument interfaces (digital panels) and transparent materials via physically based rendering (PBR).
The benchmark provides biologically grounded tasks across three difficulty levels (Easy / Medium / Hard) covering protocol operations such as opening/closing instrument lids, picking up and transferring tubes, screwing/unscrewing caps, aspirating liquid with a pipette, operating digital panels, and loading centrifuge rotors. It ships demonstration-generation infrastructure and seamless VLA integration.
Baseline evaluations with two SOTA open-source VLA models — π0 and RDT — reveal significant gaps in precision manipulation, visual reasoning, and instruction following in scientific workflows. (Per-task success numbers omitted here pending the camera-ready tables.) The simulator and benchmark are released publicly for reproducible research.
First simulation benchmark to target high-precision, multimodal professional (scientific) environments for generalist robot policies, opening a domain distinct from the domestic / tabletop tasks that dominate VLA evaluation. The custom lab-mechanism physics plugins are reusable infrastructure beyond the benchmark itself.
- arXiv: https://arxiv.org/abs/2505.14030
- OpenReview: https://openreview.net/forum?id=UUE6HEtjhu
← Back to ICLR-2026