详细信息
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation ( EI收录)
文献类型:期刊文献
英文题名:BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
作者:Du, Zhaohui[1,2]; Wang, Zhe[1,2]; Fei, Hongmei[4]; Cao, Xiwen[1,2]; Xiao, Ting[1,2]; Wang, Qi[3]; Jin, Huanbo[1,2]; Gu, Jiaming[1,2]; Lu, Quan[1,2]; Liu, Zhe[1,2]
机构:[1] Key Laboratory of Smart Manufacturing in Energy Chemical Process Ministry of Education, East China University of Science and Technology, Shanghai, China; [2] Department of Computer Science and Engineering, East China University of Science and Technology, Shanghai, China; [3] Department of Laboratory Medicine, Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, Shanghai, China; [4] School of Information Science and Technology, Shihezi University, Shihezi, China
年份:2026
外文期刊名:arXiv
收录:EI(收录号:20260287286)
语种:英文
外文关键词:Intelligent agents - Interface states - Laboratories - Molecular biology - Robots - User interfaces - Vision - Visual languages - Waste disposal
摘要:Biological laboratory automation holds promise for reducing repetitive manual work, improving workflow reproducibility, and broadening access to experimental execution. Yet building embodied agents that operate reliably in real wet-lab environments remains difficult. Biological protocols are typically expressed as unstructured natural language, common labware such as tubes and bottles are often transparent or reflective, and multi-step procedures require state-aware execution rather than one-shot instruction following. In addition, existing robotic laboratory platforms often depend on expensive hardware, dedicated instruments, fixed workflows, or robotics-oriented interfaces, which can limit their accessibility for resource-constrained laboratories and make them difficult to use or adapt for domain users without robotics expertise. Here, we introduce BioProVLA-Agent, an affordable, protocol-driven, and vision-enhanced embodied multi-agent system enabled by Vision-Language-Action (VLA) models for biological laboratory manipulation. BioProVLA-Agent uses experimental protocols as the task interface and couples protocol understanding with closed-loop-capable reasoning and embodied execution. A Tailored LLM Protocol Agent transforms unstructured protocols into executable and verifiable subtask units, while a VLM-RAG Verification Agent reasons over real-time visual observations, robot states, retrieved operation knowledge, and reference success/failure examples to assess task readiness and completion. The verified subtasks are then executed by a VLA Embodied Agent based on a lightweight VLA policy. To address wet-lab visual perturbations, we further develop AugSmolVLA, an online augmentation strategy tailored to transparent labware, specular reflections, illumination shifts, and overexposure. We evaluate BioProVLA-Agent on a hierarchical biological manipulation benchmark spanning 15 atomic tasks, 6 composite workflows, and 3 representative bimanual tasks, including centrifuge-tube loading, tube sorting, waste disposal, cap twisting, and liquid pouring. Across normal and high-exposure settings, AugSmolVLA improves execution stability over ACT, X-VLA, and the original SmolVLA, with pronounced gains in precise placement, transparent-object manipulation, composite workflows, and visually degraded scenes. These results demonstrate a practical route toward affordable, protocol-centered, and verification-capable embodied AI systems for biological laboratory manipulation. Copyright ? 2026, The Authors. All rights reserved.
参考文献:
正在载入数据...
