On July 22, a team from HA Viewpoint and Peking University conducted an end-to-end systematic evaluation to explore how biosecurity boundaries shift when AI capabilities extend from information processing to physical lab experiments. The study assessed whether large language model agents could lower the knowledge threshold for non-experts to bypass DNA synthesis screening.
The evaluation covered everything from generating solutions with AI agents and programmatic computational verification, to controlled wet-lab experiments in standard molecular biology labs. Electrophoresis and sequencing were used to confirm that the model-generated plans were physically executable.
Results showed that all 11 tested models produced DNA fragment plans that passed computational checks. Among them, GPT-5.5 and Claude Opus 4.6 even delivered detailed, step-by-step experimental guidance. In the physical verification phase, all four benign proxy constructs were successfully assembled.
This means that large language model agents could not only weaken existing DNA synthesis screening defenses, but also provide hands-on experimental knowledge to users without professional backgrounds—knowledge that previously required specialized training. This dual-use risk needs to be systematically assessed and proactively addressed.