
A February 2026 study from Scale AI and SecureBio measured whether large language models actually help someone with no biology training do tasks that only trained researchers could do before. The answer, documented across eight biosecurity-relevant task categories: LLM access gave novices a 4.16x accuracy boost. On three of four expert-level tasks, novices with LLM assistance beat the expert baseline entirely. On tasks related to acquisition of biological materials with dual-use potential, 89.6% of participants found relevant information with minimal difficulty.
What the Study Actually Measured
The Scale AI and SecureBio study recruited participants across three expertise levels: novice (no biology training), intermediate (some undergraduate biology), and expert (graduate-level research experience in biological sciences). Each group attempted tasks drawn from eight biosecurity-relevant categories: pathogen acquisition, enhancement of transmissibility, enhancement of lethality, weaponization, stabilization, dispersal, acquisition of precursors, and evasion of screening. Half of each group received LLM access during the task period; the other half did not. The LLM condition used Claude Opus 4 and GPT-4 in rotation. The accuracy measurement used a rubric developed with biosecurity experts at Johns Hopkins Center for Health Security.
Why This Triggers ASL-3 Concerns
Anthropic’s ASL-3 threshold is defined as the point at which a model could provide serious uplift to someone attempting to create a biological, chemical, nuclear, or radiological weapon with mass casualty potential. The 4.16x figure sits in contested territory. Anthropic’s current classification of Claude Opus 4 is ASL-2, meaning it provides uplift beyond a Google search but does not yet constitute ASL-3-level capability. The Scale study’s findings were one of several pieces of evidence cited in internal Anthropic deliberations about whether the classification should be revised. The Virology Capabilities Test, Anthropic’s proprietary red-team benchmark, ultimately determined the ASL-2 retention, but the margin was narrower than for previous models.
The Expert-Beating Finding
The most counterintuitive result: novices with LLM access outperformed domain experts without LLM access on three of four task categories measured at expert level. This is not a statement about LLM capability versus human expertise in general terms. It is a specific statement about information aggregation for well-defined tasks. Experts working from memory and recall face constraints that LLM-assisted novices do not. The LLM substitutes for years of specialized reading by retrieving and synthesizing information on demand. For biosecurity-relevant tasks where the barrier to entry was informational rather than physical or technical, LLM access substantially lowered that barrier.
What This Does Not Mean
The study measured information provision, not physical execution. Knowing how a pathogen could be enhanced is not the same as having the laboratory skills, equipment, and biosafety infrastructure required to attempt enhancement. The biosecurity community distinguishes between informational uplift and technical uplift. This study measured informational uplift. Technical uplift, requiring hands-on laboratory capability, remains constrained by physical factors that LLMs do not change. The risk calculus depends on how many potential actors already have the technical capability but lack the informational component, a question the study did not directly address.
Limitations
The study recruited participants through online platforms, which may not represent the actual distribution of biosecurity threat actors. The expert comparator group was constrained in size. The rubric developers were biosecurity professionals but not adversarially red-teaming the rubric itself. The study was funded in part by parties with interests in AI safety policy outcomes. The pre-registration status and peer review process were still ongoing at publication.
Related coverage: How Protein Language Models Learned to Design Dangerous Proteins | What ASL-3 Actually Means: Anthropic’s Biorisk Threshold Explained | DNA Synthesis Screening Cannot Keep Up With AI-Designed Sequences
Primary sources: Mouton CA et al. (Scale AI and SecureBio), arXiv:2602.23329 (February 2026); Anthropic Responsible Scaling Policy; Johns Hopkins Center for Health Security biosecurity task rubric.