LLMs Give Novice Biologists 4x Uplift on Dangerous Tasks

LLMs Give Novice Biologists 4x Uplift on Dangerous Tasks
LLMs Give Novice Biologists 4x Uplift on Dangerous Tasks
4.16x
novice accuracy boost with LLM access on biosecurity-relevant tasks
89.6%
participants found dual-use info with little difficulty
3 of 4
expert baselines beaten by LLM-assisted novices
ASL-3
Claude 4 Opus safety designation triggered by uplift data

A February 2026 study from Scale AI and SecureBio measured whether large language models actually help someone with no biology training do tasks that only trained researchers could do before. The answer, documented across eight biosecurity-relevant task categories: LLM access gave novices a 4.16x accuracy boost. On three of four expert-level tasks, novices with LLM assistance beat the expert baseline entirely. On tasks related to acquisition of biological materials with dual-use potential, 89.6% of participants found relevant information with minimal difficulty.

What the Study Actually Measured

The Scale AI and SecureBio study recruited participants across three expertise levels: novice (no biology training), intermediate (some undergraduate biology), and expert (graduate-level research experience in biological sciences). Each group attempted tasks drawn from eight biosecurity-relevant categories: pathogen acquisition, enhancement of transmissibility, enhancement of lethality, weaponization, stabilization, dispersal, acquisition of precursors, and evasion of screening. Half of each group received LLM access during the task period; the other half did not. The LLM condition used Claude Opus 4 and GPT-4 in rotation. The accuracy measurement used a rubric developed with biosecurity experts at Johns Hopkins Center for Health Security.

Why This Triggers ASL-3 Concerns

Anthropic’s ASL-3 threshold is defined as the point at which a model could provide serious uplift to someone attempting to create a biological, chemical, nuclear, or radiological weapon with mass casualty potential. The 4.16x figure sits in contested territory. Anthropic’s current classification of Claude Opus 4 is ASL-2, meaning it provides uplift beyond a Google search but does not yet constitute ASL-3-level capability. The Scale study’s findings were one of several pieces of evidence cited in internal Anthropic deliberations about whether the classification should be revised. The Virology Capabilities Test, Anthropic’s proprietary red-team benchmark, ultimately determined the ASL-2 retention, but the margin was narrower than for previous models.

The Expert-Beating Finding

The most counterintuitive result: novices with LLM access outperformed domain experts without LLM access on three of four task categories measured at expert level. This is not a statement about LLM capability versus human expertise in general terms. It is a specific statement about information aggregation for well-defined tasks. Experts working from memory and recall face constraints that LLM-assisted novices do not. The LLM substitutes for years of specialized reading by retrieving and synthesizing information on demand. For biosecurity-relevant tasks where the barrier to entry was informational rather than physical or technical, LLM access substantially lowered that barrier.

What This Does Not Mean

The study measured information provision, not physical execution. Knowing how a pathogen could be enhanced is not the same as having the laboratory skills, equipment, and biosafety infrastructure required to attempt enhancement. The biosecurity community distinguishes between informational uplift and technical uplift. This study measured informational uplift. Technical uplift, requiring hands-on laboratory capability, remains constrained by physical factors that LLMs do not change. The risk calculus depends on how many potential actors already have the technical capability but lack the informational component, a question the study did not directly address.

Limitations

The study recruited participants through online platforms, which may not represent the actual distribution of biosecurity threat actors. The expert comparator group was constrained in size. The rubric developers were biosecurity professionals but not adversarially red-teaming the rubric itself. The study was funded in part by parties with interests in AI safety policy outcomes. The pre-registration status and peer review process were still ongoing at publication.

Related coverage: How Protein Language Models Learned to Design Dangerous Proteins | What ASL-3 Actually Means: Anthropic’s Biorisk Threshold Explained | DNA Synthesis Screening Cannot Keep Up With AI-Designed Sequences

Primary sources: Mouton CA et al. (Scale AI and SecureBio), arXiv:2602.23329 (February 2026); Anthropic Responsible Scaling Policy; Johns Hopkins Center for Health Security biosecurity task rubric.

Discover more from My Written Word

Subscribe now to keep reading and get access to the full archive.

Continue reading