Radiology Foundation Models: What Merlin, the 22% Hallucination Rate, and ED Fracture Data Tell Us

Radiology Foundation Models: What Merlin, the 22% Hallucination Rate, and ED Fracture Data Tell Us
Radiology AI foundation model chest radiograph Merlin

Radiology AI has been dominated by narrow task-specific models trained on single imaging modalities for single findings. Merlin, developed at Stanford University and published in Nature in 2026 (following a 2024 preprint), is a 3D vision-language foundation model for abdominal CT, trained on more than 6 million images from 15,331 CT scans, over 1.8 million EHR diagnosis codes, and more than 6 million tokens of radiology reports: learn general anatomical representations first, then apply them to any downstream task with far less labeled data than task-specific models require.

What Makes Merlin Different

Most radiology AI models operate on 2D slices. Merlin processes full 3D CT volumes at native resolution, learning anatomical relationships across axial, coronal, and sagittal planes simultaneously. The pretraining objective combines reconstruction of masked anatomical regions with contrastive learning between imaging and radiology report text. This image-text contrastive pretraining follows the same architectural logic as CLIP-based vision-language models, applied to a medical domain where image-caption pairs are CT volumes paired with radiologist reports rather than web-scraped photographs paired with alt text.

On downstream tasks, Merlin matched or exceeded task-specific models while requiring substantially fewer labeled fine-tuning examples, evaluated across 752 individual tasks spanning zero-shot findings and phenotype classification, disease prediction, report generation, and 3D segmentation.

The Annotation Bottleneck

Expert radiology annotations are expensive: annotating a single CT volume for complex segmentation can take 30 to 90 minutes. Merlin-class foundation models directly address this by reducing labeled data requirements for new tasks.

What Foundation Models Still Cannot Do

Merlin’s benchmarks reflect performance on tasks and populations close to its abdominal CT training distribution, internal validation plus external validation on outside clinical CTs and two public datasets. The model has not been evaluated on rare pathology types, pediatric populations, or imaging protocols significantly different from its training sites. Distribution shift remains an unresolved problem for all radiology foundation models, and later independent benchmarks have found newer foundation models outperforming Merlin on several task categories, which is the normal trajectory for a fast-moving subfield rather than a flaw specific to Merlin.

What Happens Next

The trajectory is toward multimodal foundation models processing CT, MRI, PET, and radiograph simultaneously. The regulatory pathway for foundation model-derived radiology tools under the FDA PCCP framework is an active area of policy development.

Related coverage: AI in Radiology: Three Phases and What the Clinical Evidence Shows | FDA Clearance for AI Medical Devices: What 510(k), De Novo, and PMA Mean | Poisoning the Medical Brain: RAG Attacks and Security in Clinical AI

Primary source: Blankemeier L et al., “Merlin: a computed tomography vision-language foundation model and dataset,” Nature 2026, doi:10.1038/s41586-026-10181-8. Originally posted as a Research Square preprint, June 2024. Updated 2026-08-18 to correct the developing institution (Stanford University, not Mass General Brigham/Harvard), the journal and publication year (Nature 2026, not Nature Medicine 2024), and the training dataset size (15,331 CT scans, over 6 million images, not 110,000 CT volumes).

Discover more from My Written Word

Subscribe now to keep reading and get access to the full archive.

Continue reading