Google AI Model Matches or Exceeds Radiologist Performance in Lung Cancer Detection from CT Scans
In May 2019, researchers at Google Health and Northwestern Medicine published a deep-learning model in Nature Medicine that detected malignant lung nodules in low-dose CT scans, matching or exceeding the performance of six radiologists on a held-out dataset, with fewer false positives and false negatives when prior scans were unavailable.

Background
Reading a CT scan for lung cancer is painstaking work. A radiologist must examine hundreds of cross-sectional images from a single scan, looking for small nodules, then judge whether each one is likely to be malignant. Miss one and the consequences can be severe. Flag too many and patients go through unnecessary follow-up procedures and anxiety.
By the late 2010s, computer-aided detection tools already existed in radiology, but they worked as prompts rather than independent readers: they would highlight regions and a clinician would decide. What nobody had convincingly shown was whether a deep-learning model, a type of neural network trained on large volumes of labelled images, could assess a full CT scan end-to-end and produce a cancer risk prediction that held up against experienced specialists on a properly controlled test.
The National Lung Screening Trial, run by the National Cancer Institute, had produced one of the largest collections of annotated lung CT scans anywhere, with verified outcomes. That dataset made rigorous comparison possible. The question was whether anyone would do the work carefully enough to answer the question cleanly.
What happened
In May 2019, Daniel Tse, Shravya Shetty, Lily Peng and colleagues at Google Health, working with Mozziyar Etemadi and others at Northwestern University Feinberg School of Medicine, published their results in Nature Medicine. They had trained a deep-learning model, built on a convolutional neural network (one that processes images by scanning them in overlapping patches rather than treating each slice in isolation), on 42,290 CT scans from the National Lung Screening Trial. The model took a scan as its input and produced a single malignancy score, with no human in the middle of that process.
The team then tested it against six radiologists on a separate set of cases. When no prior scan was available for comparison, the model produced fewer false positives and fewer false negatives than the radiologists as a group. When a prior scan was available and given to the radiologists, performance was comparable. The model’s area under the ROC curve, a measure of how well a classifier separates positive from negative cases across all possible decision thresholds, reached 94.4% on the US validation set and 91.1% on an independent set of 1,139 cases from a different source.
One detail the team was careful about: they tested on data that had not been used in training, and they compared against specialist readers rather than generalists. That discipline is what made the result worth taking seriously. Google Health and Northwestern did not announce a clinical product. They published a research paper with numbers attached, which then entered a much longer conversation about how to verify and regulate tools of this kind before they reach screening programmes.
Why it mattered
The study provided one of the first rigorously controlled demonstrations that a deep-learning system could reach or surpass specialist clinician performance on a high-stakes cancer screening task, using a large and independently validated dataset drawn from the National Lung Screening Trial. It shifted the terms of debate about AI in radiology from theoretical promise to measurable clinical benchmarks, prompting broader discussion about how AI tools should be evaluated and regulated before deployment in screening programmes.
People
Daniel Tse, Atilla Kiraly, Lily Peng, Dale R Webster, Shravya Shetty, Mozziyar Etemadi, Krish Eswaran
Organisations
Google Health, Northwestern Medicine, Feinberg School of Medicine, National Cancer Institute, Nature medicine
Sources
- International evaluation of an AI system for breast cancer screening, corrected: this is the lung cancer paper: 'End-to-end lung cancer detection on CT scans using deep learning'.Nature Medicine.Primary source
- Google's AI system detects lung cancer better than radiologists in some cases.STAT News.Secondary
- Using AI to improve breast cancer screening, corrected: this refers to Google's official post on the lung cancer CT work.Google Blog.Official
Cite this page
AI Achievements. (2019). Google AI Model Matches or Exceeds Radiologist Performance in Lung Cancer Detection from CT Scans. Retrieved 2026-08-22, from https://achievements.ai/milestone/ai-outperformed-radiologist-lung-cancer
@misc{achievements_ai_outperformed_radiologist_lung_cancer,
title = {Google AI Model Matches or Exceeds Radiologist Performance in Lung Cancer Detection from CT Scans},
author = {{AI Achievements}},
year = {2019},
url = {https://achievements.ai/milestone/ai-outperformed-radiologist-lung-cancer}
}