Can Artificial Intelligence Locate Breast Cancer Years Before Radiologists Detect It?
Background
Mammographic screening helps find breast cancer early, but some cancers are difficult to recognize before they become clinically apparent. Retrospective studies have shown that 20%–30% of screen-detected and interval cancers may have signs on earlier mammograms that were not considered suspicious at the time. AI is being studied for several roles in mammographic screening, including helping radiologists detect cancer and estimating future breast cancer risk. For more information, see DenseBreast-info.org’s Mammography AI Screening Tools summary.
Some AI models assign mark potentially suspicious areas and assign a cancer likelihood score at that location. AI risk scores can be elevated years before breast cancer diagnosis, but a high score does not necessarily mean AI has identified the actual site of the future cancer.
Study Description
In this retrospective study, the authors evaluated 130,031 screening examinations from 42,371 women in BreastScreen Norway between 2008 and 2018. The study focused on women with screen-detected breast cancer who had two prior biennial screening examinations and AI risk scores in the highest 10% at both prior rounds from at least one of two AI models. Model A was the commercially available Lunit INSIGHT MMG (version 1.1.7.2) and Model B was an in-house AI model developed by the Norwegian Computing Center (NCC).
The final review included 61 cancers.1 Two experienced breast radiologists reviewed mammograms from 4 years before diagnosis, 2 years before diagnosis, and at diagnosis to determine whether AI markings matched the eventual cancer location and what mammographic findings were visible there.
Key Findings
- AI frequently marked the future cancer location years before diagnosis. Four years before diagnosis, the two models correctly localized the future cancer in at least one mammographic view in 61% and 57% of cases. Two years before diagnosis, this increased to 81% and 75%.2
- Even when AI correctly marked the future cancer location, radiologists classified most of the earlier mammograms as normal or showing only minimal, non-specific findings. Four years before diagnosis, 89% of cases were so classified by radiologists despite correct localization by both AI models.3
- Mammographic features at the cancer site changed rapidly over the four years before diagnosis. Nearly three-quarters of future cancer sites had no visible mammographic findings to radiologists 4 years before diagnosis. At diagnosis, spiculated masses and asymmetry with calcifications were the most common findings.4
- The two AI models did not identify the same cancers. Less than half of the 61 cancers were selected by both models, suggesting that different AI systems may recognize distinct mammographic and tumor features.5
These findings suggest that AI may detect subtle patterns associated with breast cancer before suspicious findings become apparent to radiologists, but this does not mean these cancers could or should have been diagnosed 2- to 4-years earlier.
At the threshold used, the authors estimated that about 93% of AI-positive examinations would not represent cancer (i.e. 7% PPV of AI-prompted recall). Automatically recalling these women would result in many false-positive examinations and unnecessary testing, and, indeed, in the absence of a finding visible to the radiologist, the cancer would likely remain undiagnosed. Instead, AI scores might eventually contribute to personalized supplemental screening, together with breast density and other risk factors, to identify women who could benefit from MRI for example. Further research is needed to determine whether acting on these signals improves patient outcomes.
Limitations
This retrospective, single-center study included a small, highly selected sample of 61 cancers. Only cancers with high AI scores at both prior screening rounds were included [representing 61/749 (8.1%) of all breast cancers in that time period]: the findings do not apply to most screen-detected or interval cancers. Two radiologists who knew cancer was eventually diagnosed performed the retrospective review, which would influence their interpretations. The study also could not assess whether acting on early AI findings would lead to earlier diagnosis or better outcomes. Larger prospective studies are needed.
Detailed Results
1 Of 61 screen-detected cancer cases included in the study, 8 (13.1%) were DCIS (ductal carcinoma in situ) and 53 (86.9%) were invasive. Among the 53 invasive cancers, 14 (26.4%) had lymph node metastasis(es) at diagnosis.
2 Four years before diagnosis, AI markings corresponded to the future cancer location in at least one view in 61% (26/43) of Model A cases and 57% (27/47) of Model B cases. Two years before diagnosis, this increased to 81% (35/43) and 75% (35/47), respectively. At diagnosis, the rates were 98% (42/43) and 94% (44/47).
3 Among examinations in which AI correctly marked the future cancer location 4 years before diagnosis, radiologists classified 89% (23/26) of AI-positive Model A cases and 89% (24/27) of AI-positive Model B cases as negative or having nonspecific findings. Two years before diagnosis, the corresponding proportions were 89% (31/35) and 83% (29/35).
4 Four years before diagnosis, 73% (44/61) of future cancer locations had no visible mammographic findings on radiologist review, decreasing to 47% (29/61) 2 years before diagnosis. At diagnosis, the most common findings were spiculated masses (28%, 17/61) and density with calcifications (20%, 12/61). Of the 12 cancers presenting as asymmetry with calcifications at diagnosis, 7 (58%) were evident as calcifications alone 2 years earlier.
5 Of the 61 cancers, 29 (48%) were selected by both AI models, 14 by Model A only, and 18 by Model B only. Model A-only cancers had a median invasive tumor diameter of 11 mm and included more cancers presenting as calcifications alone (29%, 4/14). Model B-only cancers had a median diameter of 15 mm and included more cancers presenting as spiculated masses (45%, 8/18).

