Whether or not you’ve noticed it yet, artificial intelligence (AI) is beginning to appear in medical imaging processes in trusts around the UK. Depending on where you work, you may already have access to such systems, such as AI-interpretation directly integrated into PACS (Akbar, 2025; Calderdale and Huddersfield NHS Foundation Trust, 2025; Naz, 2024). For most doctors, this raises an immediate and practical question: what does that actually mean for the report I’m reading?
This article explores how AI interprets medical images, how that differs from the way a radiologist interprets them and what you should understand about the strengths and limitations of each approach. We use prostate MRIs as a case study as it is one of the most developed areas of AI in medical imaging, with large-scale international trials (Saha et al., 2024) directly comparing AI performance to that of radiologists.
Contents
How a Radiologist Reads a Prostate MRI
The human approach to image interpretation
Before looking at what an AI is doing when “reading” a scan, it is worth considering what a radiologist is actually doing when they interpret any medical image.
Radiologists rely on a combination of visual detection, pattern recognition, memory and cognitive reasoning to reach their conclusions (Waite et al., 2017). They will assess for particular radiological signs and put them together into a likely differential that fits with the clinical context. While analytical approaches are used extensively in radiological interpretation, rapid, unconscious (i.e. “gestalt”) assessments by experienced radiologists can allow them to assess whether a scan looks globally normal or abnormal within seconds of viewing it, before consciously working through the details. Expert radiologists direct their attention more precisely to relevant areas of an image than less experienced readers, having learned through training which features matter and which can be safely disregarded (Waite et al., 2019).
Unfortunately, pattern recognition is difficult to teach explicitly and is largely gained through the volume of images a radiologist has read and subconsciously processed over their career rather than through formal instruction alone (Alexander et al., 2020). Therefore, more standardised approaches are employed to ensure that an image is comprehensively assessed in all the necessary areas.
Even with standardised approaches to radiological interpretation, humans are still vulnerable to factors that have nothing to do with the image itself: fatigue, cognitive biases, increasing workloads and distraction all influence diagnostic accuracy (Waite et al., 2017).
However, radiologists can explain their reasoning, discussing which parts of a scan appeared normal or abnormal and why (McLeod et al., 2026). When faced with rare presentations, unusual anatomy, or unexpected artefacts, they can draw on broader medical knowledge rather than relying solely on what they have previously seen. And crucially, they may integrate the clinical picture, allowing a scan to be interpreted in the context of the scanned patient.
How this applies to prostate MRI in the UK
In the UK, NICE recommends that all patients with suspected localised prostate cancer undergo multiparametric MRI (mpMRI) before biopsy (NICE, 2019). The UK has been at the forefront of adopting this approach, with over 90% of areas now offering mpMRI (Tesfai et al., 2024).
To standardise reporting, radiologists score suspicious areas on a five-point scale, with 1 being highly unlikely to be clinically significant cancer and 5 being highly likely. In the UK, NICE recommends the Likert scoring system, a flexible approach based on the radiologist’s subjective global assessment of all imaging findings (NICE, 2019). The more structured Prostate Imaging-Reporting and Data System (PI-RADS) has been more common internationally but is becoming more widely used in the UK (Zawaideh et al., 2020). It uses a rule-based approach that prioritises specific factors such as size thresholds and invasive features.
The Likert approach can give a better sense of a radiologist’s clinical suspicion, but this subjectivity also makes it dependent upon the individual radiologist’s experience and therefore far less reproducible. There is well-documented variability between readers using the Likert approach (Schieda et al., 2019). Expert uroradiologists consistently outperform generalists in this task, but not every patient in every trust has access to that level of expertise (Kang et al., 2021).
How AI Reads the Same MRI
With that understanding of the human process, the AI approach becomes easier to appreciate.
A radiologist looks at an image and applies rules they’ve learned through training: “this pattern in this location is suspicious.” An AI system does the reverse. It starts with raw pixel data and works out for itself which patterns matter, based purely on what predicted cancer in its training data. Nobody tells it what to look for. Instead, it is shown thousands of MRI scans alongside the confirmed outcomes (“the ground truth”) and it builds a finely tuned ruleset to try and identify the features that distinguish why a particular scan is either normal or abnormal.
This has two important consequences:
- First, the AI may pick up on patterns that are invisible to the human eye, such as differences in image texture or subtle relationships between imaging sequences that fall below the threshold of what a person can perceive.
- Second, the features the AI uses may not correspond to anything a radiologist would recognise or describe. It doesn’t “see” a suspicious lesion in the way a clinician would. It identifies statistical patterns in the image that its training causes it to recognise as correlating with cancer.
There is a critical trade-off here. Current AI systems typically operate on the imaging data alone. They do not know the patient’s PSA, their age, whether this is a repeat scan, or why the imaging was requested. To overcome this, the approach to data collection and training will need be significantly more complex.
AI eliminates the human reader variability discussed earlier. But AI performance drops when used in settings different from its training data, whether through atypical presentations, underrepresented patient demographics or differences in imaging protocols. Accuracy can vary between human radiologists (inter-reader variability) and even for the same radiologist (intra-reader variability), but they can recognise and adapt to different protocols, scanners and patient populations in a way that AI currently cannot.
Where a radiologist brings adaptability, wider clinical reasoning and the ability to explain their conclusions, AI brings scale and consistency.
How do they compare?
Much of the existing literature comparing the accuracy of AI systems vs human radiologists suggests a rapid improvement in AI interpretation over the last decade, in many cases surpassing the human ability. However, there are still difficulties translating these achievements into practice in real-life clinical settings.
For interpretation of prostate MRI imaging specifically, the strongest comparison comes from the PI-CAI trial, an international study that compared an AI system against 62 radiologists from 20 countries using over 10,000 scans (Saha et al., 2024). It found that at the same rate of false alarms as humans, the AI system detected 6.8% more clinically significant cancers than the radiologists. Alternatively, at the same cancer detection rate as humans, the system produced 50.4% fewer false positives.
However, when the AI was compared against expert radiologists working more in line with their standard practice: with access to the full clinical picture and peer consultation, it did not clearly outperform them.
It’s worth noting that this trial used PI-RADS scoring rather than the Likert approach favoured in the UK. A smaller 2025 multi-centre study by European Radiology focussed on UK hospitals, assessed the MRI prostate scans of 793 patients and used both PI-RADS and Likert scores (Giganti et al., 2025). It found that the AI system, already in use in NHS trusts (Integrated Care Journal, 2025), was able to match the performance of MDT-supported radiologists.
The Black Box Problem
A recurring concern with AI in imaging (and beyond) is that you often can’t see its process. Many AI systems output a probability score but don’t explain their reasoning. Some tools attempt to address this by generating visual overlays (saliency maps) showing which parts of the image influenced the decision. In practice, clinicians have found these less helpful than expected as knowing that “this area mattered” is not the same as understanding why.


Figure 1. Example of a Saliency Map for a prostate MRI. These maps can be used to reveal the areas of a scan that is causing the AI to reach its conclusions. The above example is a visual representation of an AI system that is assessing for tumour volume and uses moderate-risk (green) and high-risk (red) to differentiate concerning features from normal tissue (blue) (Brisbane et al., 2024).
This is a notable contrast with human interpretation. A radiologist can explain their reasoning; discuss the features they found significant and justify their conclusion in clinical terms to colleagues. As it currently stands, AI radiology systems can’t meaningfully contribute to MDT discussions beyond the limited assessments they provide. The inability of most AI systems to do the same is not just a theoretical concern; it has practical implications for clinical responsibility. You remain responsible for the decisions you make, including those informed by AI outputs. If an AI flags a lesion and you act on it (or ignore it), you need to be able to justify that decision. Understanding the capabilities and limitations of the system you’re relying on is part of that responsibility.
This concern extends to informed consent. If you cannot determine why an AI system has flagged a possible lesion and requires a biopsy, explaining this to a patient in a way that constitutes genuinely informed consent for the procedure becomes complicated.
What This Means for You
This article has used prostate MRI as a case study, but the principles discussed here are not specific to urology. The same dynamics of consistency versus context, accuracy and explainability will continue to be important factors wherever AI imaging interpretation is deployed. As noted earlier, AI-assisted imaging is already present in NHS trusts. If it hasn’t reached yours yet, it is likely a matter of time.
If your trust is using AI-assisted imaging, it is worth finding out what that actually involves. Is the AI system flagging findings for a radiologist to review, contributing to the report itself, or triaging worklists to prioritise urgent scans? These are meaningfully different roles and they carry different implications for how much weight you should place on the output.
A confidence score generated by an algorithm is not the same as a report by a named radiologist who has reviewed the imaging and aware of the clinical context. It is worth being honest with yourself about how you respond to AI-generated outputs. The tendency to accept a computer-generated result uncritically is a documented risk in clinical settings and is more pronounced among less experienced clinicians (Rohan Krishna et al., 2025). If you find yourself deferring to an AI output without evaluating it, that should give you pause. As discussed in Should I Use AI as a Doctor?, your professional and legal responsibility for the decisions you make does not diminish because an AI system was involved.
Because the strengths and weaknesses of AI and human readers in some ways mirror one another, combining AI with a human reader often outperforms either alone (Wagner and Kaustubh Chakradeo, 2025).
This is a rapidly evolving field. Keeping up with these developments is increasingly part of being a well-informed clinician. For those who want to go further and consider pursuing this field directly, How to Get Involved in Medical AI outlines practical routes for doctors and medical students to do so.
References
- Akbar, A. (2025). Cheshire and Merseyside Radiology Imaging Network introduces transformational AI diagnostic technology. NHS Cheshire and Merseyside.
- Alexander, R.G., Waite, S., Macknik, S.L. and Martinez-Conde, S. (2020). What do radiologists look for? Advances and limitations of perceptual learning in radiologic search. Journal of Vision, 20(10), p.17.
- Brisbane, W.G., Priester, A.M., Nguyen, A.V., et al. (2024). Focal therapy of prostate cancer: Use of artificial intelligence to define tumour volume and predict treatment outcomes. BJUI Compass, 6(1).
- Calderdale and Huddersfield NHS Foundation Trust (2025). New cutting-edge AI technology to support reviews of chest X-rays goes live next week. CHFT News.
- Giganti, F., Moreira da Silva, N., Yeung, M., et al. (2025). AI-powered prostate cancer detection: a multi-centre, multi-scanner validation study. European Radiology, 35.
- Integrated Care Journal (2025). AI matches radiologists in detecting prostate cancer in NHS-backed multi-centre study. Integrated Care Journal.
- Kang, H.C., Jo, N., Bamashmos, A.S., et al. (2021). Accuracy of Prostate Magnetic Resonance Imaging: Reader Experience Matters. European Urology Open Science, [online] 27, pp.53–60.
- Rohan Krishna, N.K., Anupama, R.S., Shivaraj, K. (2025) Artificial Intelligence in Radiology: Augmentation, Not Replacement. Cureus. 2025 Jun 17;17(6):e86247.
- McLeod, G.A., Stanley, E.A.M., Rosenal, T. et al. (2026). Distinct visual biases affect humans and artificial intelligence in medical imaging diagnoses. npj Digit. Med. 9, 62.
- Naz, H. (2024). Bradford Teaching Hospitals welcomes Artificial Intelligence (AI) technology innovation in Radiology to improve patient care. Bradfordhospitals.nhs.uk.
- National Institute for Health and Care Excellence (2019). Recommendations | Prostate cancer: Diagnosis and Management | Guidance | NICE.
- Saha, A., Bosma, J.S., Twilt, J.J., et al. (2024). Artificial intelligence and radiologists in prostate cancer detection on MRI (PI-CAI): an international, paired, non-inferiority, confirmatory study. The Lancet Oncology, 25(7), pp.879–887.
- Schieda, N., Davenport, M.S., Pedrosa, I., et al. (2019). Renal and adrenal masses containing fat at MRI: Proposed nomenclature by the society of abdominal radiology disease‐focused panel on renal cell carcinoma. Journal of Magnetic Resonance Imaging, 49(4), pp.917–926.
- Tesfai, A., Norori, N., Harding, T.A., et al. (2024). The impact of pre‐biopsy MRI and additional testing on prostate cancer screening outcomes: A rapid review. BJUI Compass, 5(4), pp.426–438.
- Waite, S., Grigorian, A., Alexander, R.G., et al. (2019). Analysis of Perceptual Expertise in Radiology – Current Knowledge and a New Perspective. Frontiers in Human Neuroscience, 13(213).
- Waite, S., Scott, J., Gale, B., et al. (2017). Interpretive Error in Radiology. American Journal of Roentgenology, 208(4), pp.739–749.
- Zawaideh, J.P., Sala, E., Pantelidou, M., et al. (2020). Comparison of Likert and PI-RADS version 2 MRI scoring systems for the detection of clinically significant prostate cancer. The British Journal of Radiology, 93(1112), p.20200298.
How useful was this post?
Click on a star to rate it!
Average rating 5 / 5. Vote count: 1
No votes so far! Be the first to rate this post.
We are sorry that this post was not useful for you!
Let us improve this post!
Tell us how we can improve this post?


