Fall 2026 Lecture Series
The Physician AI Handbook
Clinical AI
in radiology.
What works. What doesn't.
What happens when AI is wrong.
7:30 AM ET · Zoom
Four questions for this morning.
Why build a clinical AI handbook?
The gap between a published result and a decision in your department.
What does the radiology evidence support?
Trials, workflow studies, and the limits of each result.
What happens when the system is wrong?
Three teaching cases, then the controls that matter.
How do we keep evidence useful?
Source verification, version changes, and questions for the next studies.
Published accuracy is one part of the decision.
patient benefit.
Model
Can it perform the stated task?
Workflow
Does it change decisions or time to action?
Patient
Does it improve outcomes or reduce harm?
An open reference for clinical decisions.
Read the evidence
Clinical studies, their designs, and the outcomes they measured.
Evaluate a tool
Intended use, local validation, human factors, and monitoring.
Revisit the claim
Sources, versioned changes, and corrections that readers can inspect.
The sources can be peer reviewed. The handbook is an open, versioned reference.
Four questions before accepting a claim.
What is the original source?
Read the study or regulatory record, not only the summary.
What outcome was measured?
Keep detection, workflow, patient benefit, and cost separate.
Where does the result apply?
Identify the product, version, population, and clinical setting.
What can the source not establish?
Keep limitations beside the result, where they affect the decision.
Choose the endpoint for the intended use.
Triage
Time to interpretation
or action
Include delays and misses in unflagged studies.
Screening
Detection, recall,
interval disease
A higher detection rate alone does not establish mortality benefit.
Report drafting
Final-report accuracy
and editing burden
Measure consequential errors alongside time saved.
Faster thrombectomy. Patient benefit remains uncertain.
Adjusted difference in door-to-groin time
Stepped-wedge cluster-randomized trial
243 thrombectomy patients
What the trial supports
A shorter treatment-process interval after AI activation.
What remains uncertain
The trial did not demonstrate a significant improvement in 90-day functional independence.
A nonsignificant result is not proof that the outcome is unchanged.
Three studies. Three different interventions.
MASAI
Sensitivity with AI support, versus 73.8% with standard double reading.
Specificity: 98.5% in both groups.
PRAIM
Adjusted cancer detection with AI, versus 5.7 per 1,000 without.
Radiologists chose when to use AI.
AITIC
Relative increase in recall. Recall noninferiority was not met.
AI triage was also checked against standard readings.
Detection, recall, interval cancers, and workload answer different questions.
More detections do not guarantee a faster pathway.
Detection trial
Nam: actionable nodules detected in 0.59% with AI versus 0.25% without.
Prospective silent evaluation
Storey: outputs were hidden from readers. A silent study does not test autonomous live use.
LungIMPACT pathway trial
Added AI worklist prioritization did not significantly shorten time to CT or cancer diagnosis.
A version update changes the assessment.
The same 117,709 Norwegian examinations
Cancers assigned the highest AI score
Share of screen-detected cancers · 0–100% scale
The decision consequence
Different scores can send different examinations through a risk-adapted reading pathway.
The evidence limit
Retrospective scoring does not establish the net benefit of a new live workflow.
Clearance answers an intended-use question.
Read the label
What task is authorized?
For which patients and images?
With what role for the reader?
Read the evidence
What testing is described?
Was the workflow prospective?
Was the reader part of the test?
Evaluate locally
Does the evidence apply here?
What can fail?
Who can pause the tool?
Missing testing in a public document is not proof that testing never occurred.
Evaluate the reader and the model together.
Advice can mislead
Incorrect AI advice reduced radiologists’ aggregate performance in a controlled study.
Trust is an outcome
Local explanations increased reliance on simulated advice versus global explanations, including when advice was incorrect.
Sequence needs testing
Interpretation order, interface, and escalation are part of the workflow.
An explanation that sounds plausible is not a validation result.
No flag appeared. The study entered the routine queue.
A trauma head CT has an urgent clinical indication.
The AI triage tool produces no flag.
The worklist moves the examination into a routine queue.
What should have determined its priority?
The clinical indication still determines priority.
Retain the required turnaround
An absent AI flag does not establish that the study is normal.
Test the entire queue
Monitor delayed unflagged studies as well as accelerated flagged studies.
Give the department control
Define escalation and a way to pause unsafe prioritization.
Your review is benign. AI flags an asymmetry.
A screening mammogram includes a prior examination.
Your initial assessment is benign.
The CAD output identifies an asymmetry for review.
What evidence would change your assessment?
Re-review the finding. Own the final assessment.
Check the finding
Review the location, imaging features, priors, and technical adequacy.
Use the appropriate category
The vignette alone cannot determine a BI-RADS category.
Document the reasoning
Record why the AI mark changed, or did not change, the final assessment.
The scanner changed. The AI assessment did not.
A department introduces a different CT platform.
The deployed AI was evaluated on the previous acquisition setup.
The equipment change did not trigger an AI impact assessment.
Which control should connect these decisions?
Make equipment change an AI review trigger.
Assess what changed
Scanner, reconstruction, preprocessing, PACS, model, or threshold.
Test before relying on the new setup
Agree on validation, acceptance, and pause criteria with clinical owners.
Monitor after the change
Review discordant cases and alert patterns with authority to intervene.
A living reference also needs change control.
| Clinical AI question | Reference question |
|---|---|
| Does the tool apply to this setting? | Does the source support this sentence? |
| Did the deployed version change? | Did the evidence or guidance change? |
| Can a failure be detected and corrected? | Can a reader inspect the source and correction? |
Keep the source attached to the claim.
Find
Treat a new paper or report as a lead.
Verify
Read design, outcome, figure, and limitations.
Review
Inspect the exact wording and its implication.
Update
Record the change and check the published result.
Automated drafting can accelerate work. It cannot establish source truth.
A working link can still support the wrong claim.
Right DOI, wrong attribution
The link resolves, but the cited author or journal is wrong.
Right study, wrong design
An observational comparison becomes “a trial” in the summary.
Right number, wrong conclusion
A workflow improvement becomes a claim about patient outcomes.
Broad benchmarks leave clinical questions open.
MedVersa
One model across imaging and report tasks.
Developer studies require prospective validation in the intended clinical workflow.
RADAR
Broad abdominal CT finding detection.
Research evaluation across 146 findings does not establish autonomous clinical use.
Consumer models
Curated case answering is a different task.
Open-ended primary accuracy: 51.6–62.7% on the tested RSNA cases.
Task-specific errors and prospective workflow performance still matter.
Evidence belongs to the deployed configuration.
Model version
Threshold
Acquisition
Reader workflow
When a consequential component changes,
reopen the assessment.
Questions an academic department can answer.
Reporting
Do drafts reduce work without increasing consequential final-report errors?
Patient pathways
Does a prioritization change improve time to meaningful clinical action?
Training
How does independent interpretation change with sustained AI exposure?
Post-market monitoring
Which performance changes follow a new model, threshold, or scanner?
Three gates before clinical use.
Evidence fits the use
Population, setting, data, threshold, and outcome match the decision.
Risk controls are defined
Labeling, failure modes, reader responsibilities, and escalation are clear.
Monitoring can change practice
Metrics, review ownership, and pause criteria exist before deployment.
An unresolved gate calls for corrective work and reassessment.
Define monitoring before deployment.
Before
Validate the exact use.
Set acceptance criteria.
Map likely failure modes.
During
Test integration.
Train for disagreement.
Make errors reportable.
After
Track errors and drift.
Review local subgroups.
Reassess consequential changes.
Three things to keep.
Match the endpoint to the claim.
Accuracy, workflow, and patient benefit are different questions.
Tie evidence to the configuration.
Version, threshold, acquisition, and reader workflow matter.
Keep evaluating after deployment.
Detect change, review failures, and retain the ability to pause.
Questions and discussion.
Will AI replace radiologists?
Start with tasks
The evidence in this talk supports assistance or automation of selected tasks.
Do not turn scenarios into forecasts
Workforce models depend on demand, payment, staffing, and deployment assumptions.
Retain clinical functions
Interpretation, procedures, consultation, communication, and quality assurance are different tasks.
How should we validate locally?
Match the intended use
Sample relevant local patients, scanners, image types, and clinical workflows.
Define the reference standard
Use adjudication suited to the task, including consequential errors.
Set decisions in advance
Acceptance, escalation, monitoring, and pause criteria should precede the pilot.
Who is liable when AI misses?
Keep the legal claim narrow
Responsibilities depend on facts, jurisdiction, contracts, and the clinical context.
Separate authorization from liability
FDA authorization does not resolve who bears civil liability in a particular case.
Make ownership explicit
Discuss clinical duties, documentation, vendor terms, and coverage with institutional counsel.
What about resident deskilling?
Distinguish concern from demonstration
The cited longitudinal signals are from endoscopy, not radiology residency.
Measure independent performance
Assisted accuracy alone cannot show whether unassisted skill is preserved.
Test safeguards
Independent interpretation, disagreement review, and failure recognition are training options to evaluate.
Are model-drafted reports ready?
Workflow signal
A prospective matched cohort associated assistance with 15.5% less documentation time.
Important boundary
The study covered one 12-hospital system. It was not randomized patient-outcome evidence.
Measure the final report
Evaluate consequential errors, editing burden, omissions, and failure handling.
Which vendor is best?
Require the right comparison
Independent head-to-head evidence is more useful than unrelated clearance summaries.
Ask for configuration
Which model version, threshold, scanners, population, and reader workflow were studied?
Ask about failure and change
How are misses, drift, updates, downtime, and local disagreement handled?
How is the handbook checked?
Trace the source
Connect the original paper, attribution, result, and exact sentence.
Review the implication
Check that design, endpoint, setting, and limitations survive the summary.
Keep corrections inspectable
An open reference benefits from source links and a record of changes.
What is the economic value?
Define the perspective
A department, hospital, payer, and patient can experience different costs and benefits.
Distinguish modeled from measured value
Time saved is only one input, alongside integration, support, error, and follow-up costs.
Evaluate the local workflow
Estimate value using the actual task, volume, staffing, and downstream consequences.
What about subgroup performance?
Inspect who is missed
Overall performance can conceal clinically relevant subgroup errors.
Check the local population
Development-set fairness need not persist under distribution shift.
Monitor after deployment
Pair subgroup review with scanner, version, threshold, and workflow records.
Sources used in this talk.
- Martinez-Gutierrez et al., JAMA Neurol 202310.1001/jamaneurol.2023.3206
- Lång et al., Lancet Oncol 202310.1016/S1470-2045(23)00298-X
- Gommers et al., Lancet 202610.1016/S0140-6736(25)02464-X
- Eisemann et al., Nat Med 202510.1038/s41591-024-03408-6
- Elías-Cabot et al., Nat Med 202610.1038/s41591-026-04277-x
- Nam et al., Radiology 202310.1148/radiol.221894
- Storey et al., Radiol AI 202610.1148/ryai.250964
- Woznitza et al., Nat Med 202610.1038/s41591-026-04253-5
- Larsen et al., Eur Radiol 202610.1007/s00330-025-12240-6
- Sivakumar et al., JAMA Netw Open 202510.1001/jamanetworkopen.2025.42338
- Yu et al., Nat Med 202410.1038/s41591-024-02850-w
- Prinster et al., Radiology 202410.1148/radiol.233261
Sources used in this talk (2).
- Brady et al., Radiol AI 202410.1148/ryai.230513
- Lehman et al., JAMA Intern Med 201510.1001/jamainternmed.2015.5231
- Blum et al., Insights Imaging 202610.1186/s13244-026-02383-5
- Zhou et al., NEJM AI 202610.1056/AIoa2500595
- Zhang et al., Science 202610.1126/science.aec6129
- Acar et al., npj Digit Med 202610.1038/s41746-026-03169-1
- Lim et al., Eur Radiol 202610.1007/s00330-026-12648-8
- Ambinder et al., AJR 202610.2214/ajr.26.34611
- Langlotz et al., preprint 202510.64898/2025.12.20.25342714
- Mello et al., NEJM 202410.1056/NEJMhle2308901
- Budzyń et al., Lancet Gastro Hepatol 202510.1016/S2468-1253(25)00133-5
- Okumura et al., Endoscopy 202510.1055/a-2661-2624
Sources used in this talk (3).
- Pedersen et al., Endoscopy 202610.1055/a-2858-7084
- Huang et al., JAMA Netw Open 202510.1001/jamanetworkopen.2025.13921
- Lawrence et al., eClinicalMedicine 202510.1016/j.eclinm.2025.103228
- Molwitz et al., Radiol AI 202610.1148/ryai.250090
- Seyyed-Kalantari et al., Nat Med 202110.1038/s41591-021-01595-0
- Yang et al., Nat Med 202410.1038/s41591-024-03113-4
- Larrazabal et al., PNAS 202010.1073/pnas.1919012117