Yale
Diagnostic Radiology
Fall 2026 Lecture Series

The Physician AI Handbook

Clinical AI
in radiology.

What works. What doesn't.
What happens when AI is wrong.

November 18, 2026
7:30 AM ET · Zoom
Opening

Four questions for this morning.

01

Why build a clinical AI handbook?

The gap between a published result and a decision in your department.

02

What does the radiology evidence support?

Trials, workflow studies, and the limits of each result.

03

What happens when the system is wrong?

Three teaching cases, then the controls that matter.

04

How do we keep evidence useful?

Source verification, version changes, and questions for the next studies.

Genesis

Published accuracy is one part of the decision.

Model performance does not establish
patient benefit.

Model

Can it perform the stated task?

Workflow

Does it change decisions or time to action?

Patient

Does it improve outcomes or reduce harm?

Genesis

An open reference for clinical decisions.

Read the evidence

Clinical studies, their designs, and the outcomes they measured.

Evaluate a tool

Intended use, local validation, human factors, and monitoring.

Revisit the claim

Sources, versioned changes, and corrections that readers can inspect.

The sources can be peer reviewed. The handbook is an open, versioned reference.

Genesis

Four questions before accepting a claim.

01

What is the original source?

Read the study or regulatory record, not only the summary.

02

What outcome was measured?

Keep detection, workflow, patient benefit, and cost separate.

03

Where does the result apply?

Identify the product, version, population, and clinical setting.

04

What can the source not establish?

Keep limitations beside the result, where they affect the decision.

Evidence

Choose the endpoint for the intended use.

Triage

Time to interpretation
or action

Include delays and misses in unflagged studies.

Screening

Detection, recall,
interval disease

A higher detection rate alone does not establish mortality benefit.

Report drafting

Final-report accuracy
and editing burden

Measure consequential errors alongside time saved.

Evidence · Stroke triage

Faster thrombectomy. Patient benefit remains uncertain.

−11.2 min

Adjusted difference in door-to-groin time

Stepped-wedge cluster-randomized trial
243 thrombectomy patients

What the trial supports

A shorter treatment-process interval after AI activation.

What remains uncertain

The trial did not demonstrate a significant improvement in 90-day functional independence.

A nonsignificant result is not proof that the outcome is unchanged.

Evidence · Mammography

Three studies. Three different interventions.

MASAI

Randomized · Sweden
80.5%

Sensitivity with AI support, versus 73.8% with standard double reading.

Specificity: 98.5% in both groups.

PRAIM

Prospective observational · Germany
6.7 /1,000

Adjusted cancer detection with AI, versus 5.7 per 1,000 without.

Radiologists chose when to use AI.

AITIC

Paired noninferiority · Spain
+14.8%

Relative increase in recall. Recall noninferiority was not met.

AI triage was also checked against standard readings.

Detection, recall, interval cancers, and workload answer different questions.

Evidence · Chest radiographs

More detections do not guarantee a faster pathway.

01

Detection trial

Nam: actionable nodules detected in 0.59% with AI versus 0.25% without.

02

Prospective silent evaluation

Storey: outputs were hidden from readers. A silent study does not test autonomous live use.

03

LungIMPACT pathway trial

Added AI worklist prioritization did not significantly shorten time to CT or cancer diagnosis.

Evidence · Model versions

A version update changes the assessment.

The same 117,709 Norwegian examinations

Cancers assigned the highest AI score

Version 1.787.1%
Version 2.193.5%

Share of screen-detected cancers · 0–100% scale

The decision consequence

Different scores can send different examinations through a risk-adapted reading pathway.

The evidence limit

Retrospective scoring does not establish the net benefit of a new live workflow.

Evidence · Regulatory record

Clearance answers an intended-use question.

Read the label

What task is authorized?
For which patients and images?
With what role for the reader?

Read the evidence

What testing is described?
Was the workflow prospective?
Was the reader part of the test?

Evaluate locally

Does the evidence apply here?
What can fail?
Who can pause the tool?

Missing testing in a public document is not proof that testing never occurred.

Evidence · Human factors

Evaluate the reader and the model together.

Advice can mislead

Incorrect AI advice reduced radiologists’ aggregate performance in a controlled study.

Trust is an outcome

Local explanations increased reliance on simulated advice versus global explanations, including when advice was incorrect.

Sequence needs testing

Interpretation order, interface, and escalation are part of the workflow.

An explanation that sounds plausible is not a validation result.

Teaching case · Fictional

No flag appeared. The study entered the routine queue.

A trauma head CT has an urgent clinical indication.

The AI triage tool produces no flag.

The worklist moves the examination into a routine queue.

What should have determined its priority?

Teaching case · Discussion

The clinical indication still determines priority.

01

Retain the required turnaround

An absent AI flag does not establish that the study is normal.

02

Test the entire queue

Monitor delayed unflagged studies as well as accelerated flagged studies.

03

Give the department control

Define escalation and a way to pause unsafe prioritization.

Teaching case · Fictional

Your review is benign. AI flags an asymmetry.

A screening mammogram includes a prior examination.

Your initial assessment is benign.

The CAD output identifies an asymmetry for review.

What evidence would change your assessment?

Teaching case · Discussion

Re-review the finding. Own the final assessment.

01

Check the finding

Review the location, imaging features, priors, and technical adequacy.

02

Use the appropriate category

The vignette alone cannot determine a BI-RADS category.

03

Document the reasoning

Record why the AI mark changed, or did not change, the final assessment.

Teaching case · Fictional

The scanner changed. The AI assessment did not.

A department introduces a different CT platform.

The deployed AI was evaluated on the previous acquisition setup.

The equipment change did not trigger an AI impact assessment.

Which control should connect these decisions?

Teaching case · Discussion

Make equipment change an AI review trigger.

01

Assess what changed

Scanner, reconstruction, preprocessing, PACS, model, or threshold.

02

Test before relying on the new setup

Agree on validation, acceptance, and pause criteria with clinical owners.

03

Monitor after the change

Review discordant cases and alert patterns with authority to intervene.

Maintenance

A living reference also needs change control.

Clinical AI questionReference question
Does the tool apply to this setting?Does the source support this sentence?
Did the deployed version change?Did the evidence or guidance change?
Can a failure be detected and corrected?Can a reader inspect the source and correction?
Maintenance

Keep the source attached to the claim.

01

Find

Treat a new paper or report as a lead.

02

Verify

Read design, outcome, figure, and limitations.

03

Review

Inspect the exact wording and its implication.

04

Update

Record the change and check the published result.

Automated drafting can accelerate work. It cannot establish source truth.

Maintenance

A working link can still support the wrong claim.

01

Right DOI, wrong attribution

The link resolves, but the cited author or journal is wrong.

02

Right study, wrong design

An observational comparison becomes “a trial” in the summary.

03

Right number, wrong conclusion

A workflow improvement becomes a claim about patient outcomes.

Future · Generalist models

Broad benchmarks leave clinical questions open.

MedVersa

One model across imaging and report tasks.

Developer studies require prospective validation in the intended clinical workflow.

RADAR

Broad abdominal CT finding detection.

Research evaluation across 146 findings does not establish autonomous clinical use.

Consumer models

Curated case answering is a different task.

Open-ended primary accuracy: 51.6–62.7% on the tested RSNA cases.

Task-specific errors and prospective workflow performance still matter.

Future · Model change

Evidence belongs to the deployed configuration.

01

Model version

02

Threshold

03

Acquisition

04

Reader workflow

When a consequential component changes,
reopen the assessment.

Future

Questions an academic department can answer.

01

Reporting

Do drafts reduce work without increasing consequential final-report errors?

02

Patient pathways

Does a prioritization change improve time to meaningful clinical action?

03

Training

How does independent interpretation change with sustained AI exposure?

04

Post-market monitoring

Which performance changes follow a new model, threshold, or scanner?

Impact

Three gates before clinical use.

01

Evidence fits the use

Population, setting, data, threshold, and outcome match the decision.

02

Risk controls are defined

Labeling, failure modes, reader responsibilities, and escalation are clear.

03

Monitoring can change practice

Metrics, review ownership, and pause criteria exist before deployment.

An unresolved gate calls for corrective work and reassessment.

Impact

Define monitoring before deployment.

Before

Validate the exact use.

Set acceptance criteria.

Map likely failure modes.

During

Test integration.

Train for disagreement.

Make errors reportable.

After

Track errors and drift.

Review local subgroups.

Reassess consequential changes.

Closing

Three things to keep.

01

Match the endpoint to the claim.

Accuracy, workflow, and patient benefit are different questions.

02

Tie evidence to the configuration.

Version, threshold, acquisition, and reader workflow matter.

03

Keep evaluating after deployment.

Detect change, review failures, and retain the ability to pause.

Discussion

Questions and discussion.

Backup · Discussion

Will AI replace radiologists?

Back to discussion
01

Start with tasks

The evidence in this talk supports assistance or automation of selected tasks.

02

Do not turn scenarios into forecasts

Workforce models depend on demand, payment, staffing, and deployment assumptions.

03

Retain clinical functions

Interpretation, procedures, consultation, communication, and quality assurance are different tasks.

Backup · Discussion

How should we validate locally?

Back to discussion
01

Match the intended use

Sample relevant local patients, scanners, image types, and clinical workflows.

02

Define the reference standard

Use adjudication suited to the task, including consequential errors.

03

Set decisions in advance

Acceptance, escalation, monitoring, and pause criteria should precede the pilot.

Backup · Discussion

Who is liable when AI misses?

Back to discussion
01

Keep the legal claim narrow

Responsibilities depend on facts, jurisdiction, contracts, and the clinical context.

02

Separate authorization from liability

FDA authorization does not resolve who bears civil liability in a particular case.

03

Make ownership explicit

Discuss clinical duties, documentation, vendor terms, and coverage with institutional counsel.

Backup · Discussion

What about resident deskilling?

Back to discussion
01

Distinguish concern from demonstration

The cited longitudinal signals are from endoscopy, not radiology residency.

02

Measure independent performance

Assisted accuracy alone cannot show whether unassisted skill is preserved.

03

Test safeguards

Independent interpretation, disagreement review, and failure recognition are training options to evaluate.

Backup · Discussion

Are model-drafted reports ready?

Back to discussion
01

Workflow signal

A prospective matched cohort associated assistance with 15.5% less documentation time.

02

Important boundary

The study covered one 12-hospital system. It was not randomized patient-outcome evidence.

03

Measure the final report

Evaluate consequential errors, editing burden, omissions, and failure handling.

Backup · Discussion

Which vendor is best?

Back to discussion
01

Require the right comparison

Independent head-to-head evidence is more useful than unrelated clearance summaries.

02

Ask for configuration

Which model version, threshold, scanners, population, and reader workflow were studied?

03

Ask about failure and change

How are misses, drift, updates, downtime, and local disagreement handled?

Backup · Discussion

How is the handbook checked?

Back to discussion
01

Trace the source

Connect the original paper, attribution, result, and exact sentence.

02

Review the implication

Check that design, endpoint, setting, and limitations survive the summary.

03

Keep corrections inspectable

An open reference benefits from source links and a record of changes.

Backup · Discussion

What is the economic value?

Back to discussion
01

Define the perspective

A department, hospital, payer, and patient can experience different costs and benefits.

02

Distinguish modeled from measured value

Time saved is only one input, alongside integration, support, error, and follow-up costs.

03

Evaluate the local workflow

Estimate value using the actual task, volume, staffing, and downstream consequences.

Backup · Discussion

What about subgroup performance?

Back to discussion
01

Inspect who is missed

Overall performance can conceal clinically relevant subgroup errors.

02

Check the local population

Development-set fairness need not persist under distribution shift.

03

Monitor after deployment

Pair subgroup review with scanner, version, threshold, and workflow records.

Backup · Sources

Sources used in this talk.

Back to discussionMore sources
Backup · Sources

Sources used in this talk (2).

Back to discussion
Backup · Sources

Sources used in this talk (3).

Back to discussion