Start with the thing being measured. Pivot the timeline on the note and both sides go dark — nobody hydrates the context coming in, nobody tracks the inbox going out. The note itself is 12.5% of the arc.
The clinical process runs from booking to closure. Every tool sits in one narrow band of it — and three bands have nothing in them at all.
The encounterclinical process
Schedulingwhy they booked
Prechartingread the whole chart
Historythe interview
Examhands on the patient
CDSthe decision
Documentationthe note
Follow-updid it happen?
booksdays beforeIN THE ROOMafterweeks → months
Answer enginesOE · DoxGPT
the question, answeredyou gather and paste the context by hand
ScribesAbridge · Ambience
hears the room → writes the notenever reads the chart — no precharting
Nobodyno product, no eval
?
?
?
The three empty bands are not mysteries — they are where the questions already go. Classified and counted, the stream is the workflow eval those ? boxes are waiting for.
the ? boxes, answered · praxis — question stream → intent → subintent · de-identified · counts land here as the taxonomy matures
lands before the noteContext questions
“what changed since her last visit?”
“summarize the outside cardiology notes”
“which meds still need reconciling?”
interval history · chart synthesis · med rec
lands at the notePoint-of-care questions
“max metoprolol dose in CKD?”
“differential for this presentation?”
“what does the guideline say here?”
med dosing · differential · guideline lookup
lands after the noteFollow-through questions
“draft a message explaining this result”
“handle this refill request”
“write the prior-auth appeal”
result interpretation · patient messaging · admin drafting
The compute belongs on the dark sides of the note — hydrate the chart coming in, track the inbox and closure going out. Every question above is a vote for where it goes; classified and counted, the stream becomes the workflow eval the ? boxes are waiting for. And learn how clinicians actually use it before trying to teach them. One recommendation, walked end-to-end →
The chart is the whole picture. Every tool hands you a lens instead — so sweep it around and
see how little comes into focus, and how long it takes.
simulated questions
ask a question, or sweep the glass yourself
patient context 0
Empty. Whatever you hand the answer engine, this is
all of it — and right now it is nothing.
0 of 0 items in context
You uncover only what you thought to look at, in whatever shape your
sweep happened to make. No benchmark scores what never came into focus.
Run the same patient twice. Identical engine, identical question — the only difference is
whether anybody swept the chart first.
the same patient, twice · illustrative
⏺ without the chart
T+0
“Cellulitis — what antibiotic?”the same question, both times
T+0
Context assembled: 2 itemsage, sex, the complaint
T+0
The chart is not sweptthe 2019 outside ED record is never opened
2019-04 · OUTSIDE ED · amoxicillin → ANAPHYLAXIS
T+2 min
Prescribed: amoxicillin–clavulanatea guideline-standard choice for cellulitis
T+90 min
outcome
Anaphylaxis
epinephrine · ED · admitted overnight. The allergy was in the chart the whole time.
⏺ with the chart
T+0
“Cellulitis — what antibiotic?”the same question, both times
T+0
Context assembled: 38 itemsincluding the scanned outside records
T+0
The chart is sweptthe 2019 outside ED record surfaces
2019-04 · OUTSIDE ED · amoxicillin → ANAPHYLAXIS
T+2 min
Prescribed: doxycyclinethe guideline’s choice when penicillin is out
T+6 d
outcome
Resolved
no reaction, no return visit. Nothing about the model was different.
Both prescriptions are correct answers to the question that was asked. Score either
one against a rubric and it passes — the drug is guideline-appropriate for cellulitis, the reasoning is sound,
the citation is real. The harm is upstream of everything a benchmark can see.
The question that never gets asked
2
The precharting stream, filtered by the doctor, classified by praxis — and the questions that never form.
speaker notes
Every gray dot is a before-visit inquiry — unscheduled, uncounted, unpaid; most of the chart is never reviewed. The doctor is the filter: attention, time, and what they know to ask — only a fraction enters, and color is assigned on the way out. praxis classifies what emerges (intent → subintent): extraction = interval history · med rec · outside records · guideline lookup; reasoning = differential · risk synthesis · dose in context · goals of care. Both lanes end in the answer-engine log, counted as product usage (“1M consultations a day”) — the same cognition, relocated; the tool takes credit. Never formed — the red dots in the pile: “could this be ATTR amyloid?” exists only if you’ve read the 2026 guidance — an unasked question is indistinguishable from no need.
Every instrument ever built for medical AI aims at that 12.5% — the answer at the moment of the note. Here they are, era by era, and here is the same chart recolored to ask who each one is for and who gets to see what it produces.
McLaren, 2026 Austrian GP · Lukas Raich, CC BY-SA 4.0
the engine is necessary · the package sets the time
One Mercedes power unit goes into four teams in 2026. They do not finish
anywhere near each other. Every instrument on the chart grades the engine.
Four eras of measuring medical AI
3
2019–24Exam era“Are you booksmart?”
Multiple-choice recall — integrated retrieval, not practice. Saturates at 84–90%, at or above physician level; falls to 45–69% on clinical tasks. Ends as marketing.
2023–Rubric era“Can it converse safely?”
Physician-written rubrics over multi-turn conversations — a curated vignette stands in for the patient.
2024–Task era“Can it do clinical work?”
Real EHR data and agentic workflows — graded on output quality, never on adoption.
2025–Deployment era“What does deployment show — and did it change the patient?”
Real use and real outcomes, one era: benchmarks curated from live telemetry, while the outcome question runs daily as private telemetry.
patient-chat long tail, not charted: K Health · Ada · Babylon · HealthTap · Woebot · symptom checkers & wellness bots
JAN 2026
The instruments, as they arrived
January 2026 — where the exam era actually endsPatients had been asking for three years — 230 million health questions a week — before any instrument measured it. Then the labs stopped selling scores and built for that instead, in a single week: ChatGPT Health on the 7th, Claude for Healthcare on the 12th. The question stops being can it pass? and becomes what is it doing?
Two questions, one chart
4
The argument
Medicine already knows how to close a loop — it does it thousands of times a day, and chases the ones left open. The AI recommendation is the one event in the chart that never got a close. The reason is access, not rigour.
Every workflow the chart trusts is a ticket. Something opens; something specific is allowed to close it. Open-without-close isn’t a gap — it’s an error the system chases until it dies.
The encounter
Visit beginsOPEN
the visit
Note signedCLOSED
The lab order
Order placedOPEN
specimen → analyzer
Result filedCLOSED
Every open has a designated close — and the chart hunts unclosed tickets for a living.
The ticket nobody closes
6
Run the same schematic on an AI recommendation. The event happens. The action happens. And the ticket that should close it — the follow-up, the outcome — is never even opened. Not failed: unlogged.
The AI answer
RecommendationOPEN
clinician acts
Order · script · plan
…then?
Follow-up · outcomeSTILL OPEN
The close exists in the workflow — not in the record. Treat the recommendation like a lab order: open it, watch it, close it on evidence.
Everyone grades the model. No one grades the outcome.
7
Every instrument and product on the chart lives on one side of a wall: they grade artifacts of the session — answers, notes, sandboxed orders. Whether the thing was done lives on the other side, inside the chart. The literature is lopsided in exactly this shape: of 4,609 clinical LLM studies, 1,048 touched real patient data and 19 were prospective randomised trials; an earlier review found 5% used real patient-care data at all. Governance has noticed — CHAI and the Joint Commission now require monitoring for “changes in outcomes” — but a mandate is not an instrument: none of it says what to measure, over what window, against what denominator. The wall is EHR access.
Grading the model
0 instruments & products — the full timeline. everyone.
THE WALL — EHR · PATIENT-DATA ACCESS
Grading the outcome
the same timeline — one entry here; 19 prospective trials in 4,609 studies
The follow-up lives in the EHR. Almost no evaluator can see the EHR. That’s the whole story.
Era
when
who it's for
what it grades
who sees the data
the loop
Exam era 2019–24
item recall
clinician-facing
session artifacts
academia
never checked
Rubric era 2023–
conversation
patient-facing
session artifacts
frontier labs
never checked
Task era 2024–
task
clinician-facing
sandboxed actions
academia
never checked
Deployment era 2025–
usage → outcome
clinician-facing → the whole workflow
session artifacts → patient outcome
vendor streams
the only place it could close — today, privately
Product launches '22–'26
ChatGPT (Nov '22) · Doximity GPT (Feb '23) · Abridge in Epic ('23) · DAX Copilot (Sep '23) · MedGemma (Jul '25) · Doximity Scribe (Jul '25) · OpenEvidence Visits (Aug '25) · Epic Art & Emmie (Aug '25) · ChatGPT Health (announced Jan 7 '26, US GA Jul '26) · Claude for Healthcare (Jan 12 '26, JPM) · Doximity Clinical AI Suite (May '26) · EvidenceGrade (Jul '26) — the ink squares on every lens; under the data lens they join the vendor streams, under the loop they check nothing