A generated summary of a medical record and a verified record can look identical on the page. Both are readable. Both are organized. Both use professional language and appropriate hedging. Both present a chronology that hangs together.
A skilled reader handed both cannot tell them apart by reading them—and this is not a criticism of skilled readers. There is nothing in the text to tell them apart with.
The problem in one sentence
Fluency is a property of language. Evidence is a property of the relationship between a statement and a source. These are independent. A system optimized for the first can produce outputs that are indistinguishable from the second while carrying none of its guarantees.
The word “summary” hides this. A summary is generated: a natural-language synthesis describing observed evidence. It is not itself evidence, no matter how accurate it happens to be, because there is no way to check it without going back to the source—and going back to the source is precisely what the summary was meant to save you from.
Cooperative versus adversarial settings
In a cooperative setting—a clinician reading a summary, applying judgment, and moving on—a 95% accurate summary is genuinely useful. Errors cost a second look. The reader has context, catches most problems, and the system is a net gain even with a known error rate.
In an adversarial setting, the arithmetic inverts. Take a three-thousand-page file yielding two thousand extracted facts: dates, providers, diagnoses, findings, prior injuries. At 95%, that is a hundred wrong facts sitting in a record someone is about to testify from. Opposing counsel needs one—and they will read it more carefully than you did, because finding that one thing is their entire job for the week.
An expert opinion is only as defensible as its weakest sourced fact, not its average. That is why aggregate accuracy is the wrong frame for this work. Nobody challenges a report in aggregate.
Make the epistemic status visible
The requirement is not a better summary. It is a different kind of output: one where the epistemic status of every statement is visible.
Direct observation. Visible in the source: medication listed, page 184. Highest confidence and checkable in seconds.
Structured extraction. Pulled from the source and normalized, while still grounded in something observable.
Supported inference. Requires combining observations, such as a treatment timeline built across providers. Reasonable and legitimate, but interpretive.
Generated narrative. Natural-language synthesis describing the evidence. Readable, but not evidence.
Human judgment. Clinical interpretation, legal opinion, and diagnostic reasoning. Always the professional’s, never the system’s.
Most tools collapse all five into one undifferentiated register: everything presented with the same confidence, in the same voice. The user is left to guess which parts they can rely on and which parts they need to check. In practice, that means either checking everything—which defeats the purpose—or checking nothing, which is how the one wrong fact gets to a deposition.
Keeping the levels distinct costs something. Output looks less clean. Some statements carry visible caveats. The system says “I can’t source this” where a less careful one would simply assert.
That is the trade. In adversarial work, it is not close. A system that never declines to answer is not confident. It just is not checking. Ninety-five percent accurate is a great demo and a terrible deposition.
© 2026 Attunement. All rights reserved.
Built in San Francisco