- Legal
AI in Medical Record Review: What Is Actually Solved, What Is Not, and How to Tell the Difference
Table of Contents
What the market looks like now
We recently reviewed twenty vendors in this market for our own competitive research. Four kinds of providers compete for the work, and they fail in different ways.
AI-native platforms sell software. You upload, the system returns a chronology, and the pricing and speed are excellent. Verification is generally your responsibility, sometimes framed as a feature: the tool surfaces confidence scores, and you decide what to check.
Hybrid service vendors pair an extraction engine with a review team, frequently offshore. Cost is low, turnaround is competitive, and the quality of the human layer varies enormously and is difficult to assess from outside.
Legal nurse consultant networks put clinical judgment first and technology second. The review quality is often excellent, but the model does not scale well to a 20,000-page file on a litigation deadline.
Bundled document services treat record review as one line item beside retrieval, copying, and Bates stamping. Convenient, but rarely specialized.
None of those categories is wrong. They make different tradeoffs, and a firm running twelve files a year has genuinely different needs from a carrier running twelve hundred. What they share is that the extraction step is no longer where they differ.
What AI has genuinely solved
It is worth being direct about this, because vendor content tends to either oversell AI or nervously undersell it.
Machine extraction is now reliably better than human review at the mechanical parts of this work. Separating a merged PDF into discrete documents. Removing the duplication that multi-custodian requests reliably produce, where the same discharge summary arrives four times in four scan qualities. Dating records. Building a chronological spine. Reading twenty thousand pages without fatigue, at hour nine, with the same attention as at hour one.
A person doing that work is slower, more expensive, and more error-prone, because the tasks are exactly the kind humans do badly. Any argument that record review should not be automated at the extraction layer is not an argument about quality. It is nostalgia.
What AI has not solved
Verification. And the industry is watching the wrong failure mode.
Purpose-built legal tools are not exempt. In Fletcher v. Experian, counsel used vLex and CoCounsel and still filed sixteen fabricated quotes. The sanctioned Morgan & Morgan citations came from the firm’s own in-house platform rather than a consumer chatbot.
In medical summarization, hallucination is not the main failure. Omission is.
Why omission is the harder problem
Six questions worth asking any vendor
-
Who verifies the output, and what are their credentials?
Not who could verify it. Who does, on the file you are about to send. -
Is every file verified, or a sample?
Sampling is a legitimate quality control method for manufacturing. It is a weak fit for litigation, where the file that matters is the one in front of you rather than the average file. -
Can every entry be traced to a source page?
If a chronology entry cannot be clicked or cited back to the underlying document, it is an assertion rather than evidence. Under a substantial evidence standard, that distinction decides motions. -
What is your process for completeness, as distinct from accuracy?
This is the omission question, and most vendors have no good answer, ourselves included, because quality processes in this industry were built to catch errors of commission. Partial controls exist. A page-stamped index grouped by provider lets a reviewer see custodian coverage and spot where a date range breaks. Ask what the vendor actually does, and treat a confident answer with more suspicion than a candid one. -
Who owns the output format?
A defense apportionment review and a plaintiff damages narrative need the same records organized differently. A vendor that will only deliver in its own template is asking your team to adapt to its software. -
What does your published accuracy figure actually measure?
Nearly every accuracy percentage in this market is vendor-generated, unaudited, and measured against a benchmark the vendor chose. Ask what the denominator was, who scored it, and whether omissions counted as errors. The answers are usually instructive.
A note on where we stand
Frequently Asked Questions
Is AI accurate enough for medical record review?
What is the difference between a hallucination and an omission in a medical chronology?
Do purpose-built legal AI tools still hallucinate?
What should I ask a medical record review vendor before sending a file?
How should I read a vendor's published accuracy percentage?
Does a preexisting disability have to be work related?
No. Section 4750(e)(1)(B) expressly contemplates a nonindustrial impairment that could support an award of permanent partial disability, provided it meets the labor-disabling test.