AI in Medical Record Review: What Is Actually Solved, What Is Not, and How to Tell the Difference

Three years ago, telling a workers’ compensation defense firm that AI could build a medical chronology was a sales pitch. Today it is an assumption. Every serious vendor in this market has an extraction engine, and most have had one long enough to have worked the obvious bugs out of it.

Table of Contents

That changes the buying question. When every provider can process the record, the differentiator is no longer whether the technology works. It is what happens to the output between the model and your desk.
Most firms evaluating record review vendors right now are still asking the 2023 question. Here is what we think the 2026 question is, and how to test for it.

What the market looks like now

We recently reviewed twenty vendors in this market for our own competitive research. Four kinds of providers compete for the work, and they fail in different ways.

 

AI-native platforms sell software. You upload, the system returns a chronology, and the pricing and speed are excellent. Verification is generally your responsibility, sometimes framed as a feature: the tool surfaces confidence scores, and you decide what to check.

 

Hybrid service vendors pair an extraction engine with a review team, frequently offshore. Cost is low, turnaround is competitive, and the quality of the human layer varies enormously and is difficult to assess from outside.

 

Legal nurse consultant networks put clinical judgment first and technology second. The review quality is often excellent, but the model does not scale well to a 20,000-page file on a litigation deadline.

 

Bundled document services treat record review as one line item beside retrieval, copying, and Bates stamping. Convenient, but rarely specialized.

None of those categories is wrong. They make different tradeoffs, and a firm running twelve files a year has genuinely different needs from a carrier running twelve hundred. What they share is that the extraction step is no longer where they differ.

What AI has genuinely solved

It is worth being direct about this, because vendor content tends to either oversell AI or nervously undersell it.

 

Machine extraction is now reliably better than human review at the mechanical parts of this work. Separating a merged PDF into discrete documents. Removing the duplication that multi-custodian requests reliably produce, where the same discharge summary arrives four times in four scan qualities. Dating records. Building a chronological spine. Reading twenty thousand pages without fatigue, at hour nine, with the same attention as at hour one.

 

A person doing that work is slower, more expensive, and more error-prone, because the tasks are exactly the kind humans do badly. Any argument that record review should not be automated at the extraction layer is not an argument about quality. It is nostalgia.

What AI has not solved

Verification. And the industry is watching the wrong failure mode.

The legal profession’s AI conversation is dominated by hallucination, for understandable reasons. Hallucination draws sanctions, and the record is now substantial. A public database maintained by researcher Damien Charlotin has cataloged well over a thousand court decisions in which AI-fabricated material reached a court, and it grows weekly. Penalties have moved from a $5,000 fine in Mata v. Avianca in 2023 to five-figure sanctions, bar suspensions, and a canceled trial in 2026.
Two details in that record deserve more attention than the totals.

Purpose-built legal tools are not exempt. In Fletcher v. Experian, counsel used vLex and CoCounsel and still filed sixteen fabricated quotes. The sanctioned Morgan & Morgan citations came from the firm’s own in-house platform rather than a consumer chatbot.

The best evidence on this is a preregistered study by Magesh and colleagues at Stanford’s RegLab and Institute for Human-Centered AI, peer-reviewed and published in the Journal of Empirical Legal Studies in 2025. Testing more than 200 hand-scored legal queries, they found Lexis+ AI hallucinated on roughly 17 percent of queries and Westlaw AI-Assisted Research on roughly 33 percent, against 43 percent for GPT-4. Retrieval-augmented generation reduced hallucination. It did not remove it.
Two findings from that study matter more here than the headline rates. LexisNexis had marketed “100% hallucination-free linked legal citations,” and Thomson Reuters had said its tools avoid hallucinations by relying on trusted content. The researchers’ conclusion was that providers’ claims are overstated. And Thomson Reuters’s Ask Practical Law AI returned incomplete answers, meaning refusals or ungrounded responses, on more than 60 percent of queries, the highest rate of any system tested.
That last figure is the one to sit with. In the only independent, peer-reviewed evaluation of commercial legal AI, the most common failure was not invention. It was incompleteness.

In medical summarization, hallucination is not the main failure. Omission is.

A 2026 study in npj Health Systems gathered physician feedback on 147 AI-generated chart summaries from a tool integrated into an electronic health record. Reviewers flagged omissions 46 times and hallucinations 5 times, and the authors concluded that errors of omission may pose a larger threat to accuracy and usefulness than errors of commission. A University of California, San Francisco study in PLOS Digital Health evaluated GPT-4 across 100 emergency department encounter summaries: 47 percent omitted clinically relevant information, 42 percent contained a hallucination, and only a third were entirely error-free. A framework study in npj Digital Medicine, built on nearly 13,000 clinician-annotated sentences, measured a 1.47 percent hallucination rate against a 3.45 percent omission rate.
Those studies examine clinical summarization, not litigation record review, and the models tested are not current production systems. The direction of the finding is what transfers, and it is consistent across all three.

Why omission is the harder problem

A hallucination announces itself. Someone checks the citation, the case does not exist, the error surfaces. Embarrassing, occasionally sanctionable, and detectable by anyone willing to look.
An omission announces nothing. There is no entry to check. A chronology that quietly leaves out a 2011 note documenting a modified duty assignment reads as complete, reviews as clean, and survives every spot check, because spot-checking tests entries that are present. You cannot spot-check an absence.
This is why confidence scores solve less than they appear to. A confidence score tells you how sure the model is about something it produced. It says nothing about what it never produced.
In workers’ compensation and personal injury work, that is the failure that costs cases. The document that decides a claim is often not the one describing the injury. It is the older, duller record establishing what the person could and could not do beforehand.
California’s SB 171, signed in July 2026, made this concrete for anyone handling Subsequent Injuries Benefits Trust Fund claims. Eligibility now turns on substantial evidence drawn from records that existed before the subsequent industrial injury, and the statute expressly bars establishing a preexisting disability through a retroactive work restriction. A missing pre-injury document cannot be cured by an expert opinion. We covered the details in our analysis of SB 171.
Volume compounds all of it. When a file runs 20,000 pages across a dozen custodians and four decades, nobody can hold the record in mind well enough to notice what is not there. That is exactly the condition in which people stop checking and start trusting.

Six questions worth asking any vendor

These are the questions we would ask, including of ourselves.
  1. Who verifies the output, and what are their credentials?
    Not who could verify it. Who does, on the file you are about to send.
  2. Is every file verified, or a sample?
    Sampling is a legitimate quality control method for manufacturing. It is a weak fit for litigation, where the file that matters is the one in front of you rather than the average file.
  3. Can every entry be traced to a source page?
    If a chronology entry cannot be clicked or cited back to the underlying document, it is an assertion rather than evidence. Under a substantial evidence standard, that distinction decides motions.
  4. What is your process for completeness, as distinct from accuracy?
    This is the omission question, and most vendors have no good answer, ourselves included, because quality processes in this industry were built to catch errors of commission. Partial controls exist. A page-stamped index grouped by provider lets a reviewer see custodian coverage and spot where a date range breaks. Ask what the vendor actually does, and treat a confident answer with more suspicion than a candid one.
  5. Who owns the output format?
    A defense apportionment review and a plaintiff damages narrative need the same records organized differently. A vendor that will only deliver in its own template is asking your team to adapt to its software.
  6. What does your published accuracy figure actually measure?
    Nearly every accuracy percentage in this market is vendor-generated, unaudited, and measured against a benchmark the vendor chose. Ask what the denominator was, who scored it, and whether omissions counted as errors. The answers are usually instructive.

A note on where we stand

We build an AI record review product of our own, so publishing a buyer’s checklist plainly invites you to turn it on us. That seemed like the right trade. The questions above are the ones we think matter, and we would rather be judged against them than against any single number.
WHITE is our record review product, and it opens to early access shortly. When it does, the honest test is a file your team has already reviewed: compare the output, then check the entries that decide the case against the source pages.

Frequently Asked Questions

Is AI accurate enough for medical record review?

For extraction, deduplication, dating, and organization, machine processing is now more reliable than manual review, particularly at volume. The unresolved problem is verification. Peer-reviewed studies of AI medical summarization consistently find omission rates exceeding hallucination rates, and omissions are difficult to detect because there is no entry to check. Accuracy therefore depends less on the model than on what review sits between the output and delivery.

What is the difference between a hallucination and an omission in a medical chronology?

A hallucination is content the system produced that the source record does not support, such as a diagnosis or a date that does not appear in the file. An omission is relevant content the system left out. Hallucinations are detectable by checking entries against the source. Omissions are not, because there is no entry to check, which is why they carry the greater risk in litigation record review.
Yes. The preregistered Stanford study by Magesh and colleagues found hallucination rates of roughly 17 percent for Lexis+ AI and roughly 33 percent for Westlaw AI-Assisted Research, against 43 percent for GPT-4. Retrieval-augmented generation reduced the problem without eliminating it, and the researchers concluded that providers’ claims about being hallucination-free were overstated.

What should I ask a medical record review vendor before sending a file?

Who reviews the output and with what qualifications, whether every file is reviewed or only a sample, whether every entry traces to a source page, what the process is for completeness as distinct from accuracy, who controls the output format, and what any published accuracy figure actually measures.

How should I read a vendor's published accuracy percentage?

Ask what it measures before you compare it to anything. Most figures in this market are measured internally against a benchmark the vendor selected, which means two vendors’ numbers are rarely comparable even when both are honest. Ask what the denominator was, who scored it, and whether omissions counted as errors or only incorrect entries. A single percentage cannot separate hallucination from omission, and the Stanford researchers who tested commercial legal AI called for independent, preregistered, publicly updated benchmarking precisely because self-reported figures could not be verified. No such benchmark yet exists for medical record review.

No. Section 4750(e)(1)(B) expressly contemplates a nonindustrial impairment that could support an award of permanent partial disability, provided it meets the labor-disabling test.

Scroll to Top