What “Defensible” Should Mean When You’re Vetting a QME/IME Record Review Partner

The trouble with the word “defensible” is that you cannot test it at the moment you are buying.

Table of Contents

Price compares on a spreadsheet. Turnaround compares. Sample quality you can assess in twenty minutes with a coffee. Defensibility shows up in none of those columns. It shows up eleven months later, when a supplemental report gets requested, or in a deposition, when someone asks your evaluator where a specific value came from and the answer takes four seconds longer than it should have.

By then the vendor is not in the conversation. The evaluation is. And the defect is being paid for at its maximum price: in re-evaluations, in a supplemental report, in a case that slides two quarters to the right while the reserve sits open.

So the real job in vendor selection is to drag that failure forward. You want the defect to surface on a sales call instead of under oath. That takes a definition sharp enough to run a test against:

A record set is defensible when someone motivated to discredit it cannot.

Accuracy is a separate property. The gap between the two is where the cost lives, and most vendor evaluations never look at it.

Accurate and defensible are different things

A summary can be completely correct and still come apart under examination.
Say the date of injury in the summary is right. Good. Now someone asks where it came from. If answering means a person has to open the underlying record set and go hunting, that value was accurate and undefended. Its accuracy lived inside the vendor’s process rather than inside the document, and processes do not testify.
Defensibility is accuracy plus provenance plus a name, and all three have to live in the artifact itself. The artifact is what travels. The vendor does not travel with it.
That reframe earns its keep in procurement, because it converts a vague quality judgment into four properties you can test on a call. Can I trace this. Who verified it. How do they know their accuracy claim is true. And when something turns out wrong, what happens next.
Four questions. The rest of this is how to ask them and what the answers tell you.

First, retire "AI-powered" as a filter

Nearly every vendor in this category runs AI somewhere in the pipeline now. The label has stopped separating anyone from anyone, which means it cannot carry weight in an evaluation even though it still occupies the top of most homepages.
Two vendors can both be accurately described as AI-powered and be nothing alike. One runs extraction and ships the output the same hour: fast, inexpensive, nobody’s name attached. Another runs extraction, has a licensed reviewer verify the output against the source records before delivery, and tells you who that reviewer was. The second produces a record set. The first produces a draft that looks finished. They often price within a few dollars of each other and describe themselves in nearly identical language.
One question sits above all four of the ones below. Does the product reach conclusions? If any part of the pipeline generates opinions on causation, apportionment, compensability, or work restrictions, accuracy is no longer your primary concern. Those are determinations. They belong to the physician and the adjuster. A record product that reaches them has introduced an unattributed opinion into the file with no licensed author standing behind it, and that is the kind of thing opposing counsel builds a whole afternoon around.

Question 1: Can you trace every value back to a page?

Why it matters. Traceability is the property that survives cross-examination. “Where did this date come from?” should be answerable in seconds by anyone holding the document, including the person trying to break it. A summary that only the vendor can verify is a summary you cannot defend without the vendor, and the vendor is not available at 3pm on a Thursday in a deposition.

The test. Ask for a real redacted production deliverable before you sign anything. Not the demo file. Pick five entries at random: a date of injury, a medication, a work restriction, a diagnosis as stated by a treater, a treatment gap. Ask them to show you the source page for each one, live, while you are on the call.

Then push one level further. Does the citation use the same pagination the evaluator will have in front of them? A reference to an internal document ID that nobody outside the vendor can resolve is not a citation. It is a receipt for work you cannot inspect.

A good answer is someone opening the deliverable and the source file side by side without preamble.

A bad answer sounds like “our QA process catches that,” “we can pull that and send it after the call,” or “every summary is reviewed before delivery.” All three are answers about process. You asked about the document.

Question 2: Whose name is on it?

Why it matters. “Reviewed by experts” is a noun with nobody behind it. When an evaluation gets challenged, the question of who verified the underlying record set either has an answer or it does not, and finding out which at that moment is expensive.

The test. Four things, asked in this order:
  • What credential or license does the reviewer hold
  • Does their name appear on the delivered document
  • Which service levels include that review, and which do not
  • What happens to that step when the queue is backed up
The last one produces the real information. Most vendors have a verification step on the org chart. Fewer have one that survives a Friday afternoon with a Monday evaluation. Ask specifically what happens under deadline pressure, and listen for whether you get a process or a reassurance.

A good answer: volunteers the service-level distinction before you ask for it. If the fast tier is organization only and the summary tier includes clinician verification, a serious vendor puts that in writing unprompted, because the distinction protects them as much as it protects you.

A bad answer uses “AI-reviewed” and “expert-reviewed” interchangeably within the same twenty minutes. A vendor who lets that blur in a sales call will let it blur in delivery.

Question 3: How do you measure accuracy, and against what?

Why it matters. A headline percentage with no denominator is close to meaningless, and worse than meaningless if you repeat it to anyone. If you have told your carrier that your record partner runs at 99.4% and opposing counsel produces three errors in one file, you have handed them a second argument at no charge.

The test. Four follow-ups. Any vendor who has genuinely measured something answers them without stalling:

  • Accurate at what level: characters, fields, documents, or whole summaries scored pass or fail?
  • Measured by whom: internally, or by someone independent?
  • On what sample: a curated set, or live production work?
  • How often is it measured, and can you see the results?
Per-field measurement is the version that helps you most, because dates, provider names, medications, and work restrictions fail in different ways and at different rates. A blended average can conceal exactly the fields that carry the most weight. An overall accuracy figure of 98% can still mean that one in fifteen work restrictions was recorded incorrectly. That is the kind of error that matters most when a work restriction is the field someone may ultimately rely on or read aloud in a room.

A good answer is a method, with the denominator named, and an offer to run the same audit against your files rather than theirs.

A bad answer leads with a number and cannot produce what sits behind it. Treat that number as copy. It will not help you in a deposition, and it can be used against you in one.

Question 4: What happens when something is wrong?

Why it matters. At volume, errors are a given. Vendors are separated by the correction path, not by whether errors occur. A vendor claiming otherwise is describing a sample size rather than a capability.

The test. Ask what the correction mechanics look like:
  • Who do you contact, and what is the typical correction turnaround?
  • Is a corrected deliverable versioned and dated, or does it quietly replace the original?
  • If they discover an error internally after delivery, do they tell you, or only respond when you catch it?
  • Is the correction recorded anywhere you can review afterward?
Versioning matters more than it sounds. If your QME relied on version one and version two overwrote it without a trace, you have lost the ability to reconstruct what the evaluator actually saw. Occasionally that reconstruction is the entire argument.

A good answer includes a turnaround commitment and a versioning practice described without hesitation, because the vendor has needed it before.

A bad answer is “that hasn’t really come up.”

The California wrinkle: provenance, not just accuracy

In California workers' compensation, what reaches a QME is governed as tightly as what the QME writes. Labor Code section 4062.3 sets the rules on what information may be provided to an evaluator, how it must be served on the other party, and the objection process that follows. Records get objected to. Records get excluded. Sometimes that happens after they have already gone out.

This has a direct consequence for how you evaluate a record partner. If a summary blends the full record set into one narrative with no document-level provenance, and a document in that set is later struck, you cannot demonstrate what the summary would have looked like without it. The evaluator's reliance becomes an open question, and your record review product becomes the thing being litigated.

A defensible deliverable keeps each source document identifiable inside the output. That is a structural requirement rather than a formatting preference, and it is worth raising specifically with any vendor whose experience sits mostly outside California.

What the finished artifact has to contain?

Independent of which vendor you choose, hold the output to this. Each item is here because its absence is a specific way the record set fails later.
Requirement What it prevents
A page citation on every extracted value, using the pagination of the record set as served A value nobody can verify without calling the vendor
A scope statement: pages received, date range, what was excluded and on what basis An argument about whether something was missed or deliberately left out
Visible duplicate handling, with duplicates marked instead of silently removed Three copies of a report with different signature pages quietly becoming one
Gaps stated explicitly, with illegible pages and missing intervals flagged A treatment gap caused by a scanning failure being read as a treatment gap
A provider and date index A reader who cannot reorient inside the original records quickly
Separation between record content and interpretation, with conclusions attributed to whoever made them An unattributed opinion travelling inside a document that looks factual
The reviewer’s name and the service level printed on the document “Who checked this” having no answer at the moment it is asked
A version and date stamp Two copies of the same file being mistaken for each other

How WHITE AI answers the four?

We built the product around these questions, so our answers belong on the record.

Traceability. Output values cite back to their source page in the record set you provided. Anyone holding the deliverable can verify a value without contacting us, which is the entire point of the exercise.

Accountability. On Sort & Summarize, AI handles sorting, indexing, and the first-pass summary. That is document indexing rather than medical coding: WHITE AI does not assign ICD or CPT codes. A licensed medical reviewer then verifies the output against the source before it is delivered, so you are never receiving raw, unchecked AI output. Sort only is an organization service level and does not include that verification step. We state the difference plainly rather than letting the two blur together in a sales conversation.

Measurement. We do not publish a headline accuracy percentage. Accuracy is measured per field and is auditable, which is the claim we can stand behind when someone asks how it was measured and on what sample.

Boundaries. WHITE AI organizes and summarizes records. Clinical, coverage, and causation determinations stay with the physicians and adjusters qualified to make them.

Then stop taking our word for it. Send us a file you already have a deliverable for, from us or from anyone else. We will return ours, and you run the five-value test on both.
General information for medical-legal and workers’ compensation professionals, not legal advice. Requirements differ by jurisdiction and by evaluation type. Verify statutory, regulatory and fee schedule requirements against current authority before relying on them in an active matter.
Scroll to Top
soc2-logo
ISO 27001
HIPAA Compliant