- Legal
What “Defensible” Should Mean When You’re Vetting a QME/IME Record Review Partner
Table of Contents
By then the vendor is not in the conversation. The evaluation is. And the defect is being paid for at its maximum price: in re-evaluations, in a supplemental report, in a case that slides two quarters to the right while the reserve sits open.
A record set is defensible when someone motivated to discredit it cannot.
Accurate and defensible are different things
First, retire "AI-powered" as a filter
Question 1: Can you trace every value back to a page?
Why it matters. Traceability is the property that survives cross-examination. “Where did this date come from?” should be answerable in seconds by anyone holding the document, including the person trying to break it. A summary that only the vendor can verify is a summary you cannot defend without the vendor, and the vendor is not available at 3pm on a Thursday in a deposition.
The test. Ask for a real redacted production deliverable before you sign anything. Not the demo file. Pick five entries at random: a date of injury, a medication, a work restriction, a diagnosis as stated by a treater, a treatment gap. Ask them to show you the source page for each one, live, while you are on the call.
A good answer is someone opening the deliverable and the source file side by side without preamble.
A bad answer sounds like “our QA process catches that,” “we can pull that and send it after the call,” or “every summary is reviewed before delivery.” All three are answers about process. You asked about the document.
Question 2: Whose name is on it?
Why it matters. “Reviewed by experts” is a noun with nobody behind it. When an evaluation gets challenged, the question of who verified the underlying record set either has an answer or it does not, and finding out which at that moment is expensive.
- What credential or license does the reviewer hold
- Does their name appear on the delivered document
- Which service levels include that review, and which do not
- What happens to that step when the queue is backed up
A good answer: volunteers the service-level distinction before you ask for it. If the fast tier is organization only and the summary tier includes clinician verification, a serious vendor puts that in writing unprompted, because the distinction protects them as much as it protects you.
A bad answer uses “AI-reviewed” and “expert-reviewed” interchangeably within the same twenty minutes. A vendor who lets that blur in a sales call will let it blur in delivery.
Question 3: How do you measure accuracy, and against what?
Why it matters. A headline percentage with no denominator is close to meaningless, and worse than meaningless if you repeat it to anyone. If you have told your carrier that your record partner runs at 99.4% and opposing counsel produces three errors in one file, you have handed them a second argument at no charge.
The test. Four follow-ups. Any vendor who has genuinely measured something answers them without stalling:
- Accurate at what level: characters, fields, documents, or whole summaries scored pass or fail?
- Measured by whom: internally, or by someone independent?
- On what sample: a curated set, or live production work?
- How often is it measured, and can you see the results?
A good answer is a method, with the denominator named, and an offer to run the same audit against your files rather than theirs.
A bad answer leads with a number and cannot produce what sits behind it. Treat that number as copy. It will not help you in a deposition, and it can be used against you in one.
Question 4: What happens when something is wrong?
Why it matters. At volume, errors are a given. Vendors are separated by the correction path, not by whether errors occur. A vendor claiming otherwise is describing a sample size rather than a capability.
- Who do you contact, and what is the typical correction turnaround?
- Is a corrected deliverable versioned and dated, or does it quietly replace the original?
- If they discover an error internally after delivery, do they tell you, or only respond when you catch it?
- Is the correction recorded anywhere you can review afterward?
A good answer includes a turnaround commitment and a versioning practice described without hesitation, because the vendor has needed it before.
A bad answer is “that hasn’t really come up.”
The California wrinkle: provenance, not just accuracy
In California workers' compensation, what reaches a QME is governed as tightly as what the QME writes. Labor Code section 4062.3 sets the rules on what information may be provided to an evaluator, how it must be served on the other party, and the objection process that follows. Records get objected to. Records get excluded. Sometimes that happens after they have already gone out.
This has a direct consequence for how you evaluate a record partner. If a summary blends the full record set into one narrative with no document-level provenance, and a document in that set is later struck, you cannot demonstrate what the summary would have looked like without it. The evaluator's reliance becomes an open question, and your record review product becomes the thing being litigated.
A defensible deliverable keeps each source document identifiable inside the output. That is a structural requirement rather than a formatting preference, and it is worth raising specifically with any vendor whose experience sits mostly outside California.
What the finished artifact has to contain?
| Requirement | What it prevents |
|---|---|
| A page citation on every extracted value, using the pagination of the record set as served | A value nobody can verify without calling the vendor |
| A scope statement: pages received, date range, what was excluded and on what basis | An argument about whether something was missed or deliberately left out |
| Visible duplicate handling, with duplicates marked instead of silently removed | Three copies of a report with different signature pages quietly becoming one |
| Gaps stated explicitly, with illegible pages and missing intervals flagged | A treatment gap caused by a scanning failure being read as a treatment gap |
| A provider and date index | A reader who cannot reorient inside the original records quickly |
| Separation between record content and interpretation, with conclusions attributed to whoever made them | An unattributed opinion travelling inside a document that looks factual |
| The reviewer’s name and the service level printed on the document | “Who checked this” having no answer at the moment it is asked |
| A version and date stamp | Two copies of the same file being mistaken for each other |
How WHITE AI answers the four?
Traceability. Output values cite back to their source page in the record set you provided. Anyone holding the deliverable can verify a value without contacting us, which is the entire point of the exercise.
Accountability. On Sort & Summarize, AI handles sorting, indexing, and the first-pass summary. That is document indexing rather than medical coding: WHITE AI does not assign ICD or CPT codes. A licensed medical reviewer then verifies the output against the source before it is delivered, so you are never receiving raw, unchecked AI output. Sort only is an organization service level and does not include that verification step. We state the difference plainly rather than letting the two blur together in a sales conversation.
Measurement. We do not publish a headline accuracy percentage. Accuracy is measured per field and is auditable, which is the claim we can stand behind when someone asks how it was measured and on what sample.
Boundaries. WHITE AI organizes and summarizes records. Clinical, coverage, and causation determinations stay with the physicians and adjusters qualified to make them.