T.06Knowledge Base
Measuring coding model accuracy correctly — what vendor figures typically measure vs what organisations need
Aggregate F1 conceals variance by document type and specialty — the variance most material to procurement.
One number, many populations
A vendor accuracy figure is a weighted average over a document mix. The weights are the vendor's evaluation set composition, not yours, and they can move the headline number by several points without any change to the underlying model. A set weighted toward straightforward office visits will report a materially higher figure than the same system evaluated on a set weighted toward inpatient surgical encounters.
This makes cross-vendor comparison of headline figures close to meaningless unless the evaluation composition is disclosed. Two systems quoting the same number may have been measured on populations that differ more than the systems do, and the better-looking figure frequently belongs to whoever chose the easier mix.
The first question in any evaluation is therefore not how accurate the system is, but on what population the figure was computed and how that population compares to the one the system will actually see.
What unit of accuracy is being counted
Accuracy in coding can be counted at several units, and the choice changes the number substantially. Per-code accuracy asks whether each emitted code is correct. Per-encounter accuracy asks whether the full code set for an encounter is correct, which is a stricter test because one wrong code fails the encounter. Character-level partial credit asks whether the code is close, which is not a real-world outcome — a code that matches to four characters when seven are required is denied, not partially paid.
Encounter-level exact match is the unit that corresponds to a clean claim, and it is the unit most vendor figures avoid because it is unflattering. A system with 96% per-code accuracy on encounters averaging six codes has a considerably lower clean-encounter rate than the headline suggests.
Sequencing and modifiers must be inside the unit of measurement as well. A code set that is complete but incorrectly sequenced, or missing a required modifier, is not a correct output even though every individual code appears in the reference set.
Precision, recall, and the asymmetry F1 hides
F1 collapses precision and recall into one number, which is convenient and, in this domain, misleading. The two error types have different consequences. A missed code is lost revenue and, where it affects risk adjustment, understated acuity. An extra code that the documentation does not support is an overcoding exposure, and where it recurs systematically it is the pattern that draws audit attention and repayment demands.
Because the consequences are asymmetric, the acceptable operating point is asymmetric too, and it differs by organisation and by risk posture. Reporting only F1 removes the ability to see where a system sits on that trade-off, and two systems with identical F1 can have very different precision-recall splits.
Report both figures, and report them for the error classes that matter separately: false positives arising from negated or hedged findings, false positives arising from copy-forward historical content, and false negatives on codes with revenue or risk-adjustment impact.
Autonomous accuracy versus accuracy with review
Almost every production deployment routes low-confidence output to human review, which means the number that matters operationally is not the model's standalone accuracy but the accuracy of the model-plus-review system at a stated review rate. Those are different figures and both are legitimate — but they are frequently conflated.
The honest presentation is a curve rather than a point: accuracy as a function of the proportion of encounters sent to human review. A system that reaches 97% while routing 40% of encounters to a coder has a different economic profile than one that reaches 95% while routing 8%, and neither figure is interpretable without the review rate attached.
Ask for direct-to-bill rate alongside accuracy — the proportion of encounters that clear without human touch — and ask what confidence policy produced it. A vendor quoting high accuracy without a review rate is quoting the wrong number.
Stratify before comparing
The variance that matters is inside the aggregate. Stratify by document type — office visit, emergency encounter, operative report, inpatient discharge summary, radiology report — and by specialty, and report each stratum separately with its sample size. This is where systems differentiate, and it is where the aggregate figure actively conceals information.
Sample sizes per stratum need to be large enough to be meaningful. A specialty represented by thirty encounters produces a confidence interval wide enough to accommodate almost any claim, and stratum-level figures reported without sample sizes should be treated as unmeasured.
Establish the reference standard explicitly as well. Coder-adjudicated ground truth with a documented inter-rater agreement figure is a defensible reference; historical billed claims are not, since billed data contains the very error patterns you are trying to detect and will reward a system for reproducing them.
The OIG threshold and what to require in procurement
The Office of Inspector General treats 95% accuracy as the working compliance threshold for hospital coding programs subject to external audit, which sets the floor for any automated system operating without full human review. That threshold is a floor, not a target, and it applies to the population being coded rather than to a favourable subset of it.
In procurement, require: the evaluation population and its composition; the unit of accuracy; precision and recall reported separately; the accuracy-versus-review-rate curve with direct-to-bill rate; stratification by document type and specialty with sample sizes; the reference standard and its inter-rater agreement; and a pilot on your own documents with your own coders adjudicating.
The pilot is the part that cannot be substituted. Every figure above is a claim about a population; only a pilot on your own encounter mix, adjudicated by your own coding staff, measures the system on the population you will actually run it against.