T.10Knowledge Base

Specialty-stratified accuracy benchmarks — F1 by specialty and document type

Specialty readiness mapped against documentation structure, code set depth, and payer rule density.

SpecialtyBenchmarks

Specialty is a proxy, not a cause

Accuracy varies by specialty, but specialty is not itself the causal variable. What varies with specialty is a bundle of underlying properties: how structured the documentation is, how deep and how granular the applicable code set is, how dense the payer rule surface is, and how much of the coding decision depends on information outside the note being coded.

This distinction is practical rather than pedantic. If specialty were causal, the only remedy would be specialty-specific models. Because it is a proxy, the remedy is frequently upstream: a template change that moves laterality into a structured field, or a data feed that makes prior encounters visible, can move a specialty's figure more than a model change would.

It also means the ordering transfers between organisations while the values do not. Two health systems will see the same specialties at the top and bottom of their distributions, with different absolute numbers, because their documentation practices differ.

Driver one: documentation structure

Specialties that produce templated documents against a fixed reporting structure — radiology, pathology, dermatology procedure notes, ophthalmology — give a system a reliable place to find the diagnostic conclusion and a bounded set of ways it can be phrased. Section headers are stable, the impression is explicitly labelled, and negation follows recurring forms.

Specialties working in free narrative — internal medicine, psychiatry, emergency medicine — spread the coding-relevant content through the note, mix considered-and-excluded conditions with treated ones, and rely on the reader to infer which conditions were actually addressed at this encounter. Assertion status becomes the dominant error source rather than code selection.

Structure is also the most modifiable of the three drivers. Requiring laterality as a discrete field, separating the assessment from the plan, and suppressing copy-forward into the assessment are template decisions that raise measured accuracy without touching the model.

Driver two: code set depth and granularity

Orthopaedics illustrates depth. Fracture coding requires bone, subsite, laterality, displacement, open or closed, and a seventh character encoding encounter type and healing status. The combinatorics are large, the distinguishing attributes are frequently absent from the dictation, and the difference between two adjacent leaf nodes is a single character with real reimbursement consequence.

Oncology adds a second axis. Coding requires primary site, morphology, behaviour, laterality, and staging context, and must distinguish active malignancy from personal history of malignancy — a distinction that turns on treatment status rather than on any phrase in the note. Cardiology adds device and procedural specificity layered onto diagnosis coding.

Where depth is high, the failure mode is characteristic: the category is right and the leaf is wrong. That fails as a claim exactly as completely as an unrelated code would, which is why per-character partial credit is not a meaningful metric in these specialties.

Driver three: payer rule density

Some specialties carry a rule surface disproportionate to their code complexity. Behavioural health has time-based codes, session frequency limits, and plan-specific authorisation requirements. Physical and occupational therapy have unit-based billing with the eight-minute rule, plan-of-care documentation requirements, and therapy caps. Sleep medicine, infusion therapy, and durable medical equipment all sit behind heavy coverage determination and prior authorisation logic.

In these specialties, code accuracy and first-pass acceptance decouple sharply. A system can emit exactly the right codes and still see a high denial rate because the constraint lives in the payer layer, not in the code set. Evaluating such specialties on coding accuracy alone systematically overstates readiness.

The corollary is that improvement in these specialties comes from rule-layer investment and intake workflow, not from model work. Prior authorisation failures in particular originate before the encounter is documented at all.

Reading the readiness tiers

Production-ready specialties score low on all three drivers: structured documentation, shallow code sets, light rule surfaces. Radiology, pathology, laboratory, screening and preventive encounters, and routine dermatology procedures fall here. They support high direct-to-bill rates with modest review, and they are the correct first deployment.

Strong-with-tuning specialties score high on one driver only. Orthopaedics is deep but well structured; general surgery is rule-dense but narratively conventional; primary care is narratively loose but shallow in code set. Each responds to a targeted intervention — an ontology traversal layer, a rule pass, a template change — rather than to a general capability increase.

Advanced specialties score high on two or three. Oncology, cardiology, inpatient internal medicine, and complex multi-procedure surgery belong here. They warrant dedicated models, a heavier review policy from day one, and an explicit decision about whether automation targets full coding or coder assistance.

Using stratified benchmarks in procurement and rollout

In procurement, refuse aggregate figures. Require F1 by specialty and document type, each with a sample size and a stated reference standard, and require that the vendor's specialty mix be disclosed alongside so you can reweight to your own. A vendor whose evaluation set is 60% radiology will quote a headline number that says nothing about your cardiology volume.

In rollout, sequence by tier and set confidence thresholds per specialty rather than globally. A global threshold either wastes automation on the easy specialties or under-reviews the hard ones. Route by expected difficulty, and let the low-tier specialties fund the review capacity the high-tier ones consume.

Then re-measure locally and continuously. Specialty figures drift as templates change, as clinicians rotate, and as payer policies update. A stratified dashboard reviewed monthly catches those drifts as localised movement in one stratum, which an aggregate accuracy number would absorb and hide for a quarter.

← Back to the knowledge base