Picking a diagnosis code off a fax looks like pure data entry. Someone reads a referral, types a code, moves to the next one. Nothing about it feels consequential.
But code it one digit wrong, K50.90 for Crohn's instead of K51.90 for ulcerative colitis, and you've quietly switched the patient to a different disease, one with its own policy criteria and its own preferred agent pathway. The keystroke that did it looked exactly like the thousand correct ones around it.
I spent years on the payer side, close to how coverage decisions get made and where they go wrong. The pattern I kept seeing: most healthcare leaders look for risk in the wrong place. They watch the tasks that obviously involve clinical judgment and trust the ones that look like paperwork. But the errors that actually hurt patients tend to hide in the clerical looking work.
Why "administrative" doesn't mean low risk
The administrative versus clinical split is useful, and it's worth being clear about what it's good for. It tells you who should be doing what. Clinicians should spend their time on clinical judgment, not chasing faxes and retyping lab values, and pulling them back to the top of their license is the right goal.
The trouble starts when "administrative" gets quietly read as "low stakes." In practice, that's not always true, and it depends on the kind of operation. It's not that the split is wrong, it's that it answers a different question than people think. It tells you who should staff the work, not how much damage the work can do if it's wrong. Those two questions get treated as one, and that's where the risk hides.
Everything before the patient is in the chair gets filed as paperwork: intake, benefit verification, prior authorization, billing. But the data flowing through that paperwork is almost entirely clinical. Which means the risk is too.
Take the fields that feel most like routine entry. A patient's height and weight are the most boring boxes on the form, right up until you remember that for a chemotherapy regimen, those two numbers are the dose. Fat finger the weight and you've either underdosed the cancer or overdosed the patient. Nothing about the task looked clinical while you were doing it, and downstream, the error doesn't announce itself either. It's baked into a regimen that gets ordered, dispensed, and administered, and it surfaces only if the patient's response is off or something goes wrong, long after the keystroke that caused it.
The same trap hides in softer fields. Noting that a patient stopped a drug is administrative. Knowing why is clinical. "Quit because it didn't work" and "quit because it nearly killed them" get typed into the same box, and they point to opposite coverage decisions. Most specialty drug policies require step therapy: try and fail a preferred agent before the payer covers the one you actually want. Lack of efficacy means trying the next drug in the pathway. A severe adverse event usually means waiving step therapy entirely, since you can't ethically ask someone to repeat a drug class that already hurt them. Same fact on the surface, discontinued, opposite consequences underneath. The form doesn't care which one you meant. The payer does.
One mistake, all the way down
What ties these examples together is a single move: dense clinical meaning compressed into a field too small to hold it. A drug failure with a whole story behind it becomes a checkbox. A dose becomes two integers. A diagnosis becomes five characters. The compression loses information, and the loss is invisible at the moment it happens, which is why nobody slows down.
Follow any of those losses downstream and you see the same shape. The wrong diagnosis code routes the case down the wrong policy pathway, producing a denial, or worse, an approval on grounds that were never true. The claim gets paid, then reversed months later when an auditor works back through the file, and the patient's continuity of care becomes a question mark. The mistyped weight rides quietly inside a dose calculation until a bad response forces someone to trace it back. The flattened discontinuation reason locks a patient into a step therapy sequence they should have skipped, and the denial that follows looks procedurally correct even though the premise was wrong. Three fields, three failures, all traceable to a keystroke nobody thought was worth a second look.
That's the shape of the risk. Not a dramatic failure of judgment, but a quiet one of transcription, sitting undetected inside work that looked finished the moment it was typed.
The question worth asking
The useful question was never "is this administrative or clinical." That tells you who to staff, not how much it can hurt you. The better move is naming which specific fields carry that downstream weight and routing only those for clinical level verification, still consistent with keeping clinicians at the top of their license.
The sharper question: what decision is being made here, what happens if it's wrong, and can a qualified person verify the basis for it? Ask it field by field, and the boxes needing the most scrutiny turn out to be the ones that looked like data entry all along: height, weight, the reason a drug was stopped, the single digit in a diagnosis code.
None of these fields look dangerous. That's what makes them dangerous. Work that announces its stakes gets careful attention by default. Work that hides them behind a clean, clerical surface is where the costly, silent errors get made.
The fix isn't more vigilance training. That bet is already failing, since vigilance doesn't scale and these fields don't look like they need it. The fix has to match the shape of the field, and that's also where AI earns its place, not as a blanket layer of automation, but as a different check for each kind of error.
Diagnosis codes are finite and enumerable, so this is a pattern matching problem. A model doesn't need to understand Crohn's versus ulcerative colitis to help. It needs to know K50.90 and K51.90 are one edit apart, that they route to different pathways, and whether the referral note supports or contradicts the code typed. A mismatch is worth a human look before the case routes anywhere.
Height and weight are numeric and checkable against context, so the useful move is cross-referencing, not judgment. A weight that's genuinely changed since the last visit isn't an error, but a weight that disagrees with a second mention of it elsewhere in the same fax or note is a strong signal something got mistyped, since both numbers describe the same patient at the same moment and have no legitimate reason to differ. That kind of cross-check within a single document is exactly what a person retyping one field at a time can't do, and exactly what's cheap to catch before the number becomes a dose.
The discontinuation reason is different. There's no wrong keystroke to catch, the loss happens because the real reason gets squeezed into a checkbox that's too small for it. Here AI's job isn't validation, it's preservation: read the note, sort the reason into the categories the payer actually uses, and keep the original sentence attached so the classification stays traceable to source.
None of this asks AI to make the clinical judgment call itself. Its job is to hold more context at once than a person moving field by field can, and to flag disagreement rather than resolve it silently. You don't need to audit everything. You need to know which boxes are load bearing, and you need a check that matches the shape of each one: a rule for what's enumerable, cross-referencing for what's numeric, preserved source language for what compresses a clinical story into a box too small to hold it.
The three moves
Framed this way, the operators who get this right make three moves. They automate aggressively where the answer is defined. They prepare intelligently where the material is ambiguous, giving the expert a faster path to the decision. And they never let either one touch the judgment of whether a treatment is necessary and safe for the patient in front of them.
That's the philosophy behind how Mandolin builds AI for healthcare operations: aggressive automation on the deterministic work, a clinician kept firmly on every decision that affects whether a treatment is right and safe. Done that way, the human in the loop isn't a constraint on automation. It's what makes trusting automation possible at all.




