Why human review still matters in AI-assisted document processing
The better an AI tool gets at reading clinical documents, the more a practice needs a person looking at its output, not less. That sounds backwards. It is the most important thing to understand before switching one on.
The usual case for human review is that AI makes mistakes. That is true, and it is the weaker argument. A tool that was wrong a third of the time would never be trusted, so every output would be checked. The risk arrives when the tool is right almost all of the time, because that is when people stop looking.
Researchers call this automation bias, and Australia’s coordinated AI guidance for general practice names it specifically. It is not a character flaw. It is what any sensible person does with a colleague who has been right nine hundred times in a row. The trouble is that the errors a good document tool still makes are not the obvious ones. They are the ones that look exactly like a correct answer.
What survives a good tool
Five errors that look right
Each of these produces a tidy, well-formatted, plausible result. None of them is caught by a glance at the form. All of them are caught in seconds by a person who can see the document next to it.
- 01
The right name on the wrong patient
A father and son with the same name, two patients born a year apart, a married name the sender has not caught up with. The extraction is letter-perfect and the match is still wrong. Nothing about the output looks unusual.
- 02
Two patients in one PDF
A pathology provider batches a day’s reports into one file, or a hospital discharge summary attaches a family member’s consent form. A tool that reads the first page confidently files the whole thing under one name.
- 03
A digit lost to the scanner
A faxed report, a photographed letter, a third-generation photocopy. A 3 becomes an 8 in a date of birth or a Medicare number. The field is populated, formatted correctly, and wrong.
- 04
The letterhead’s date, not the report’s
Documents carry several dates: collected, reported, dictated, printed, received. Pick the wrong one and a result appears older or newer than it is, which changes whether anyone chases it.
- 05
The category that is almost right
A specialist letter that embeds a pathology table, a referral that is really a discharge summary. Filed as correspondence when it should be an investigation, it never appears where the doctor looks for results.
Notice what these have in common. The tool did not hallucinate. In most of them it read the page correctly. The error is in the world around the document, in the practice’s records, in the scanner, in the sender’s habits, and only a person with the context can see it.
The standard
What “reviewed” has to mean
RACGP Standards for general practices, Criterion GP2.2, requires that incoming pathology, imaging, investigation reports and clinical correspondence are reviewed, noted electronically, acted on where needed and added to the patient record, and that the review itself is recorded. The criterion predates AI and does not mention it. It does not need to. It describes an act performed by a person and evidenced in the record, and a surveyor will ask to see that evidence.
That sets the bar for what a review step in an AI-assisted workflow has to be. A confirm button pressed without the source in view is not a review. It is a signature on someone else’s work. The practice has taken responsibility for the output without checking it, which is the worst of both worlds: the speed of automation with none of its auditability, and the liability of manual processing with none of its attention.
Review is the moment the practice takes ownership of what the software prepared. Everything before it is a draft.
Medical defence organisations make the same point from the other direction. Avant’s position on AI in healthcare, and the RACGP’s guidance on AI tools, are consistent that using software does not move clinical or professional responsibility to the vendor. If a wrongly filed result reaches a patient late, the question asked afterwards will not be whether the AI was accurate. It will be who reviewed it, what they could see, and whether that was recorded.
Design, not discipline
Making a review people actually do
Telling staff to “always check the output” does not survive contact with a Monday inbox. Review has to be designed so that doing it properly is faster than not doing it. Four rules do most of the work.
- 01
Evidence beside the answer
Every extracted field should point at the words on the page it came from. Checking a name against a highlighted name takes a second. Checking it against a whole document takes a minute, so it stops happening.
- 02
Uncertainty on the surface
When the tool is unsure, it should say so, in the field, before anyone confirms. A quiet best guess is the most dangerous output a system can produce, because it looks identical to a confident one.
- 03
Review measured in seconds
A review step that takes longer than the manual process it replaced will be skipped on the first busy morning. The tool’s job is to make a proper review faster than a rubber stamp.
- 04
The record as a by-product
Who reviewed, what they changed and when it was filed should be written by the system as the review happens. If someone has to type the audit trail afterwards, it will be the first thing to go.
There is a corollary practices sometimes resist: a tool that is built for review will occasionally be slower on a single document than one that simply files it. It will stop on the two-patient PDF and the ambiguous date of birth and wait. That pause is the product working. The alternative is a system that is faster on every document and silently wrong on the ones that matter.
The thirty-second review
Human review does not mean re-reading the document. It means checking the six things that decide where it goes and what happens next. With the evidence highlighted beside each field, this list takes about half a minute.
- 1Two identifiers, not one: name and date of birth both match the record, checked against the document, not the form.
- 2One patient: the document is about one person from first page to last.
- 3The right document type: correspondence, investigation or administrative, judged by the content, not the sender.
- 4The right date: the clinically relevant date, not the print date.
- 5The right recipient: the requesting or usual practitioner, and a delegate if they are away.
- 6The flag, if any: an urgency marker is seen and passed on, not decided on, by the person filing.
Everything on that list is administrative. Deciding what a result means, and what to do about it, is clinical and belongs to the practitioner it is routed to. For where that line sits across the whole workflow, see What should, and shouldn’t, be automated in a medical practice?
The Australian position
What the guidance actually asks for
Verify before it informs care
The coordinated guidance from the RACGP, the TGA and the Australian Commission on Safety and Quality in Health Care asks practices to verify all AI-generated output before it is relied on, to guard against automation bias, and to keep patient information out of general-purpose chatbots. A review step with the source in view satisfies the first two by construction.
Record the review
GP2.2 wants evidence that each incoming item was reviewed and by whom. The cleanest way to produce it is for the tool to write the trail as a side effect of the review, with no patient content in the log itself.
Document the governance
Practices are expected to be able to say which AI tools they use, what those tools are allowed to do, who checks their output and how. A one-page policy that names the review step and the thirty-second checklist covers most of what a surveyor will ask.
This article is general information about practice workflow, not legal or clinical advice. Check the current RACGP, TGA and OAIC guidance for your circumstances.
Frequently asked questions
If an AI tool is accurate, does a practice still need to review every document?
Yes. Australian guidance from the RACGP, the TGA and the Australian Commission on Safety and Quality in Health Care requires AI-generated content to be verified by a person before it informs care, and RACGP Standards Criterion GP2.2 requires the review of every incoming result and item of correspondence to be recorded. Accuracy changes how long the review takes, not whether it happens.
What is automation bias, and why does it matter at the scanning desk?
Automation bias is the tendency to accept a system’s output without checking it, especially when the system is usually right. In document processing it shows up as clicking “confirm” without looking at the source. It matters because the errors that survive a good tool are precisely the ones that look correct, such as a plausible match to the wrong patient.
Does human review mean re-reading the whole document?
No. It means checking the specific things that determine where the document goes and what happens next: the patient’s identity against two identifiers, the document type, the relevant date, the recipient and any urgency marker. A well-designed tool shows the evidence for each of those next to the field, so a proper review takes about thirty seconds.
Who is responsible if an AI-assisted filing error reaches a patient?
The practice and the treating practitioner. Medical defence organisations and the RACGP are consistent that AI tools do not shift clinical or professional responsibility to the vendor. That is the practical reason the review step has to be real: it is the point at which the practice takes ownership of what the software prepared.
How MEDsort applies this
Built for the thirty-second review.
MEDsort shows every extracted field beside the words on the page it came from, colour-coded by section, so checking a name or a date is a glance rather than a search. An email with several documents is reviewed one document at a time. Uncertain matches are surfaced, not guessed. Nothing reaches a Best Practice record until a member of your team has confirmed it, and that confirmation is recorded without storing any patient information.