Can You Trust What Your AI Scribe Wrote?

Current as of August 20, 2026. Next check-in: February 2027 — to see whether the one clinician-edit-rate study behind this post has cleared peer review, and whether an SNF-specific accuracy study has been published yet.
The honest answer is: it depends what you’re measuring, and no one has measured it in a skilled nursing facility yet. Independent, peer-reviewed studies of AI scribes in outpatient care put the error rate anywhere from roughly 1 in 4 notes (missing or wrong key clinical elements) to 7 in 10 notes (containing at least one error of any kind) — two real numbers, not one blended range, and neither was measured in a SNF. A peer-reviewed comparison of one commercially available scribe also found it hallucinated at a statistically higher rate than a human-authored note. None of that means the tool can’t be trusted — real-world deployment data shows clinicians are still editing the large majority of AI drafts before they sign.
It does mean your review step, not the AI’s output, is what actually protects you.
So How Many AI-Drafted Notes Actually Contain an Error?
An AI scribe — sometimes called an ambient AI scribe — listens to a clinical encounter and drafts the note for a clinician to review and sign, instead of the clinician typing or dictating it. The pitch is obvious: less time on the keyboard, more time with the resident. The question every facility asks next is whether the draft it hands back is actually right.
Two of the most rigorous studies on this gave two different, both-correct answers — and conflating them is not quite right. Biro et al., writing in the Journal of Medical Internet Research (2025), ran two commercial ambient scribe products against 11 real outpatient encounters and found that 31 of 44 draft notes — 70% — contained at least one error, averaging 2.9 errors per note. Most of those were omissions: something the clinician said that never made it into the note. Separately, Anderson et al., in Mayo Clinic Proceedings: Digital Health (October 2025), tested five ambient scribe platforms against 14 simulated encounters and found a mean of 26.3% of key clinical elements were omitted or captured incorrectly — a different unit of measurement (percentage of expected clinical content missed or wrong) than Biro’s “did this note have any error at all.”
Note that these studies were conducted in 2025 which even if it sounds recent is old given the pace of AI advancements and LLM capabilities.
What both studies agree on: vendor-to-vendor variance is at least as important as any single error rate. Anderson et al. found one platform performed significantly worse than the other four (p<.0053), and that only 35.8% of clinical elements were captured consistently across all five platforms tested — meaning even picking a “good” vendor doesn’t fully close the gap.
Has Anyone Tested This in a Skilled Nursing Facility?
No — and that’s worth saying plainly rather than glossing over. Every study cited in this post, including the ones below, was conducted in ambulatory or outpatient primary/specialty care: academic medical centers, simulated office visits, outpatient EHR pilots. None involved a nursing facility, a long-term care setting, or the note types — skilled nursing progress notes, care plans, MDS assessments — that actually drive SNF reimbursement and survey risk. That’s a genuine gap in the research base.
The best available evidence — all of it from outpatient care — puts ambient AI scribe error rates somewhere in the range above. Treat that as the closest available proxy for a SNF, not a SNF-validated number. If your facility has taken a documentation-driven citation in a recent survey cycle — an incomplete care plan, a missing signature, a deficiency where the note itself was the problem rather than a staffing number — you already know how unforgiving a surveyor is about what’s actually in the chart.
Is a Named, Commercially Available Scribe Actually More Error-Prone Than a Human?
In at least one peer-reviewed, head-to-head comparison, yes. Palm et al., in Frontiers in Artificial Intelligence (October 2025), evaluated Suki AI’s ambient documentation system against 97 clinical encounters (194 paired notes) and found a 31% hallucination rate in the AI-drafted notes versus 20% in gold-standard human-authored notes — a statistically significant gap (p=0.01). On overall documentation quality, the human notes scored marginally higher, though the AI notes actually scored better on thoroughness and organization.
Worth being precise about what this does and doesn’t show: it’s one peer-reviewed comparison of one named, currently marketed vendor against a human baseline.
Can The Error Rate Be Engineered Down?
Yes, error rates can be engineered down — which is the genuinely encouraging half of this story. Asgari et al., in npj Digital Medicine (2025), built an error-detection framework (CREOLA) and tested 18 experimental LLM configurations against nearly 13,000 clinician-annotated sentences, landing on a 1.47% hallucination rate and 3.45% omission rate — an order of magnitude below the ambulatory scribe studies above. It’s real evidence that “AI scribes are inherently this error-prone” isn’t the right conclusion — implementation and product quality moves the number a lot.
What About Residents Who Rely on an Interpreter?
This compounds the accuracy question rather than sitting apart from it. Rabotin et al., in JMIR Medical Informatics (July 2026), scripted five English/Spanish clinical encounters with 20 deliberate interpreter errors built in, then ran them through two ambient AI scribes. One propagated 55% of those interpreter errors into the resulting note; the other propagated 60%. The asymmetry is the more useful finding: errors that started in what the patient said were propagated 80–100% of the time, versus only 20–30% for errors that started in what the clinician said. In other words, the scribe was far more likely to faithfully preserve a clinician’s words than a patient’s, once an interpreter had already introduced a mistake.
In one documented example, an interpreter’s omission of a patient reporting hearing trouble led both AI scribes to chart “no other hearing problems” — an unsupported negative finding invented on top of the interpreter’s error, and corrected by neither tool.Rabotin et al., JMIR Medical Informatics (2026)
For a California SNF with residents who don’t primarily speak English, “can you trust what your AI scribe wrote” isn’t only a monolingual accuracy question. If your facility already handles interpreter-mediated care — see our related piece on interpreter access at your California SNF — this is the specific failure mode to build a review step around: don’t assume the AI caught what the interpreter missed.
Does a Human Actually Catch These Errors Before They Hit the Chart?
At real deployment scale, yes — more than you might expect. Guo et al. (UC Irvine Health), posted as a medRxiv preprint in January 2026, analyzed 23,760 real-world notes containing AI-drafted sections across two vendor deployments and 225 clinicians. Only 15.6% of notes were signed with zero edits — 84.4% were edited by the clinician before signing, with editing concentrated most heavily in the Assessment & Plan section, the highest-stakes part of the note. Individual clinician habits drove far more of that variation than specialty norms did.

Who’s Actually Accountable if an AI-Drafted Note Turns Out Wrong?
That question is already answered, and this post isn’t re-litigating it. Our companion post, Who’s Legally Accountable for an AI-Drafted Nursing Note?, covers it in full: CMS’s single-signer model treats an AI scribe exactly like a human scribe — the treating clinician’s signature authenticates the entry, not a second co-signer — and California law names the person who makes the entry as the record’s “author of record,” not the tool that drafted it. What this post adds on top of that settled answer is the practical layer: given real error rates in the range documented above, what does your review step actually need to look like to make that signature mean something, rather than a formality you’re exposed on later?
Building a Note-Review Workflow That Actually Catches Errors
None of the studies above argue for abandoning AI scribes — several argue the opposite, that careful engineering and an active review habit close most of the gap. Here’s where to focus a review workflow given what the evidence actually shows:
- Weight your review time toward the Assessment & Plan section first — it’s both the highest-stakes part of the note and where other clinicians already concentrate their edits.
- Treat omissions, not just wrong facts, as the primary error type to hunt for — cited studies found omission was the dominant error category across both products tested, not fabricated content.
- If your facility serves residents who need an interpreter, add an explicit check for patient-reported symptoms and history — patient-speech errors propagate far more often than clinician-speech errors, which makes this the specific spot an AI scribe is least likely to self-correct.
- Keep your audit trail intact regardless of which tool drafts the note — an individualized signer, a system timestamp, and an immutable entry are what a surveyor asks for either way, AI or not.
Frequently Asked Questions
What percentage of AI-drafted clinical notes actually contain errors?
It depends what’s being measured, and there’s no single agreed-on number. Peer-reviewed studies in outpatient care found anywhere from 26.3% (key clinical elements omitted or wrong, Anderson et al.) to 70% (notes containing at least one error of any kind, Biro et al.). No study has measured this specifically in a skilled nursing facility.
Is a specific AI scribe more likely to hallucinate than a human writing the same note?
In at least one peer-reviewed, head-to-head study, yes — Palm et al. found Suki AI’s ambient scribe hallucinated in 31% of notes versus 20% for human-authored notes, a statistically significant difference. That’s one study of one named vendor, not a claim that applies to every AI scribe on the market.
Has anyone studied AI scribe accuracy specifically in a nursing home?
Not yet, as of this writing. Every peer-reviewed study reviewed for this post was conducted in ambulatory or outpatient care. Treat the error rates above as the closest available proxy for a SNF setting, not a SNF-validated figure.
Do clinicians actually catch AI-drafted errors before signing, or is it a rubber stamp?
Real-world data says it’s not a rubber stamp: one health system’s deployment data found 84.4% of AI-drafted notes were edited by the clinician before signing. That measures editing activity, not whether every important error was caught — but it’s evidence against the assumption that sign-off is automatic.
Who’s legally responsible if an AI-drafted note turns out to be wrong?
The treating clinician who signs it — not the AI vendor. See our companion post for the full legal breakdown.
Who’s Legally Accountable for an AI-Drafted Nursing Note?Can an AI scribe make an interpreter’s mistake worse?
A small pilot study found ambient AI scribes propagated more than half of scripted interpreter errors into the resulting note, and were especially likely (80–100% of the time) to propagate errors that originated in what the patient said rather than what the clinician said.
Disclaimer: This post is informational, not medical or legal advice. The studies cited here were conducted in outpatient settings, not skilled nursing facilities — confirm any change to your facility’s documentation review workflow with your own clinical and compliance leadership before acting on it.
Sources
- Biro et al., “Accuracy and Safety of AI-Enabled Scribe Technology: Instrument Validation Study,” Journal of Medical Internet Research 2025;27:e64993 (published Jan 27, 2025).
- Anderson, Mohan, Dorr, Ratwani, Biro & Gold, “Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters,” Mayo Clinic Proceedings: Digital Health (Oct 9, 2025).
- Palm, Manikantan, Mahal, Belwadi & Pepin, “Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe,” Frontiers in Artificial Intelligence (Oct 22, 2025).
- Asgari et al., “A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation,” npj Digital Medicine (2025).
- Rabotin, Aguilar, Sandoval Gonzalez, Iniguez, Jung, Lee, Bell & Arroyo, “Propagation of Interpreter Errors by Ambient AI Scribes: Study Using Simulated Clinical Encounters,” JMIR Medical Informatics 2026;14:e88734 (published Jul 28, 2026).
- Guo, Hu, Zhou, Lyu, Sutari, Tam, Chow, Perret & Pandita, “From Conversation to Chart: An Analysis of Clinician Edits to Ambient AI Draft Notes,” medRxiv preprint (posted Jan 13, 2026 — unreviewed preprint, not yet peer-reviewed).
- Relic Care, Who’s Legally Accountable for an AI-Drafted Nursing Note? (companion post; CMS MLN905364 and Cal. Code Regs. tit. 22 §72543(f) analysis carried over from that post’s own research, not re-verified here).
- Relic Care, Interpreter Access at Your California SNF (companion post).
Where Relic Care Fits In
If you’re evaluating an AI scribe or already piloting one, the tool handling your documentation should make it easy to see exactly what the clinician changed before signing off — not just that they signed. See how Notes Scribing handles that review trail for long-term care documentation.
If your residents need an interpreter as part of care, see how Interpreter Assistant keeps that conversation itself accurate and hands-free — before it ever reaches a chart.
And if tracking documentation, audit-trail, and review obligations across everything your facility runs feels like a compliance spreadsheet nobody has time to maintain, that’s what Compliance is built for.
More for the People Running Your Facility

Who's Legally Accountable for an AI-Drafted Nursing Note?
CMS's single-signer model and California's author-of-record rule put accountability on the signing clinician — not your AI scribe vendor.

AB 843 Is Dead: What Actually Governs Interpreter Access at Your California SNF
AB 843 never bound SNFs directly, and it's dead. See what Section 1557 actually requires for interpreter access at your California facility.


