Ambient AI can capture the conversation. It still can’t understand the patient.

Learn the difference between “capture” and “comprehension,” and why organizations must build or adopt clinical AI with a governed clinical context layer.
Published
Written by
Picture of John Laursen
SVP of Product for Health Tech

Ambient AI scribes are, by some measures, the fastest-adopted technology healthcare has seen in a generation.

Clinicians who spent years typing through visits are now talking through them instead. Health systems are signing contracts by the dozen, and vendors are racing to meet the demand. The keyboard is quietly disappearing from the exam room.

None of that is hype, but it raises a broader issue. A medical note that faithfully reflects what was said in the room isn’t the same as data the health system can trust.

While ambient AI has solved documentation burden, it has not solved clinical understanding.

The gap between the two will soon become one of the most significant problems in health tech.

Why ambient AI is real, and why that’s a problem

Ambient AI’s breakthrough is genuine. Ambient scribes give clinicians more time with patients, ease after-hours charting burnout, and produce notes faster than dictation ever did. Health systems are seeing real returns, and clinicians are asking for the technology rather than resisting it.

That is exactly why the stakes have changed. Ambient output is no longer just a clinician convenience. It is increasingly the substrate for coding, billing, quality reporting, population health analytics, and the training data feeding the next generation of clinical AI.

When the output of one AI system becomes the input to every downstream system that follows, “good enough to read” stops being good enough. Much of the issue stems from the differences between “capture” and “comprehension.” 

The two words sound similar, but they solve different problems. Capture means transcribing and summarizing what was said in the room. Comprehension means understanding what it clinically means once the recorder stops running. 

Consider what gets lost in a capture-only workflow. “Shortness of breath” is not the same clinical fact as acute hypoxic respiratory failure, but a transcript can flatten the two into the same phrase. A suspected condition and a confirmed one can sound identical yet carry different treatment and risk implications.

Chronicity and disease progression rarely get narrated aloud in a 15-minute visit. And the laterality, staging, and specificity that coders and payers require often live in clinical judgment, not the transcript. 

Probabilistic AI is excellent at producing fluent, plausible-sounding language. The problem is that clinical decisions, reimbursement, and analytics require something different: deterministic specificity. Those are different jobs, and conflating them is where the risk hides.  

Generative AI doesn’t eliminate data quality problems. It just makes them easier to hide. 

Capture means transcribing and summarizing what was said in the room. Comprehension means understanding what it clinically means once the recorder stops running. 

The risk compounds when ambient notes feed more than the chart, including downstream coding and clinical documentation integrity workflows, population health and quality measures, risk adjustment, and training data for the next wave of clinical AI models.  

When ungrounded notes become training data, ambient AI scales its own mistakes faster than any human reviewer can catch them. The gaps this creates are not just linguistic but contextual and may include historical trends, prior labs, treatment response, and observations never spoken aloud in the room. 

The scale of that reasoning gap shows up in the research. For example, an April 2026 JAMA Network Open study tested 21 leading large language models against 29 clinical vignettes. Every model posted a failure rate above 80% on differential diagnosis, even as final-diagnosis failure rates dropped below 40%.  

Producing a plausible answer and reasoning reliably toward it are different skills, and most models have mastered only the first. 

The missing layer: Clinically governed context 

A recent analysis in npj Digital Medicine found that while one-on-one conversations transcribe well, ambient scribes struggle to integrate data from multiple sources asynchronously, largely because of limited interoperability across EHR platforms, and that inadequate summarization leaves notes bloated with non-essential information.  

A separate study on coding accuracy found three recurring discrepancies between notes and visit diagnosis lists: diagnoses documented but omitted from the code list, diagnoses coded but never documented, and specificity mismatches, where less detailed codes replace more precise ones in ways that can degrade complexity measures like Hierarchical Condition Categories (HCCs).  

That research found operationalizing ambient AI required more than plug-and-play deployment, including a governance structure spanning technical integration, documentation compliance, and analytics-based decision rules to preserve coding integrity

What both studies point to is a missing layer, not a missing feature. A governed context layer looks like a clinician-maintained knowledge graph, refreshed continuously as medicine evolves, mapped to standards like ICD-10-CM, SNOMED®, LOINC®, and RxNorm®, and connected into the EHR through FHIR so clinical meaning travels cleanly across documentation, coding, analytics, and model training.  

This is infrastructure, not a feature bolted onto a scribe. No one builds a payments system without a ledger, and no one should build clinical AI without governed clinical context underneath it. 

This is infrastructure, not a feature bolted onto a scribe. No one builds a payments system without a ledger, and no one should build clinical AI without governed clinical context underneath it.

A data strategy instead of another documentation tool 

For health tech leaders building on ambient output, treat clinical semantics as a platform decision, not a downstream cleanup step, and judge AI output on more than fluency. Does it produce data a customer’s systems can use?   

For health systems buying ambient AI, ask vendors what happens to the note after the encounter ends, and how specificity, coding integrity, and analytics readiness survive that handoff. Don’t buy a documentation tool. Buy a data strategy. 

Capturing the conversation turned out to be the easy part. Understanding the patient, at scale and with clinical fidelity, is the work still ahead. The next era of healthcare AI won’t be won by whoever listens best. It will be won by whoever understands what was said. 

Building or buying ambient AI? Chat with us to learn more about IMO Health’s clinical context layer and how it supports semantic understanding. 

RxNorm® is a registered trademark of the National Library of Medicine.

SNOMED and SNOMED CT are registered trademarks of SNOMED International. 

Related Content

Latest Resources​

E/M downcoding issues often fly under the radar. Understand why this is, how to identify issues, and what steps can be taken
Most healthcare analytics tell us what happened. Real-time clinical intelligence can help us understand what is happening now.
Maintaining surgical dictionaries has become increasingly complex. Learn how surgical data governance protects revenue.
ICYMI: BLOG DIGEST

The latest insights and expert perspectives from IMO Health

In your inbox, twice per month.