Blog · FEB 5, 2026 · 4 min read
AI Note-Taking for Doctor Visits: What It Gets Right and Wrong
Our AI turns doctor visits into structured, compliance-ready notes most of the time. Here is where it still needs a human in the loop.
Every verified visit on Connex generates an AI-extracted note: a structured summary of what was discussed, along with any compliance flags worth a second look. We think this is one of the more genuinely useful pieces of the product, and we also think it is worth being specific about where it holds up and where it does not.
Where it works well
For visits conducted in a single language, with clear audio and a reasonably structured conversation, extraction quality is strong. The notes capture the substance of what was discussed, correctly identify samples or materials referenced, and flag language that brushes up against UCPMP boundaries around gifting or promotional claims with good precision. Reps consistently tell us the generated notes are more complete than what they would have written themselves after the fact, mostly because nothing depends on memory anymore.
Where it still struggles
- •Code-switching mid-sentence between Hindi or a regional language and English, common in real clinical conversations, occasionally produces extraction errors that a native listener would not make
- •Overlapping speech, common when a clinic is busy or when more than two people are in the room, degrades transcription quality before it degrades the note itself
- •Highly specialized medical terminology outside a rep's usual product line is sometimes transcribed phonetically rather than correctly, which can distort a note's clinical detail even when the compliance-relevant content is fine
- •Sarcasm or hedging language, where a doctor politely declines something in a way that is not a direct no, is the single hardest category for the model to flag correctly
How we handle the gap
Every note carries a confidence score, and anything below our threshold routes to human review rather than being treated as final. We would rather a compliance team see fewer, more reliable flags than a flood of noisy ones that trains them to stop reading.
We do not think AI note-taking will ever be perfect for a domain this linguistically and clinically varied, and we are not chasing perfection. We are chasing a system where the failure mode is a human catching what the model missed, never a compliance issue slipping through because everyone assumed the AI had it covered.