Tech & AI News
Hacker News

LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes

Ambient AI scribes produce clinical notes that frequently omit information established during encounters. Standard LLM judges detect added or altered content with 0.79–0.94 accuracy but identify omissions only 0.50–0.63, failing to flag missing facts reliably. A restructured task—listing transcript facts and checking each against the note—yields a per-fact pipeline with 2.7% false alarms and a single-call prompt that detects 36.9% of omissions at 6.2% false alarms, outperforming all eight baseline judges.