Auditing Evidence Use in Medical LLM Diagnosis
Abstract
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately.
We present a behavioral audit of evidence use in medical diagnosis.
For each case, we decompose patient information into evidence units, score candidate diagnoses under controlled evidence subsets, and mine low-order interactions in diagnostic margins.
Because medical evidence is diagnosis-relative, the audit separates interaction discovery from failure assignment: large or negative interactions can reflect plausible differential diagnosis, while suspicious interactions require robustness checks and clinical review.
We evaluate five open-weight LLMs on DDXPlus, CupCase, and MedCase.
Across datasets, faithful support and differential conflict or cancellation account for most interaction strength, showing that many evidence interactions are clinically plausible rather than failures.
In a DDXPlus-focused blinded five-reviewer 130-item enriched review sample, invalid or shortcut-like cases concentrate in negated or absent findings and clinically local evidence.
These results show that accuracy can hide candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation.
이 뉴스, 어떠셨어요?
탭 한 번으로 반응 · 로그인 불필요