AI x Science Seminar: Mert Sabuncu
Ways to handle distribution shift and missingness in AI for medical diagnosis
Amy Gutman Hall, Room 414
Medical diagnosis can be naturally framed as a classification problem: inferring an underlying pathology from observed (e.g., imaging) data. A common failure mode in classification is shortcut learning, where models exploit spurious or confounding correlations. Shifts in patient populations across sites can therefore alter both labels and inputs, making naive robustness or invariance strategies insufficient. I will introduce Conditional and Generalized Prevalence Adjustment (CoPA [1] and GPA [2]), our proposed approach for explicitly modeling and adapting to prevalence shifts rather than suppressing them. I will then show how practical deployment of prevalence adjustment requires robustness to missing or high-dimensional confounders, and present Knockout [3], a simple, theoretically grounded training strategy.
Presented by
Penn AI, Innovation in Data Engineering and Science (IDEAS), and the Data Driven Discovery Initiative (DDDI)