Why ROC-AUC can mislead on rare-positive problems and how calibration and cost-based thresholds fit in.
A fraud model is trained on data with 0.5% positives. It shows ROC-AUC of 0.97 but stakeholders complain about too many false alarms in production. Which analysis and fix is MOST appropriate?