A bank's transaction monitoring system raises tens of thousands of anti-money-laundering alerts a month, and about 92% of them are cleared as false positives after a human investigation that costs roughly $50 each. The compliance team cannot hire its way out. You are building the triage model that decides which alerts an analyst sees first and which are auto-closed. Send too few and confirmed laundering slips through, which is a regulatory fine. Send too many and the queue is the same queue you started with. Your job is to compute the metrics honestly, beat the legacy score, and defend an operating point in dollars.
Establish how much of the legacy alert queue is noise, and which rule produces the most of it.
Implement explore_alerts(train_df) returning a dict with:
The legacy false-positive rate is the number the compliance lead quotes in every meeting. The last two numbers are why you must not impute the missing history away.
Evaluated server-side against a hidden test set.