99% False Positives is effectively a compliance tax
Sanctions screening false positive rates of 95–99% are treated as normal. Why legacy name matching drowns compliance teams, and how AI-driven contextual screening fixes it
99% False Positives is effectively a compliance tax
Somewhere in your compliance team, right now, an analyst is clearing an alert explaining that your investor Mohammed Khan is not the sanctioned Mohamed Kahn, born thirty years earlier in a different country.
They cleared a nearly identical alert last month. And the month before. The file even says so. But the system alerted again, so the ritual repeats: open, compare, document, dismiss. Multiply by every common name on your book, every list update, every rescreening cycle, and you arrive at the industry's most normalised absurdity - screening operations where the overwhelming majority of alerts, routinely 95% and up, are noise.
We've stopped even flinching at this. Ask why, and you'll hear that false positives are the price of safety. That's the assumption worth attacking, because it has the logic exactly backwards.
Noise is risk
The defence of over-alerting is that it's conservative - better a thousand false alarms than one miss. This would be true if humans weren't in the loop. But they are, and human attention doesn't scale with alert volume.
An analyst clearing forty alerts a day, thirty-nine of which are obviously nothing, is being trained - methodically, daily - to expect nothing. The true hit, when it finally comes, looks exactly like the noise it's buried in, and it gets the same ninety seconds of attention as the Mohammed Khan ritual. Every enforcement action that begins with "the alert was generated but dismissed" is this mechanism playing out.
Why the matching is this bad
The root cause is that legacy screening asks a primitive question: does this string resemble that string? Fuzzy matching on names - across transliterations, spellings, and truncations - inevitably fires on the world's millions of Mohammed Khans and Maria Gonzalezes.
But here's the thing: the disambiguating information almost always exists. Your file knows the investor's date of birth, nationality, address, occupation. The list entry carries its own identifiers. A human analyst resolves the alert precisely by comparing this context - that's what the ninety seconds consist of. The legacy system simply doesn't use any of it at matching time. It matches on strings, then hands the context work to a person.
Read that again: the industry's screening cost problem is a system generating work it already has the information to resolve. The 99% false positive rate is an architecture decision from an era when using context automatically wasn't feasible.
It's now feasible
This is one of the cleanest applications of modern AI in compliance, because the task is exactly what the technology is good at: reading a potential match in full context: the file, the list entry, the secondary identifiers, the prior dispositions, and reasoning about whether these are plausibly the same person, the way a good analyst would, at machine speed and cost.
Note what this is not: it's not raising the fuzzy-match threshold and hoping. Cranking the sensitivity dial down reduces alerts by reducing looking, which regulators rightly hate. Contextual adjudication is different in kind - the net stays wide, catching every candidate match, and the intelligence moves to disposition, where each candidate is investigated rather than merely pattern-scored.
And because every disposition is written down - here's the alert, here's the evidence compared, here's why it cleared or escalated - the auditability improves over the human baseline, where the reasoning too often lives in a tired analyst's head and a two-word case note.
What to do about yours
1. Measure the ritual
Pull the numbers: alerts per month, true-positive rate, average handling time, repeat-alert share. Most firms have never computed the fully-loaded cost per meaningful alert. It's a clarifying number to put in front of a budget owner.
2. Attack repeat alerts first
The same investor alerting against the same list entry every cycle, cleared identically every time, is pure waste - and the easiest category to fix with disposition memory. If your system can't remember last month's decision, that's your first vendor question.
3. Insist on evidence-based clearance, not threshold tuning
When evaluating tooling, the question isn't "how much does it cut alerts?" - anything can cut alerts by going blind. The question is "show me the reasoning trail for each cleared alert." Regulators will ask the same thing.
4. Redeploy the hours where judgement pays
Every analyst hour reclaimed from name-ritual is an hour for the work that actually needs humans: the genuine matches, the complex structures, the escalations. Screening reform is a reallocation story.
Where Steward fits
Steward screens the way a great analyst would if they had infinite time: every alert investigated in full context, every disposition evidenced, every prior decision remembered across sanctions, PEPs and adverse media. Our customers' teams see the handful of alerts that deserve human judgement, with the investigation already assembled. The Mohammed Khan ritual ends.
A 99% false positive rate is an architectural problem wearing a prudence costume, taxing your team and degrading your detection at the same time. The context needed to resolve almost every alert was in your files all along. The technology to use it now exists in Steward.
Stop paying the tax.
Related Insights

The Best PEP and Sanctions Screening Vendors in 2026
Compare PEP and sanctions screening vendors: which sell data, which sell platforms, and how to run a screening proof of concept
Sanctions Screening Before a Fund Distribution
A UK fund distribution is due and an investor is flagged. Follow the decisions on identity, ownership, payment controls, escalation and evidence.

Former PEP Due Diligence: Reassessing the Investor
How UK investment firms can reassess former PEPs, document continuing risk and update investor controls after a public role ends