Sanctions Screening Thresholds: How to Tune Without Drowning

Screening thresholds decide how many alerts you see and which true matches you miss. How to set, tune, test, and document them defensibly.

Updated On August 20, 2026
Sanctions Screening Thresholds: How to Tune Without Drowning

A screening threshold is the match score at or above which your system treats a candidate as an alert a human must review. Set it too high and you miss true matches you were obliged to catch. Set it too low and your team stops reading alerts properly, which produces the same outcome by a slower route.

No regulator publishes a number. What they publish is an expectation that whatever number you pick derives from your risk assessment, is tested before and after you change it, is monitored on an ongoing basis, and is documented well enough that someone else can reconstruct why. This guide is current as of August 2026.


What is a sanctions screening threshold?

A screening threshold is the minimum match score at which a screening system raises an alert for human review. Below it, a candidate is discarded as noise. At or above it, someone has to look. It is the single control that decides both how much work your compliance team receives and how much sanctions exposure you carry.

Two things follow, and both get missed constantly.

Scores are not comparable across vendors. A 75 on one engine and a 75 on another are not the same claim about the same name. Each scores a different set of signals with different weights on a scale it defined itself. Copying a threshold from a peer's programme tells you nothing useful.

A threshold is a display and action rule, not a matching rule. It does not change how well the engine recognises a name. Matching quality is decided upstream, in how the engine handles transliteration, name order, and spelling variance.


What threshold should I start with?

Start from your risk profile, not from a default. On a 0-100 scale where 100 is an exact name match, a standard regulated onboarding flow starts around 70. Higher-risk cross-border and trade activity starts lower, around 60-65, because missing a match costs more than reviewing one. Genuinely low-risk domestic activity can start higher, at 75-80.

Risk profileStarting point (0-100)WhyWhat you accept
Higher risk: cross-border payments, correspondent relationships, trade finance, customers near sanctioned jurisdictions60-65A missed designation is a strict-liability violation. An extra alert costs analyst minutes. The asymmetry is extreme, so buy recallA materially larger queue, and the staffing to clear it without rubber-stamping
Standard regulated onboarding: fintechs, EMIs, payment institutions, lenders70Around this level an exact all-token name match still clears the bar after one contradicting attribute (a wrong birth year on a list record), while a sound-alike-only hit with no corroboration does notSound-alike matches with zero corroborating data are dropped
Lower risk: domestic-only, single jurisdiction, low volume75-80Where local KYC already establishes identity and the customer base has little sanctions nexus, a weak fuzzy hit adds littleYou will not see borderline variants. Document why that is defensible
Back-book bulk screen: one-off screening of an existing customer fileStart at 75-80, then sweep downA first bulk screen of an existing book generates a one-time queue far larger than steady-state onboarding. Working the strongest band first resolves real exposure before the tailNothing, provided you run the lower sweeps. Stopping after the top band is the failure mode

These are starting points derived from what the score bands mean, not measured alert-volume benchmarks. Anyone quoting an industry-standard threshold without naming the engine it applies to is quoting a number that cannot mean anything.


Why is a single global threshold the wrong design?

Because one number has to serve situations with different legal consequences and different base rates of coincidental similarity. A sanctions designation and a PEP position are not the same finding. One triggers an asset freeze and a reporting obligation. The other triggers enhanced due diligence. Collapsing both into one cut-off means either drowning in PEP alerts or under-screening sanctions.

The Wolfsberg Group frames this as a list-management decision, asking firms to "determine which lists should be deployed within the screening filter on an exact match basis, and which would use fuzzy matching" [Wolfsberg 2019, §6]. The FFIEC BSA/AML Examination Manual frames it as risk calibration: "The screening criteria used by banks to identify name variations and misspellings should be based on the level of OFAC risk associated with the particular product or type of transaction" [FFIEC, OFAC section].

That does not mean you need five thresholds. In most systems the better answer is one cut-off plus differentiated handling above it.

DimensionWhy one number failsWhat to vary instead
List (SDN vs EU vs UN vs PEP)An exact hit on a designation list is a potential freeze. The same score against a PEP source is a due-diligence triggerA list-aware escalation rule: force exact designation-list matches to top severity regardless of numeric score
Entity type (individual vs organization)Company names are assembled from generic words. "Global Trading Solutions Ltd" matches every candidate sharing that boilerplateDown-weight matches built only from generic corporate tokens. Require a distinctive token or an identifier
PEP seniority and domesticityA local official at 85 and a foreign head of state at 85 are not the same riskCap and escalate by PEP tier and domestic/foreign status after scoring
Vessels and aircraftVessel names change with reflagging and are frequently generic. The IMO number is the stable signalPenalise name-only matches with no identifier corroboration
Customer risk ratingThe same score should not produce the same workflow for a high-risk and a low-risk customerRoute and staff differently. Change review depth, not visibility

The pattern is consistent: vary what happens to an alert, not whether the alert exists. Suppressing a match from view is the one adjustment that creates undocumented exposure, because nobody can review what was never surfaced.


How do I tune a threshold defensibly?

Run a documented loop: measure, sample the band you intend to change, adjust, test, then record the rationale before the change ships. The documentation step is not overhead. It is the entire difference between a defensible tuning decision and a finding.

1. Measure. Establish a baseline first: alerts per 1,000 screenings, true-match rate, median review time. Without it you cannot show the change did what you claimed.

2. Sample the band, do not count it. Moving from 70 to 75? Pull every alert that scored 70-74 over a meaningful window and adjudicate the full sample. Counting the band tells you the workload you would save. Only reading it tells you what you would have missed. One true match in that band takes the change off the table in its current form.

3. Adjust one variable at a time. A threshold change and a list-scope change shipped together cannot be attributed afterwards.

4. Test before and after. New York's Part 504 requires "end-to-end, pre- and post-implementation testing of the Filtering Program, including, as relevant, a review of data matching, an evaluation of whether the OFAC sanctions list and threshold settings map to the risks of the institution" [23 NYCRR 504.3(b)(3)]. It is a good standard even where it does not bind you.

5. Document the rationale. Wolfsberg is explicit that "a governance framework should contain the documented rationale for risk based decisions, such as those made in support of the creation of screening rules and threshold settings" [Wolfsberg 2019, §3.1]. Part 504 separately requires documentation of "the underlying assumptions, parameters, and thresholds" [23 NYCRR 504.3(a)(6), (b)(5)].

6. Re-review on a schedule. Part 504 requires ongoing analysis of "the threshold settings to see if they continue to map to the risks of the institution" [23 NYCRR 504.3(b)(4)]. A threshold set once stops matching the business within a year.

A complete change record answers six questions: what changed, from what to what, who approved it, what analysis supported it, what testing was performed, and what effect was expected. If your platform does not write that record automatically, write it by hand.


Which metrics should I track?

Track alert volume, true-match rate, median review time, escalation rate, decision-reversal rate, and control-name recall. The first five describe the cost of your threshold. The last describes what it misses, which is why it is the one most programmes never measure.

MetricWhat it tells youWarning sign
Alerts per 1,000 screeningsBaseline workload and trendA step change not explained by a known config or list change
True-match rate (confirmed / total alerts)Precision of the current settingNear zero across a quarter suggests the threshold or list scope is wrong. Very high suggests you only see obvious cases and the borderline band is being cut
Median review time per alertWhether alerts arrive with enough context to decideRising time usually means missing corroborating data, not a bad threshold. Tightening will not fix it
Escalation rate (alerts reaching EDD, case, or report)Whether alerts are decision-usefulA rate near zero over time means the queue is largely noise
Decision reversal or reopen rateAdjudication qualityRising reversals after a change mean it moved review pressure, not risk
Control-name recall (known positives caught)Direct measurement of what you missAny drop after a change. This is what stops tuning from being a guess

Wolfsberg lists metrics reporting as a core screening function, specifically "the number of alerts generated by list, by jurisdiction, by business, or the identification of unintended data and list omissions" [Wolfsberg 2019, §3.3]. Split alert counts by list early. It is usually one or two sources driving most of the volume.


When does threshold tuning become a regulatory finding?

When the reason for the change is the size of the queue rather than the quality of the matches. Two identical changes, one supported by a documented sample showing the band held no true matches and one made because reviewers were behind, look the same in the settings screen and are completely different under examination. The analysis separates them, and it only exists if you wrote it down at the time.

Both directions are examinable, which is the part programmes underestimate.

Tightening too far suppresses matches nobody ever sees. Wolfsberg sets out the consequence: where a sanctions data point may have gone previously undetected as a result of a name variation, a firm should consider both whether configuration changes are warranted and whether a lookback over already-processed activity should be performed [Wolfsberg 2019, §7]. A lookback is expensive, public, and slow.

Loosening without capacity is also a finding. The FFIEC manual is direct that a high volume of false hits may indicate a need to review the interdiction programme [FFIEC, OFAC section]. A queue nobody can clear is not a conservative posture. It is an unreviewed backlog with an obligation attached to each item.

OFAC's own framework names "sanctions screening software or filter faults" among the root causes behind enforcement actions, including failures to account for alternative spellings of designated parties [OFAC 2019]. Screening configuration is a named enforcement risk, not a background setting.

The legitimate ways to cut review load, none of which cut recall:

  • Improve corroborating data. Capturing date of birth and nationality at onboarding lets the engine separate two people with the same name instead of escalating both.
  • Whitelist confirmed false positives. Wolfsberg describes rules "for automatically eliminating potential hits caused by the interaction of certain list terms and frequently encountered data, for example, customer names which have already been confirmed as false positives" [Wolfsberg 2019, §6]. That removes repeat work, not coverage.
  • Do not screen weak aliases. On low-quality aliases (very short strings, aliases containing digits, nicknames such as "Ahmed the Tall"), Wolfsberg is blunt: "it is not expected, nor is it typically productive, to screen against weak aliases" [Wolfsberg 2019, §6.3].
  • Gate automated case creation, not visibility. Restricting which alerts open a case automatically is a workload control. The matches stay recorded and reviewable, so nothing is hidden.

The red line is simple. Any change that reduces what your system surfaces needs evidence that the removed band held no true matches. Any change made for workload reasons alone must change how alerts are handled, not whether they appear.


How do I test a threshold change before shipping it?

Maintain a control set of known-positive names and run it before and after every change. Draw the names from the lists themselves, then add deliberate variants: alternative transliterations, reversed given and family name order, single-character typos, dropped honorifics, and a common name that should produce noise. Compare results name by name against the current configuration.

Three practices make the test worth running. Include known negatives, since a set of only true positives tells you about recall and nothing about precision. Run in parallel where you can, scoring against both configurations on live traffic rather than switching and hoping. Refresh the set when lists change, because a control name that gets delisted stops testing anything.

Do not validate a change by watching production alert volumes for a week and declaring success because the number moved the way you wanted. That measures workload, which was never the thing in question.


How does DeRisk Hub handle screening thresholds?

DeRisk Hub uses one configurable match threshold per organisation, defaulted to 70 on a 0-100 scale and adjustable between 50 and 100. Scores below 50 are excluded platform-wide as low-confidence noise, so the floor is not something an administrator can tune away. Above the threshold, differentiation is handled by rules rather than extra cut-offs.

  • Risk bands classify each match by severity: Critical at 90 and above, High at 80-89, Medium at 60-79, Low below that. The band drives triage and routing. It does not decide visibility.
  • A list-aware rule promotes an exact name match on the OFAC SDN list to Critical regardless of where the numeric score lands.
  • PEP matches are capped and escalated by tier and domesticity after scoring. A tier 1 foreign position with an exact name match escalates to Critical. A tier 3 domestic position is capped at Low.
  • Scoring penalises weak evidence rather than leaving it to the threshold: matches built only from generic corporate tokens score far lower than those carrying a distinctive token, and matches with no name signal, or with a contradicting birth year, nationality, or country, are penalised.
  • Critical-only case creation restricts automatic case creation to Critical matches while everything else above the threshold stays recorded and reviewable. That is the safe version of "reduce the queue."
  • Every settings change is written to the audit log with the previous value, the new value, and the user who made it, so the change record Part 504 and Wolfsberg ask for exists without anyone remembering to create it.

Where we are honest about the limits: today that is one threshold, not one per list or per entity type. Differentiation lives in scoring and severity instead, which we think is the better design for most programmes and is certainly the harder one to misconfigure. If your programme genuinely needs independent cut-offs per list, that is a legitimate requirement and we would rather hear it than have you assume the answer is no. Tell us what your programme needs.


Frequently asked questions

What is a good sanctions screening match threshold? There is no universal number, because scores are not comparable between engines. On a 0-100 scale, 70 is a reasonable starting point for standard regulated onboarding, 60-65 for higher-risk cross-border activity, and 75-80 for low-risk domestic-only businesses. Whatever you start with must be measured against your own alert outcomes and adjusted with a documented rationale.

Do regulators specify a screening threshold? No. No major regulator publishes a required match score. They require that the setting derives from your risk assessment, is tested before and after implementation, is subject to ongoing analysis, and is documented. New York's Part 504 is the most prescriptive on process and still specifies no number, noting that it does "not mandate the use of any particular technology, only that the system or technology used must be reasonably designed to identify prohibited transactions" [23 NYCRR 504.3(b), n.4].

Should PEP and sanctions screening use the same threshold? They can share a numeric threshold, but not an outcome. A sanctions match is a potential asset-freeze and reporting event. A PEP match triggers enhanced due diligence and is not a prohibition. Differentiate with severity rules and workflow, by PEP tier and by domestic versus foreign position, rather than two cut-offs reviewers have to reconcile.

Can I lower my threshold to reduce false positives? No, and the question usually reflects a reversed mental model. Lowering the threshold admits weaker matches and increases alert volume. Raising it reduces volume by discarding the weakest band, which is exactly where genuine variant spellings live. If false positives are the problem, address them through corroborating data, whitelisting confirmed false positives, and better alert presentation, not by cutting recall.

How often should thresholds be reviewed? At minimum annually, and after any material change to your customer base, product mix, geographic footprint, list coverage, or matching engine. A sharp move in alert volume with no known cause is itself a trigger, in either direction.


Citations


Set it, measure it, write it down

A threshold is the one screening setting where the safe-looking choice and the defensible choice can point in opposite directions. The number matters less than whether you can show how you arrived at it and what you tested before you moved it.

DeRisk Hub gives you a configurable match threshold with a hard low-confidence floor, severity banding and list-aware escalation above it, and an audit record of every configuration change, so the tuning history examiners ask for already exists. See the platform overview, how fuzzy name matching produces the scores you are thresholding, or what sanctions screening involves end to end. Start your free trial, or go to DeRiskHub.com.

This article is informational and does not constitute legal advice. It describes publicly available regulatory material as verified on 19 August 2026; confirm your specific obligations with qualified counsel.