Direct Answer to Payment Fraud Risk Thresholds

There is no universal payment fraud risk threshold that is correct for every merchant, card scheme, geography, or product. A practical starting point is to treat a transaction score below roughly 30 as low risk, 30–69 as medium risk, and 70 or above as high risk, but those numbers only make sense if they come from the platform’s own calibrated model. For card-not-present payments, a score of 60–69 may justify step-up authentication, while a score of 80 may justify blocking outright; those are operational defaults, not industry rules. The better approach is to combine a model score with transaction value, customer history, device quality, authentication method, and expected loss. As of 28 September 2026, modern risk systems increasingly use explainable machine learning, and schemes such as Mastercard continue strengthening scam defenses. A merchant should set thresholds from observed outcomes, monitor approval and loss rates daily, and revise them by market rather than adopting a single global cutoff.

Also worth reading: What Should Consumers and Merchants Secure Before Using Payment Apps in 2026? · What Is the Safest Payment Processor Migration Checklist for Merchants Switching in 2026? · How Do Digital Payments Workflow Guides Help Merchants Choose Wallets, Gateways, and Payment Tools?

A score cannot stand alone because fraud losses and customer friction have different economics. Blocking every transaction above 50 might reduce some fraud while also rejecting many legitimate customers, whereas blocking only transactions above 80 can leave common account-takeover or friendly-fraud patterns below the cutoff. Practical thresholds should therefore be divided into actions such as approve, review, authenticate, delay, or decline, with separate limits for card-present, card-not-present, new-device, and high-value activity. The objective is not the lowest fraud rate; it is the lowest total cost of fraud plus acquisition loss, chargeback expense, support cost, and customer inconvenience.

How Payment Fraud Scoring Actually Works

Payment fraud risk scores are generally produced from several layers of information. The first layer contains transaction attributes, such as amount, currency, merchant category, time, billing and delivery address, and whether the card and browser belong to a previously seen customer. The second contains identity and behavioral signals, including login age, account changes, typing speed, navigation paths, session history, and deviations from normal purchasing patterns. The third contains network and authentication signals, such as device reputation, IP address, proxy or virtual-private-network use, biometric or one-time-password results, and prior confirmed fraud. Models then estimate probability or assign a score, although vendors do not always reveal the precise probability behind a 0–100 risk rating.

A merchant should distinguish three outcomes that are sometimes incorrectly collapsed into one “fraud rate.” Authorization fraud is a declined or unauthorized card transaction; disputed fraud involves a cardholder later denying a charge; and merchant or consumer abuse may include bots, stolen accounts, fake accounts, or excessive transaction attempts. Each outcome has a different prevention mechanism and reporting delay. A rule that prevents a stolen card from authorizing may have little effect on a later dispute, while a customer who receives an unwanted but technically authorized payment may never trigger a model block. This distinction matters because a weekly dashboard cannot fully evaluate a decision until disputes and customer reports have had time to appear.

Vendors such as Stripe Radar, Mastercard, and other payment platforms offer proprietary scoring, but customers still need an economic policy around it. Radar can evaluate signals available to Stripe, while merchant systems may know whether an order is being shipped to a different country, whether the coupon was newly created, or whether the customer changed an email address immediately before checkout. The platform sees a fast and broad data set; the merchant sees business-specific context. Neither view is complete, which is why rules based on internal data can add value to a vendor score rather than supposedly replacing it.

Recommended Thresholds by Transaction Type

The safest initial thresholds depend heavily on where the payment occurs. For established, low-risk card-on-file transactions, a score of 30 or less might be approved, 31–69 approved with monitoring, and 70 or more reviewed. For card-not-present orders with new devices or addresses, review may begin near 50–60, and declines may begin around 80–90. High-value orders, cryptocurrency purchases, travel tickets, gift cards, and electronics deserve separate models because their loss and dispute patterns differ. These ranges are starting points for testing, not promises, and a merchant should avoid applying one checkout threshold to all card categories.

An illustrative policy might approve transactions below 45, automatically request 3-D Secure for scores from 45 to 69, and route scores from 70 to 84 to manual review or a stronger challenge. Scores of 85 or more could be declined, subject to a recovery path for trusted customers. Manual review should have a service-level target, such as reviewing 95% of flagged orders within 15 minutes during staffed hours, because a “review” that sits untouched for six hours is only a delay rather than a decision. High-value orders may use proportional controls—for example, requiring additional authentication for an order equal to five times the customer’s previous maximum—rather than relying only on an absolute dollar cutoff.

FeatureLow-Risk TransactionMedium-Risk TransactionHigh-Risk Transaction
Illustrative model score0–2930–6970–100
Typical actionApprove and monitorStep up or reviewDecline or strict review
AuthenticationExisting account and normal sessionOTP, passkey, or 3-D SecureAdditional identity evidence or decline
Useful contextKnown customer, trusted device, low valueNew address, device change, moderate valueVelocity spike, impossible travel, high value
Review priorityNoneMinutes to a few hoursImmediate where operations allow
Optimization targetPreserve acceptanceBalance friction and fraudMinimize expected loss
These score bands should be paired with velocity limits. A common starting policy might allow no more than three failed payment attempts per card, five per device, and ten per account in 24 hours. Another practical control is to stop a buyer from changing the shipping address, email, password, and payment method repeatedly within one short session. Exact limits should reflect customer volume, average order value, and staffing, since a five-attempt limit that is harmless to most subscribers may block a family or business traveler.

Turning Scores Into Practical Decisions

A useful decision formula considers expected fraud loss, authentication cost, handling cost, customer lifetime value, and the probability that a genuine customer will abandon. If a transaction is worth $30 and the platform’s chargeback handling cost is $15, a $2.40 expected loss may justify a low-cost challenge, but not necessarily a high-cost manual review. If the transaction is $2,000 and expected loss approaches $400, a small added verification cost can be rational. Merchants should not treat 3-D Secure success as proof of fraud prevention: the protocol improves issuer and cardholder authentication, but account takeover, social engineering, and certain disputes can survive it.

The best workflow is to connect every decision to an outcome. Record the risk score, threshold band, rule fired, action taken, authentication result, final authorization, and later dispute status. Then segment the data by country, card brand, device, customer tenure, order category, and authentication method. After sufficient observations, lower a threshold if a segment has low confirmed loss and high approval, or raise it if there are repeated losses with little fraud reduction. A useful target might be to keep confirmed fraud below 0.1% of revenue for a low-risk subscription product, while accepting that travel, ticketing, or high-resale markets may require a higher rate and stronger controls.

Changes should be gradual. A simultaneous cut from 70 to 50 might look effective in the first day because high-score traffic is blocked, but later dispute data may reveal unnecessary rejections. Test one major policy change at a time, use a holdout group when volume permits, and set a predetermined review date. Stripe Radar can be tuned, and broader industry work on explainable supervised models supports the use of understandable feature effects, but “explainable” does not guarantee correct calibration or freedom from bias. The merchant remains responsible for testing how its own customers and markets behave.

Comparison of Fraud Prevention Alternatives

There is no single control that dominates the others. A velocity rule is cheap and transparent but cannot recognize a sophisticated fraud pattern. A machine-learning score can evaluate many signals but may be opaque, poorly calibrated, or blind to current events. 3-D Secure adds issuer authentication, yet it introduces friction and does not cover every scam. Manual review can examine unusual orders closely, but it is expensive, inconsistent, and unavailable around the clock. A tokenized card-on-file environment can reduce exposure to stored card details, while strong customer authentication and account security address unauthorized access more directly.

FeaturePlatform Risk ScoreCustom Rules3-D SecureManual Review
Main advantageEvaluates many live signalsFast, transparent, and customizableAdds issuer/cardholder authenticationApplies human judgment
Main weaknessProprietary calibration and blind spotsBrittle if rules are too broadFriction and incomplete scam coverageCostly and inconsistent
Typical usePrimary risk rankingHard limits and overridesMedium- or high-risk checkoutAmbiguous or high-value cases
Operating costOften included or usage-pricedLow technical costUsually priced per attempt or contingent on outcomeHighest direct labor cost
ExplainabilityVaries by vendorHighHighHigh, but not perfectly repeatable
Rather than asking which alternative is “best,” merchants should ask which failure each control handles. Radar-like scoring is useful for first-pass ranking. Custom rules should enforce non-negotiable limits, such as blocking a newly added card after repeated account changes. Strong authentication should protect account logins and high-risk payment changes, while manual review should be reserved for cases where evidence is incomplete and expected loss justifies the labor. These tools work better as layers, although adding every layer indiscriminately will turn a checkout into an obstacle course and can increase abandonment without preventing targeted fraud.

Tokenization and recurring billing provide another alternative to repeatedly evaluating a new public card entry. A stored credential can lower checkout friction and improve conversion, but it can increase the damage from weak account recovery. Customers should still face strong login controls before saving, viewing, or changing a card. Merchants that emphasize subscriptions may prefer lower risk thresholds and account-level monitoring, while high-ticket one-time sellers may use stricter step-up authentication only above a value or behavioral trigger.

Common Mistakes in Setting Fraud Thresholds

The most common mistake is copying another merchant’s thresholds. Fraud patterns differ by average order value, geography, product, and customer behavior, so a cutoff tuned for a $20 digital subscription may be ineffective for a $3,000 international electronics order. A second mistake is focusing on the model’s raw score as though it were a stable probability. Scores may be reordered after model updates, and a value of 80 in one system need not represent the same probability in another. A third mistake is declining without a recovery option; many legitimate customers mistype addresses, switch networks, travel, or use new devices.

Another error is overfitting to recent attacks without measuring customer impact. A sudden block in one country can appear effective because fraud is low for several hours, even though honest travelers are being lost. Conversely, waiting for finalized chargeback data can make tuning months behind events. Merchants should use provisional fraud reports, customer trust-and-safety tickets, account-takeover signals, and confirmed disputes together, while remembering that observed loss is itself an imperfect label. Not every report becomes a liability, and some fraud never generates a formal dispute.

Finally, teams often set thresholds once and never revisit them. Payment infrastructure changes: new authentication methods appear, bot tactics adapt, and schemes respond to scam campaigns. Mastercard’s ongoing scam-defense work and dLocal’s reported use of Oscilar for AML-related compliance show that the broader industry is investing in intelligence, but merchants still need scheduled governance. At minimum, review the policy monthly, test major changes against a holdout group, and document the rationale for each threshold. Governance should also record who can override a block, because inconsistent emergency overrides can quietly erase the value of the model.

When Merchants Should Act Immediately

Immediate tightening is appropriate when there is a verified account-takeover spike, a known stolen-card campaign, a bot attack, or a sudden fall in authentication success. For example, a merchant seeing 25 payment failures and eight account-takeover alerts in ten minutes may reasonably impose a short, account- or device-level block. The response should be scoped, time-limited, and reviewed. A permanent geographic restriction based only on anecdotal suspicion can create regulatory, customer, and revenue problems that outweigh the avoided fraud.

Act more cautiously when a new model version, card scheme rule, or authentication change takes effect. Run shadow scoring first when possible, compare the proposed decisions with the current policy, and estimate lost authorization revenue and fraud exposure. A threshold adjustment should have an owner, effective date, expected effect, and rollback condition. If expected fraud loss falls by $1,000 but the team expects 1,000 additional legitimate orders worth $25 each, the apparent savings could be negative after contribution margin and chargeback costs are considered.

Seasonal merchants should prepare before predictable peaks rather than responding after a launch. Travel sellers may need stronger rules for last-minute, high-value itinerary changes; ticket resellers may see abnormal fans and payment changes; gift-card sellers should monitor rapid card testing. A merchant can lower friction for customers with long account tenure and a consistent device while increasing verification after recent email, password, or payment changes. That narrow rule may be more defensible than raising the threshold for the entire customer base.

Cost is relevant, but pricing should be compared on total operating expense rather than a headline percentage. A vendor’s fee may be a per-transaction, per-report, or plan-based charge, and acquirers may price premium risk intelligence differently. Manual review at 300 reviewed orders per day, three minutes per order, and a fully loaded labor cost of $25 per hour costs about $375 daily. A scoring or authentication product that costs less than $375 may be economical, but only if it actually prevents sufficient loss or saves equivalent labor. Merchants should include software fees, implementation, false declines, customer support, chargebacks, engineering time, and reporting in the comparison.

A Defensible Threshold-Setting Policy

Begin with a written policy separating low, medium, and high risk, then map each band to an action. Start with 0–29, 30–69, and 70–100 score bands, but calibrate the underlying thresholds using at least several months of transaction and dispute data. Add separate policies for high-value orders, new customers, new devices, credential changes, and rapid attempts. Review the thresholds weekly during a major incident and at least monthly in stable periods, while requiring approval for permanent changes from both payments operations and a suitable risk or compliance owner.

Measurement should include more than gross fraud dollars. Track authorization rate, approval rate by value, step-up challenge completion, manual-review time, confirmed fraud basis points, dispute loss, account takeover, repeat attacker recurrence, and customer complaints. One useful reporting structure is 100 basis points of revenue, or 0.1%, rather than mixing percentages across unrelated measures. A business that reduces confirmed fraud from 0.20% to 0.08% is eliminating 0.12% of revenue, but it should also state how much revenue was rejected and what contribution margin was lost. Without that denominator, improvement can be overstated.

The defensible policy is therefore not “block above 60” or “block above 80.” It is a documented sequence of evidence, actions, and performance measures that can be tested as conditions change. A trusted customer with a normal history should generally experience fewer interruptions, while an account making rapid credential changes, repeated card attempts, and shipping anomalies should receive a stronger challenge. By 2026, machine-learning scores and scheme-level scam defenses can improve prioritization, but they do not replace sound economics, operational review, or customer-friendly recovery. The best thresholds are those that prevent more expected loss than the friction they create—and remain that good six months later.