How Many Good Customers Did Your Fraud Controls Turn Away?
I've spent a lot of my career working on the parts of ecommerce customers rarely see: fraud reviews, marketplace policies, returns, chargebacks, fulfillment rules and the operational decisions sitting between someone clicking Buy and actually receiving their order.
And there's one question I keep coming back to:
How many good customers did we stop while trying to stop the bad ones?
Most merchants can tell you how many chargebacks they had last month.
They can usually tell you how much fraud they prevented.
But ask how many legitimate customers their controls declined, delayed or frustrated enough to lose the sale, and the answer gets much harder.
The problem is structural.
When an order is cancelled for suspected fraud, it never ships. It never gets the opportunity to prove that it was legitimate. There's no successful delivery, no retained customer and no clean transaction afterward.
The control worked according to the dashboard.
But did it actually make the right decision?
That's a different question.
Key takeaways
When fundamentally different problems get compressed into one number, the decision becomes harder to inspect.
Merchant-level fraud systems see transactions. Platforms can see networks.
Risk avoidance and risk management are not the same thing. Operations is usually about finding the lowest-cost acceptable response, not eliminating every possible downside.
Trust should reduce friction, not eliminate verification.
The point isn't to allow fraud. The point is to measure the cost of a decision that would otherwise never be tested.
The workflow around the decision matters as much as the decision itself.
Before giving a system broad authority to cancel transactions, help merchants recover losses they're already entitled to dispute.
The system's job is finding the best operating point between how much loss you'll accept and how much friction you'll impose.
A merchant doesn't need every internal calculation. But they should understand the evidence, the available actions, and the tradeoff the system made.
Measure. Observe. Human-reviewed actions. Automate last.
Field Research — 10 Chapters
Loss isn't one problem
“Fraud” is a terrible catch-all term.
One thing years of ecommerce operations taught me is that “fraud” is a terrible catch-all term.
A stolen card, a false delivery claim, return abuse and promotional gaming may all cost the merchant money, but they're not the same problem.
I generally think about commerce loss in separate lanes:
- Payment fraud. Stolen cards or unauthorized payment methods.
- Return abuse. Swapped products, empty packages, excessive patterns or other misuse of returns.
- Delivery claims. “Never arrived” claims where available evidence may suggest otherwise.
- Promotional abuse. Multiple accounts, repeated new-customer offers or systematic exploitation of incentives.
Each requires different evidence.
Each has different economics.
And each should have different responses.
A signature requirement might help with repeat delivery claims.
It does almost nothing to determine whether a payment method was stolen.
A photo requirement can be useful for certain return claims.
It doesn't belong in the checkout decision.
Yet many systems still try to compress these different behaviors into one score.
Risk score: 82. Decline.
That's where useful context starts disappearing.
I've learned the same lesson building catalog, compliance and commerce systems: when fundamentally different problems get compressed into one number, the decision becomes harder to inspect.
Risk is no different.
Key Insight
When fundamentally different problems get compressed into one number, the decision becomes harder to inspect.
Sometimes the pattern is bigger than one order
The dangerous stuff doesn't always arrive with a giant fraud score attached.
Some of the most meaningful fraud patterns I've dealt with didn't become obvious from a single transaction.
Over several years, across multiple Shopify stores, I kept encountering the same kind of coordinated abuse tied back to Doral, Florida.
- Different customer names.
- Different accounts.
- Different addresses.
- Different stores.
But the same geography and many of the same underlying behavioral patterns kept resurfacing.
Looked at one order at a time, some of those transactions could pass as ordinary ecommerce noise.
Looked at across stores and over time, the pattern was much harder to ignore.
I helped identify and disrupt that activity by connecting the history, tightening the operational controls around it and making sure the signals weren't evaluated in isolation.
That experience changed how I think about ecommerce risk.
The dangerous stuff doesn't always arrive with a giant fraud score attached to it.
Sometimes it looks completely normal until you connect enough context.
And it also raises a much bigger question.
A merchant sees one store.
Shopify can potentially see the network.
If the same IP ranges, device characteristics, shipping destinations, account behaviors and timing patterns repeatedly appear in confirmed abuse across otherwise unrelated merchants, that is a much stronger signal than anything a single store can see on its own.
That doesn't mean geography or a shared IP should ever equal guilt.
Doral has legitimate businesses and customers like anywhere else. Shared networks, VPNs, freight forwarders and dense commercial areas can create false patterns very quickly.
But when multiple independent signals keep converging across stores, the platform has an opportunity individual merchants simply don't.
I would love to see Shopify use more ecosystem-level intelligence here.
Not another opaque blocklist.
A merchant-facing signal that says, in effect:
We've seen this infrastructure or behavior associated with confirmed abuse across multiple stores. Here's the evidence, here's the confidence level, and here's why we're surfacing it.
That would be far more useful than treating every merchant as an island.
Key Insight
Merchant-level fraud systems see transactions. Platforms can see networks.
A verdict isn't a decision
There's a lot of room between approve and decline.
Most fraud systems ultimately produce some version of:
Approve or decline.
But ecommerce operations has a lot of room between those two outcomes.
You can:
- Require a signature.
- Hold an order for review.
- Verify information with the customer.
- Remove promotional eligibility.
- Require supporting evidence for a claim.
- Restrict an account without cancelling the current transaction.
- Re-review an order when important information changes.
Those actions have very different costs.
That's why I think the better question is:
What's the least expensive action that protects the business without unnecessarily creating friction for a legitimate customer?
Imagine a $120 order with some measurable risk of payment fraud and delivery dispute.
A simplified decision model might look something like this:
| Action | Approx. expected cost |
|---|---|
| Ship normally | $16 |
| Require signature | $16 |
| Hold for quick review | $11 |
| Hold and verify customer | $14 |
| Cancel | $70 |
Illustrative figures combining residual loss, intervention cost, customer friction and lost future value—not production results.
The interesting part is the last row.
Cancelling may create the least immediate exposure.
But it can still be the most expensive decision.
If most flagged customers are legitimate, unnecessary cancellations cost revenue, future purchases and customer trust.
That distinction matters.
Key Insight
Risk avoidance and risk management are not the same thing. Operations is usually about finding the lowest-cost acceptable response, not eliminating every possible downside.
Good customers should build credit
Trust should reduce friction, not eliminate verification.
Risk systems are generally very good at accumulating negative evidence.
- New device.
- New card.
- Address mismatch.
- High order value.
- Multiple accounts.
- Unusual location.
- High return rate.
What interests me just as much is the positive evidence.
A customer who has successfully completed twelve orders has told you something.
- Their orders were paid.
- They shipped.
- They arrived.
- They stayed kept.
That history should matter.
Customers also don't behave in perfectly predictable ways.
- They buy gifts.
- They travel.
- They use a new device.
- They ship to a hotel.
- They change cards.
- They send something to a family member.
Those things may look unusual to a rules engine while making complete sense in real life.
The challenge is distinguishing ordinary change from genuinely meaningful change.
And trust shouldn't be permanent.
If a longtime customer suddenly changes their password, email address, payment method and shipping destination at roughly the same time, the history that previously reduced friction may need to be discounted.
Key Insight
Trust should reduce friction, not eliminate verification.
Measuring the customers you stopped
The cleanest way to estimate the counterfactual.
This brings us back to the original problem.
If an order is stopped, how do you know whether stopping it was correct?
You usually don't.
The cleanest way to estimate the counterfactual is to randomly release a small, capped sample of eligible, lower-exposure orders that would otherwise be held.
Then measure what actually happens.
Maybe nearly all of them complete normally.
That suggests your current controls may be creating unnecessary friction.
Or maybe a meaningful share becomes confirmed loss.
Now you have evidence that the intervention is doing valuable work.
Either result teaches you something.
The important part is that the release has to be controlled.
- Low exposure.
- Randomized.
- Capped.
- Disabled during known attacks or unusual conditions.
The point isn't to “allow fraud.”
The point is to measure the cost of a decision that would otherwise never be tested.
That idea is uncomfortable because businesses are trained to celebrate prevented loss.
But preventing every possible loss can itself become expensive.
Key Insight
The point isn't to allow fraud. The point is to measure the cost of a decision that would otherwise never be tested.
Sometimes the biggest risk isn't the model
The workflow around the decision matters as much as the decision.
Some of the most important vulnerabilities I've seen aren't sophisticated.
They're operational.
Imagine a fraud system that takes four minutes to evaluate an order.
But the fulfillment workflow releases that order to the warehouse after two.
The model can be excellent.
The package may already be moving.
Another example:
A customer places an order.
The payment clears.
The risk check approves it.
Then the shipping address changes.
If the order isn't reassessed after that change, the original decision may no longer be relevant.
The attacker didn't have to beat the model.
They only had to wait until after approval.
This is why commerce risk can't be designed only around scoring.
The workflow around the decision matters just as much as the decision itself.
- When does fulfillment release the order?
- What happens after an address change?
- What evidence is retained?
- Which changes trigger another review?
- Who can override the decision?
- What happens after that override?
None of this is glamorous.
But it's where a lot of risk actually lives.
Key Insight
The workflow around the decision matters as much as the decision itself.
Evidence should already exist
Saying “we're legitimate” isn't enough.
Marketplace operations taught me another lesson that carries directly into fraud and chargebacks:
Evidence matters.
When something gets challenged, saying “we're legitimate” isn't enough.
You need records.
- Order history.
- Delivery information.
- Product data.
- Account history.
- Communication.
- Device or session information where appropriate and permitted.
- Changes made after purchase.
If a chargeback arrives, the merchant shouldn't have to reconstruct the transaction manually across five different systems.
The evidence package should already be available.
That's one of the first areas I'd automate.
Before giving a system broad authority to cancel transactions, I'd rather see it help merchants recover losses they're already entitled to dispute.
The downside is lower.
The return is easier to measure.
And it builds the data foundation needed for more sophisticated decisions later.
Key Insight
Before giving a system broad authority to cancel transactions, help merchants recover losses they're already entitled to dispute.
Merchants should control the dial
These are economic constraints, not software settings.
Another problem with traditional risk tooling is that merchants often configure the system in technical language.
- Thresholds.
- Scores.
- Rules.
- “Conservative.”
- “Aggressive.”
But those aren't really the questions leadership cares about.
The business question is closer to:
How much loss are we willing to accept, and how much friction are we willing to impose on legitimate customers to achieve it?
Those are economic constraints.
The system's job should be finding the best operating point between them.
Sometimes both targets won't be achievable.
That's useful information too.
I'd much rather see:
We cannot reduce expected monthly losses below this level without pushing legitimate-order review above your stated limit.
than:
Risk threshold: 72.
One tells an operator something they can make a business decision about.
The other mostly describes the software.
Key Insight
The system's job is finding the best operating point between how much loss you'll accept and how much friction you'll impose.
Explain the decision
Show your work.
I've spent a lot of time recently thinking about explainable commerce systems.
I keep coming back to a simple rule:
Show your work.
If an order gets held, tell me why.
If it gets released, tell me why.
If signature confirmation is the lowest-cost intervention, show why.
If cancelling an order is expected to cost more than accepting the remaining risk, show that too.
Because eventually something will get through.
A fraudulent order will ship.
Someone will ask:
Why did we allow this?
If the answer is simply:
“The model approved it.”
trust disappears very quickly.
And it probably should.
A merchant doesn't need to see every internal calculation.
But they should be able to understand the evidence, the available actions and the tradeoff the system made.
Key Insight
A merchant doesn't need every internal calculation. But they should understand the evidence, the available actions, and the tradeoff the system made.
Prove it before it acts
Automate last.
AI makes it incredibly easy to build something that produces an answer.
That doesn't mean it should immediately be given permission to act on a customer.
I'd roll commerce risk out in stages.
First, measure.
Understand what the business is actually losing today across payment fraud, delivery claims, returns and promotional abuse.
Put those losses in context against sales and operational cost.
Sometimes the right answer may be that the problem isn't large enough to justify another complex system.
That's a valid result.
Then observe.
Run the decision logic silently.
- What would it have recommended?
- What actually happened?
- Where was it right?
- Where was it wrong?
Then introduce human-reviewed actions.
- Hold.
- Signature.
- Verification.
- Evidence request.
- Account restriction.
Automate last.
Automate the narrow categories where the outcomes have been measured well enough to understand both the benefit and the downside.
That's how I'd approach almost any system that can interfere directly with a legitimate customer transaction.
Key Insight
Measure. Observe. Human-reviewed actions. Automate last.
The number I want to know
After years working across direct ecommerce, Amazon, international marketplaces, Shopify, fulfillment, returns and customer operations, my view of commerce risk has become pretty simple.
The goal isn't to catch as many fraudsters as possible.
It's to make the best economic decision on each transaction while creating as little unnecessary friction as possible for legitimate customers.
And then measure whether the decision was actually right.
A merchant shouldn't only be able to say:
“We stopped $200,000 in fraud.”
They should also be able to answer:
“How many good customers did we stop with it?”
How many good customers did we stop with it?
That's the number I want to know.
Compare notes
I'm looking to compare notes with Shopify operators managing payment fraud, delivery claims, returns or promotional abuse. If you're open to discussing anonymized workflows and loss patterns, get in touch through LogicShop.io.
Get in touchViews are my own. Examples have been generalized where appropriate to protect confidential information, and all numerical examples are illustrative rather than production results.