Orchestrating payments across multiple PSPs. 6 chapters and a final quiz.
The day-to-day craft of spreading payment acceptance across several providers, step by step. Work out the volume at which a second PSP pays for itself, write a routing rule that weighs cost and approval rates in the same unit, wire up failover without creating double charges, normalize decline codes into an action table that complies with Visa and Mastercard rules, and schedule retries that stay clear of excessive-attempt fees. Then measure the gain against a control group and put a price on what the orchestration layer really costs, token vault included.
Calculate the volume at which a second provider pays for itself, fixed costs included
Write a routing rule that weighs cost and authorization rate in a common unit
Wire up safe failover with a circuit breaker, idempotency keys, and reversals, with no risk of double authorization
Normalize decline codes into an action table aligned with Visa's categories and Mastercard's Merchant Advice Codes
Chapter 1. Deciding when to go multi-PSP.
A multi-provider setup is one in which a merchant uses several payment service providers at once to accept payments, rather than a single one. Merchants adopt it to remove a specific constraint: coverage of local payment methods, authorization rates, negotiating leverage, or service availability. Each additional provider brings a contract, a settlement account, a reconciliation feed, a dispute queue, and an annual testing cycle. That cost is fixed: it is the same whether the provider processes one euro or ten million. The head of payments' first job is therefore to name the constraint, then to size it against that fixed cost.
Four triggers that justify a second provider
🌍
Missing coverage
The incumbent provider does not offer the payment method that dominates a target market. No authorization rate makes up for a method missing from checkout. This is the most common trigger, and the easiest to prove.
📉
Slipping approval rates
The authorization rate collapses on a specific corridor, defined by issuing country, card type, and amount. A local acquirer makes the transaction domestic in the issuer's eyes, and no configuration setting can substitute for that.
💶
No leverage
Providers don't adjust their pricing for a captive merchant. Volume that can actually be moved is the only argument that reopens a negotiation, and only if it can be moved within days.
🛡️
Single point of failure
When a sole provider goes down, payment acceptance stops. Size this risk by the number of hours the company can tolerate collecting nothing, not by the probability of an outage.
79.8B
Pix transactions in Brazil in 2025; 54.7% of retail transactions in the second half
Banco Central do Brasil, 2025–2026
2.9B
BLIK transactions in Poland in 2025, up 21% year over year
Polski Standard Płatności, February 2026
≈ 62 %
of Dutch online spending paid with iDEAL; scheduled to be phased out on December 31, 2027, in favor of Wero
EPI Company, 2026
Architecture
What it adds
What it costs
When to choose it
Single provider
One integration, one reconciliation, one point of contact when something breaks
No failover, no negotiating leverage, local methods limited to the provider's catalog
Modest volume, one or two markets, small payments team
Two providers managed in-house
Failover and routing decisions under control; the merchant owns the routing data
Two integrations, two token vaults, and routing logic to write and maintain over time
An established payments team, and volume that pays back the project in under 12 months
Orchestration platform
A single API, prebuilt connectors, a shared vault, configurable routing rules
An extra per-transaction fee; the dependency moves rather than disappears
Many markets to launch quickly, few developers available on the merchant side
Independent vault and direct connections
Portable card data, full routing freedom, provider-by-provider negotiation
A compliance scope to maintain, multiple integrations, operations run by the merchant
Very high volume, and a strategic requirement to be able to switch providers
Four acceptance architectures, four different bills
Players you will meet in an RFPAdyenStripeWOWorldpayDLdLocalEBEBANXYUYuno
⚠️
When volume splits, so do the pricing tiers
Acquiring rates are tiered by volume, so splitting traffic across two providers means losing the tier reached with the first. Unit costs then rise on both sides. Standard practice is to negotiate on the company's consolidated volume, with a minimum commitment and an annual review clause in each contract. Without these clauses, the merchant's effective rate goes up from the very first statement, before the second provider has delivered any gain in approvals.
Expected approval gain: annual volume of the target cohort × observed gap in authorization rate × average order value × margin rate
Fee savings: volume that can actually be moved × gap in effective rate, including FX and scheme fees
Annual fixed cost: integration, testing, extra reconciliation, support, compliance audit, management time
Decision: a second provider is justified when the first two items exceed the next two over 12 months, using figures measured in your own business, never figures from a sales deck
🎯 Quick question
A merchant splits its traffic evenly between two acquirers and sees its effective rate rise at both. What should it check first?
Chapter 2. Routing by cost and approval rate.
Routing is the choice, for each payment attempt, of the path through which it is submitted. A route is the combination of provider, acquiring entity and its country, network or rail, and requested authentication method. Two routes at the same provider can produce opposite results, because the issuer sees a different transaction depending on the acquiring country and the authentication method. Routing that thinks in providers rather than routes leaves most of the available gain on the table.
Routing pursues two conflicting goals: lowering variable cost and raising the authorization rate. The cheapest route is not the one with the most approvals, and the route with the most approvals often costs more. To trade one against the other, you need to express both in the same unit: net margin per attempt.
🔑
One point of authorization vs. a tenth of a point in fees
A route's value is p × (order value × margin − variable cost), where p is its authorization probability measured on the cohort. On an €80 order, one point of authorization gained at a 25% margin earns €0.20 per attempt; a tenth of a point saved in fees earns €0.08. The balance tips at a margin rate of exactly 10%, whatever the order value. Above it, approvals come first; below it, fees win. The threshold therefore follows from the business's margin rate, and you set it before writing any routing rules.
Segment before you route
A cohort is a homogeneous subset of payment attempts. It is defined by issuer country, product type, amount band, channel, currency, presence of a network token, and authentication result. An authorization rate computed across all traffic lumps together cohorts that behave differently, and it hides exactly the gaps routing acts on. Segmentation balances two opposing constraints. A cohort that is too broad hides the gap you are looking for; one that is too narrow no longer has enough observations to be reliable. Standard practice is to set a minimum number of observations per cohort and to fold any cohort below it back into its parent.
Criterion
Data source
Refresh rate
Pitfall
Authorization rate by cohort
Attempt log, normalized network response code, authentication result
Weekly recalculation over a rolling window
A cohort below its minimum observation count yields noise, not signal
Variable cost of the route
Detailed acquiring statements, scheme fees, FX fees, decline fees
Monthly, when statements arrive
The contract rate is never what you are billed: only the statement counts
Route availability
Technical probes, transport error rate, observed latency
Real time
An issuer decline is not an outage: confusing the two triggers needless failovers
Local regulatory constraints
Market-by-market map, maintained with the legal team
Whenever a rule changes
A local legal requirement overrides the economic optimum, no exceptions
Settlement capability
Payout currency, payout timing, destination bank account
Quarterly
Routing to a provider that doesn't settle in the right currency brings back an FX conversion
What a routing rule needs, and where to find it
An explicit routing rule, versioned with the code
routes:
- id: eu-domestic-debit
if:
issuer_country: [de, es, fr, it, nl]
product: debit
max_amount_cents: 25000
to: psp_a
acquiring: local # domestic transaction for the issuer
token: network
- id: cross-border-credit
if:
issuer_country: ["*"]
product: credit
to: psp_b
authentication: exemption_requested
default: psp_a
fallback:
triggers: [timeout, transport_error, breaker_open]
to: psp_b
never_on: [issuer_decline] # a decline is not an outage
max_failovers: 1
What local law requires of routing
United States. The Federal Reserve Board's Regulation II requires every debit card to be enabled on at least two unaffiliated networks, and the merchant chooses which one to use. The final rule of October 3, 2022, extended the requirement to card-not-present transactions, effective July 1, 2023. In the US, debit routing is a right, not a favor granted by the acquirer.
European Economic Area. Article 8 of Regulation (EU) 2015/751 lets the merchant set a priority brand on a co-badged card, but never in a way that stops the payer from overriding it. A routing rule that locks in the cardholder's choice is non-compliant.
India. Since October 1, 2022, only the issuer and the network may store card data, under a Reserve Bank of India mandate. Merchants handle only tokens, each specific to one token requestor. Routing is constrained from enrollment onward.
Brazil. Pix does not run on card rails. There is no issuer authorization, no decline code to normalize, and no retry: the order is either paid or it fails. Card routing and Pix routing are two separate machines, not two branches of the same rule.
🎯 Quick question
A merchant has an average order value of €80 and an 8% margin. Route B approves one point more than route A but costs a tenth of a point more in fees. What does the math say?
Chapter 3. Failing over when a provider goes down.
Failover is the automatic rerouting of a payment attempt to another route when the original route has not responded. Triggering it correctly depends on telling apart two events that systems often, and wrongly, treat the same way. An issuer decline is a response: the message reached the cardholder's bank, and the bank ruled against the transaction. No response, by contrast, leaves the authorization status unknown, and failover applies only to this second case. Confusing the two causes needless failovers, double authorizations, and, on card networks, excessive-attempt fees.
One attempt and its outcomes, with controlled failover
Orchestrator
Sends the attempt on the selected route
Stable order reference; idempotency key specific to this route and attempt number
➜
Provider A
Responds within the timeout
Approval or decline: the response is logged with its raw, untranslated network code
➜
Orchestrator
Receives nothing before the timeout
The authorization status is unknown, which is not the same as failed
➜
Orchestrator
Sends a reversal to provider A
Cancels any authorization that may have been approved and releases the hold on the cardholder's funds
➜
Orchestrator
Replays the attempt at provider B
New idempotency key, same order reference, failover counter incremented
➜
Reconciliation
Reconciles both logs the next day
Any orphan authorization is voided or refunded before it expires
The three signals that allow a failover
Transport error: a 5xx response, a failed TLS handshake, a DNS resolution failure. The message never reached its destination, or the response was lost.
Timeout: no response within the window set for this route. The outcome is unknown and must be resolved with a reversal.
Error rate above threshold: over a rolling window, the share of the two signals above exceeds the configured value. The circuit breaker opens and the route is skipped with no further attempts.
No issuer decline code appears on this list. That is the most useful rule in this chapter.
A circuit breaker is a software mechanism that suspends a route as soon as its error rate crosses a configured threshold. It has three states, and its settings govern the transitions between them. When closed, attempts go through normally. When open, the route is skipped without being called, which avoids piling up timeouts during an outage. When half-open, only a sample of attempts is let through, to check that the service has recovered. Four settings define it: the error threshold, the length of the observation window, how long it stays open, and the size of the recovery sample. Document these four values and replay them in testing; otherwise you will have no way to explain a failover that happened six months earlier.
Failover: what the function must do, and what it must never do
# One attempt = one idempotency key, specific to the route
key = fingerprint(order_reference, route.id, attempt_number)
response = route.psp.authorize(amount, token, key, timeout = route.max_timeout)
match response:
APPROVED -> capture; log the raw network code
ISSUER_DECLINE -> DO NOT fail over; apply the retry policy
TIMEOUT | ERROR_5xx -> reverse(route, key) # outcome unknown
if failovers < max_failovers:
route = next_route() # new key
start over
BREAKER_OPEN -> route = next_route() without calling the open route
# Explicitly forbidden
# · reusing the same idempotency key at another provider
# · failing over after an issuer decline
# · marking a timed-out attempt as "failed" without sending the reversal
⚠️
Double charges are created at the moment of timeout
When the timeout expires, the authorization may have been approved without the response reaching the orchestrator. Replaying the attempt elsewhere without a reversal then creates two holds on the cardholder's account, two settlement lines, and an almost certain dispute. The idempotency key protects against a replay at the same provider, which recorded it. Another provider never received it and cannot recognize it. Only two mechanisms close this risk: the reversal sent to the timed-out route, and the next day's reconciliation, which matches the logs of both providers.
Switching providers does not change the issuer, which is still the bank that issued the card. A card declined at provider A is declined at provider B, with the same response code. The decision comes from that bank, not from the acquiring chain. Replaying a decline on another route only adds an attempt to the network's counter and brings the merchant closer to the excessive-attempt thresholds, without raising the chance of approval. Failover handles outages; the retry policy handles declines.
🎯 Quick question
An authorization is sent to provider A, and no response comes back before the timeout. What is the correct sequence?
Chapter 4. Normalizing decline codes to drive retries.
Normalizing decline codes means mapping the labels that different parties give to the same authorization decline onto a single vocabulary. A decline reaches the merchant in three overlapping forms. The network returns an authorization response code, to which Mastercard adds a Merchant Advice Code intended for the merchant rather than the acquirer. The provider then layers its own label on top, often translated and sometimes stripped of detail. The same decline therefore has three names, depending on where you read it. No retry policy can be applied until those three vocabularies have been reduced to one.
Visa's four response categories
Category 1: the issuer will never approve. No further attempts are allowed. Code 14 (invalid account number) is in this category: the merchant must never reattempt on the same number.
Category 2: the issuer cannot approve at this time. At most 15 reattempts over 30 days. Insufficient funds falls into this category; the decline is temporary by nature.
Category 3: data quality. The data sent is wrong or incomplete. Reattempting without fixing it achieves nothing and feeds the monitoring counters.
Category 4: generic codes. Same cap as category 2: 15 reattempts over 30 days. The issuer gave no reason for the decline, which makes retries less predictable.
Sources: Visa Rules, article AI10325 of September 3, 2020, effective April 17, 2021; Visa Merchant Business News Digest for the April 13, 2024, reclassification.
Mastercard's Merchant Advice Codes
Code
What the issuer is saying
What the orchestrator should do
01
New account information is available
Use an account updater service, then retry only once
02
Temporary decline, try again later
Schedule a deferred retry, within the network's attempt cap
03
Do not try again
Permanent stop on this payment method; offer the customer another method
04
Token requirements not met for this token type
Fix the enrollment or the cryptogram before any new attempt
21
Payment canceled
Permanent stop; the merchant closes the related subscription
24 to 30
Retry after 1 hour, 24 hours, or 2, 4, 6, 8, or 10 days
Use exactly the wait time given: it replaces the in-house retry schedule
41
Single-use virtual card number
Ask for a new payment method; retrying is pointless
Mastercard Merchant Advice Codes, as listed in integration documentation, 2026
Normalization table: one internal family, one action, one cap
{
"insufficient_funds": {
"network": ["51"], "mac": ["02", "24", "25"],
"action": "retry", "delay_hours": 24,
"max_attempts": 15, "window_days": 30
},
"account_not_found": {
"network": ["14"], "mac": ["03"],
"action": "stop", "max_attempts": 0,
"note": "Visa category 1: never reattempt on the same number"
},
"card_lost_or_stolen": {
"network": ["41", "43"], "mac": ["03"],
"action": "stop", "max_attempts": 0,
"escalate": "fraud_review"
},
"data_needs_update": {
"network": ["54"], "mac": ["01"],
"action": "update_then_retry", "max_attempts": 1
}
}
# Check these mappings network by network, against the current
# specifications. This excerpt shows the shape, not the truth.
April 2020
Visa groups response codes into four categories
Issuers must return a descriptive code. Acquirers and merchants base their reattempts on the category, not the label.
April 17, 2021
The reattempt rules take effect
Category 1 bars any further attempts; category 2 allows at most 15 over 30 days. Codes 03, 62, 78, and 93 move from category 1 to category 2 (Visa Rules, article AI10325 of September 3, 2020).
April 13, 2024
Three codes reclassified
Codes 39, 52, and 53 move from category 4 to category 2, code 14 stays in category 1, and code Z5 is created in category 2 (Visa, Merchant Business News Digest).
February 1, 2026
Resubmitting a decline gets much more expensive
The Mastercard Card-Not-Present Advice Decline Fee rises from $0.05 to $0.78 per resubmission of a decline on the same card within 30 days, according to the schedule TD Merchant Solutions published in November 2025.
⚠️
Excessive-attempt counters never stop running
The networks define excessive attempts as 10 or more declines on the same card in 24 hours, or 20 or more over 30 days; beyond either threshold, the acquirer charges a fee. TD Merchant Solutions' November 2025 schedule lists $0.74 per transaction for the Mastercard Compliance Integrity Fee, and $0.15 domestic and $0.23 cross-border on the Visa side. An orchestrator that retries without tracking these counters per card crosses the thresholds without noticing, and the retry queue becomes a cost center before it becomes a monitoring case.
The normalized table maps each internal decline family to an action the retry system can execute; a family with no action is useless when the time comes to decide on another attempt. Three columns are enough: the action to take, the wait before the next attempt, and the maximum number of attempts allowed. The counter is kept per card and per merchant, not per order, because that is how the networks count. The raw, untranslated decline code stays available in the log, since provider labels change without notice and break any mapping built on them.
🎯 Quick question
A subscription payment fails with Merchant Advice Code 03. The in-house retry schedule calls for three attempts over 10 days. What does the orchestrator do?
Chapter 5. Measuring the real gain.
Measuring a routing gain relies on a counterfactual: an estimate of what the same population of orders would have produced without the change. A team that draws conclusions from a simple before-and-after comparison gets it wrong nine times out of ten. The traffic mix shifts between the two dates. Promotions, seasonality, a market launch, and acquisition campaigns bring in different cardholders, whose issuers behave differently. An authorization rate that rises in the days after a rollout therefore measures all of these shifts combined, not the effect of routing alone.
The denominator almost always lies
The authorization rate is calculated per attempt, and a retry policy generates extra attempts, many of which fail. The ratio then falls mechanically, just as collections improve. The metric to base decisions on is the share of orders ultimately collected, across all attempts. The per-attempt rate is still useful for diagnosing a specific route, but not for managing the whole.
Indicator
Definition
Pitfall
Success rate per order
Orders collected ÷ orders submitted for payment, across all attempts
An order abandoned before the first attempt must come out of the denominator; otherwise you are measuring the checkout funnel
Authorization rate per attempt
Authorizations ÷ authorization attempts
Every added retry pushes this ratio down, even when it brings in money
Variable cost per collected order
Processing fees, scheme fees, FX, decline fees, and orchestration fees, divided by orders collected
Decline and excessive-attempt fees show up on the statement a month later
Dispute rate
Disputes ÷ collected transactions, on the routed cohort
An approval gain won by loosening checks gets paid for two months later
Net margin per order submitted
The net of everything above, per order entering payment
The only number that settles it; the other four explain it
Five metrics, their exact definitions, and what they hide
A protocol that stands up to the finance team
Keep a permanent control group: 5 to 10% of traffic stays on the reference route indefinitely. That is the cost of measurement, and it is far lower than the cost of a wrong decision.
Assign the variant to the order, never to the attempt: otherwise a retry on a control order goes out on the tested route and ruins the comparison.
Cover a full cycle: for a monthly subscription, at least two cycles; otherwise the long retries have not yet played out.
Analyze by cohort: a positive average gain can hide an outright loss on one issuing country, which the next shift in mix will expose.
Set the stopping rule in advance: duration, minimum volume, and the gap that counts as significant. A test stopped on the day it looks favorable measures nothing.
Then work out the net result line by line. From the approval gain, subtract the difference in variable cost, retry and excessive-attempt fees, the cost of the orchestration layer, and the observed impact on disputes. A point of approval won by shifting the mix toward more expensive cards can cost more than it earns. Do the calculation per order submitted, over the same period, using the statements you actually received rather than contract rates.
🔑
A gain that doesn't survive the control group doesn't exist
A routing change rolled out to all traffic at once, without a control group, can no longer be evaluated or defended, because no comparable population remains on the reference route. Keep the control group after the decision: it catches model drift, a provider's pricing changes, and shifts in issuer behavior that nothing else flags.
🎯 Quick question
After a retry policy goes live, the per-attempt authorization rate falls by 3 points while collected revenue rises. Which reading is correct?
Chapter 6. Costing an orchestration layer in full.
An orchestration layer is a software intermediary between the merchant and its providers. It exposes a single interface and distributes attempts across prebuilt connections. Its advertised pricing is a fee on each transaction processed. The full cost includes six other items, several of which only surface when the merchant leaves the platform. Price all seven before signing the contract, not after.
The per-transaction fee, charged on every transaction processed, including failed ones under some contracts.
The token vault: hosting, detokenization calls, and above all re-tokenization if the tokens are not portable.
The compliance scope: depending on the architecture, card data does or does not pass through the merchant's systems, which changes its self-assessment questionnaire.
Added latency: one more network hop in the checkout flow, measured at the 95th percentile, not just on average.
Displaced dependency: an outage at the single provider becomes an outage at the orchestrator, which then controls every route at once.
Multiplied reconciliation: one settlement feed per provider, plus the orchestrator's own records, to be matched every day.
Commitments: minimum volume, fixed term, exit penalty, and billing for modules added mid-contract.
The vault determines whether you can leave
A network token is a reference issued by the network to an identified requestor and restricted in use. That is how EMVCo defines it in its tokenization specification, version 2.4 of which was published on July 9, 2026. A token is restricted to a merchant, a device, or a payment scenario. That sets it apart from an encrypted card number, which carries no such restrictions. The identity of the token requestor therefore determines which entity can still use the tokens after a change of provider.
Regime
Who holds the token
Portability
What to negotiate
Provider vault
The provider, under its own token requestor ID
None without an assisted migration, which the outgoing provider has no incentive to arrange
An encrypted export to a certified third party, with timing, format, and no charge written into the contract
Certified third-party vault
A vault provider independent of the payment providers
Good: the providers connect to the vault, not the other way around
Compliance scope, the third party's own exit terms, cost per stored token
Requestor ID in the merchant's name
The merchant, enrolled separately with each network
Best: the network token survives a change of provider
Network-by-network enrollment, lifecycle management, and the associated operating workload
India, since October 1, 2022
The issuer and the network, and no one else in the chain (Reserve Bank of India)
The token belongs to the requestor: switching providers requires re-tokenization
The cardholder re-consent flow, to be prepared market by market
Three vault models, three degrees of freedom
Clauses to write in before signing
Vault export within 30 days, in a documented format, to a third party of the merchant's choosing, with no exit fee.
Token requestor ID enrolled in the merchant's name where the network allows it, stated explicitly in the contract.
Access to raw decline codes, untranslated, in the API and in exports: without them, the normalization table cannot be verified.
Exportable attempt log, timestamped, with the route ID: the raw material for any later measurement.
Service level commitment on availability and 95th-percentile latency, with penalties and an incident procedure.
No routing exclusivity: the contract must not prohibit connecting a competing provider directly.
🔑
Exit costs are negotiated on the way in
A merchant that cannot export its vault, retrieve its logs, and move its tokens still has a single-provider architecture, with an intermediary added on top. The RFP should therefore focus on the technical and contractual terms of an exit 18 months out. A clear written answer, incorporated into the contract, says more about how easily you can actually leave than any product demo.
🎯 Quick question
A merchant wants to be able to switch providers without making its subscribers re-enter their cards. Which setup best serves that goal?