Why merchants stop relying on a single provider
A multi-provider setup means a merchant contracts with several acquirers to process sales from the same catalog. It shows up as soon as the business crosses its second border. A seller that stays in one country can collect everything through a single provider. The first reason to add providers is coverage, meaning the set of payment methods a provider can actually process. No acquirer connects equally deeply to Pix in Brazil, UPI in India, QRIS in Indonesia, BLIK in Poland, and M-PESA in Kenya. A catalog advertising 200 payment methods covers many of them through resale agreements with third parties, not through a direct connection to the rail.
Three other reasons follow, in this order. The approval rate varies from one provider to the next on identical traffic, because the issuer sees a different presenter and different data. Cost varies because domestic acquiring avoids cross-border fees and currency conversion. Resilience comes last, and it is the only one of the four that can be measured only after an incident. Until a route actually goes down, the merchant has no measure of how its backup setup behaves.
| Market | What customers expect | Rail operator | What card-only acquiring doesn't cover |
|---|---|---|---|
| Brazil | Pix, then credit cards in parcelado (installments) | Banco Central do Brasil, via the SPI infrastructure (2020) | Pix has no authorization and no chargebacks: the life cycle is completely different |
| India | UPI, and RuPay on the issuing side | NPCI (UPI; RuPay since 2012) | The payer's app belongs to a licensed third party, outside the merchant's control |
| Indonesia | QRIS, bank virtual accounts, wallets | Bank Indonesia with ASPI (QRIS, 2019) | The central bank mandates the QR standard; acquiring stays local |
| Poland | BLIK | Polski Standard Płatności (2015) | A six-digit code generated in the banking app, with no card involved |
| Netherlands | iDEAL, migrating to Wero | Currence iDEAL B.V., a subsidiary of EPI Company (2005) | Scheduled for shutdown on December 31, 2027: the integration has an expiration date |
| Spain | Bizum | Sociedad de Procedimientos de Pago S.L. (2016) | 105.6 million online purchases in 2025, up 82.1% (Bizum, January 2026) |
| Kenya | M-PESA | Safaricom (2007) | A nonbank rail: no IBAN, no card scheme, no conventional interbank clearing |
| Mexico | Cards, then deferred cash payment | FEMSA for OXXO Pay, Paynet network (2015) | The customer pays later at a convenience store: the order waits for an event, not a response |
Anatomy of an orchestration layer
An orchestration layer is a software intermediary between the merchant's checkout and its payment providers. It gives the merchant a single interface and translates each request for the selected rail. Routing is only one of its six functions. Its core deliverable is the abstraction contract it offers the merchant: a single payment object, a single life cycle, and a single set of statuses, whatever rail sits underneath. Every other function flows from that promise, which is hard to keep because each rail has its own life cycle.
- Method catalog: what exists, where, in which currency, through which connector, and what is unavailable at the moment of display;
- Credential vault: encrypted cards, network tokens, direct debit mandates, wallet aliases, all held outside the providers;
- Rules engine: routing, cascading, exemption requests, per-method limits, all editable without redeploying the application;
- Normalization: an object model and a status taxonomy shared by every rail, with the original code always attached;
- Observability: a log of every route considered, not just the one selected;
- Reconciliation: aggregating heterogeneous settlement files and matching them all the way to the bank payout received.
Routing by cost and by approval rate
Routing picks, for each transaction, the combination of provider and parameters that maximizes an objective function the merchant defines. That function is expected net margin: the probability of approval times the order amount, minus the cost of the route and the expected fraud loss. Approval rate is only one of its three terms. Two routes with a 92% approval rate do not produce the same margin if one costs 30 basis points more than the other. A route whose approval rate climbs two points because it lets fraud through destroys value, since dispute fees come on top of the lost order amount. Managing to approval rate alone leads to bad decisions.
| Rail family | What the merchant chooses | What it doesn't choose | Examples |
|---|---|---|---|
| Cards (pull payments) | Acquirer, acquiring entity and country, brand presented on a co-badged card, exemption requested | The issuer's decision, and the data it uses to score the transaction | Visa, Mastercard, RuPay (NPCI, 2012), Elo (Elo Serviços S.A., 2011), Verve (Verve International, an Interswitch subsidiary, 2009) |
| Instant payments (push) | The provider that generates the request or QR code, the receiving account, the validity period | The payment path: the payer pushes from their banking app | Pix (BCB, 2020), UPI (NPCI), PromptPay (National ITMX, 2017), DuitNow (PayNet, 2018), PayNow (Association of Banks in Singapore, 2017) |
| Direct debits and mandates | The presenting creditor, the collection date, the applicable rulebook | Returns, which arrive after the fact, for weeks | SEPA Direct Debit Core and B2B (European Payments Council, 2009), DuitNow AutoDebit (PayNet), PayTo (NPP Australia, 2022) |
| Wallet | The aggregator holding the contract, and the channel: app, QR code, or web redirect | The wallet's internal rules, opaque by design | Alipay (Ant Group, 2004), WeChat Pay / Tenpay (Tencent, 2005), GCash (G-Xchange, 2004), Mercado Pago (MercadoLibre, 2004) |
| Interoperable QR | The acceptance provider and the enrolled merchant ID | The code standard, set by the central bank or the national operator | QRIS (Bank Indonesia and ASPI, 2019), DuitNow QR (PayNet, 2019) |
| Deferred cash and vouchers | The network of payment locations and the code's validity period | When the customer pays, which is entirely up to them | OXXO Pay and Paynet (2015), Boleto Bancário (Nuclea, 1993) |
# E[margin] = P(approval | route, bin, amount) x (order_amount - route_cost)
# - P(net fraud | route, signals) x (order_amount + dispute_fees)
candidates = []
for route in eligible_routes(method, currency, issuer_country):
if not route.health_ok: # circuit breaker: technical errors
continue # above threshold over 5 min -> route removed
p_ok = model.approval_probability(route, bin, amount, local_hour)
cost = route.interchange + route.scheme_fees + route.markup + route.fx
p_frd = model.fraud_probability(route, device_signals, account_signals)
score = p_ok * (order_amount - cost) - p_frd * (order_amount + dispute_fees)
candidates.append((score, route))
chosen = max(candidates)[1]
log(candidates) # ALL routes, not just the chosen one
# Without a log of the discarded routes and their scores, no counterfactual
# can be computed later: you will never know what the other path would have
# returned, and the routing model can never be honestly re-evaluated.Three configuration decisions deliver most of the gain, before any statistical model comes into play. The first is to acquire in the country where the card was issued. The second is to settle in the cardholder's currency. The third is to present the domestic brand when the card carries one. All three are matters of contract and configuration, not machine learning. A model only comes in afterward to break ties between routes that are already comparable: it fine-tunes a decision made elsewhere and is not a first-order lever.
Failover, cascading, and deferred retries
Failover, cascading, and deferred retries sound alike but are three distinct responses to a failed payment attempt. Failover responds to an unavailable route and switches the attempt to a backup route. Cascading responds to an issuer decline and immediately re-presents the transaction through a second provider. A deferred retry responds to a decline with an economic cause, such as insufficient funds: the next attempt is scheduled for a later date, since the cause may go away once the account is funded. Mixing them up leads to two mirror-image mistakes. The merchant either abandons sales that would have gone through after a wait, or piles up attempts that trigger the card networks' penalties.
| Type of failure | Typical signal | Appropriate response | Classic mistake |
|---|---|---|---|
| Route unavailable | Timeout, server error, card response code 91 or 96 | Immediate failover to the secondary route, with an idempotency key | Hammering the failed route with retries |
| Authentication required | Code 65 at Mastercard, 1A at Visa | Replay the identical transaction with 3-D Secure | Treating the code as a hard decline and dropping the sale |
| Insufficient funds | Card code 51, or AM04 on a SEPA direct debit | Deferred retry, timed to a payroll cycle or a day of the month | Retrying within the minute, with nothing changed on the payer's side |
| Expired or closed instrument | Card code 54, or AC04 for a closed account | Update the credential (Visa Account Updater, Mastercard Automatic Billing Updater), then retry | Re-presenting the same card number without refreshing it |
| Hard decline | Codes 41, 43, 59; Merchant Advice Code 03 or 21 | Stop, and flag the credential as unusable | Cascading to another acquirer in hopes of a different answer |
| Push payment not completed | The customer didn't scan, or the code expired | Generate a fresh request with a fresh ID | Resending the original request, at the risk of a real double settlement |
- Count attempts per card and per merchant, across all providers: network limits apply across acquirers, never acquirer by acquirer. Visa caps attempts at 15 per card over a rolling 30 days;
- Read Mastercard's Merchant Advice Code before deciding:
01means updated account information is available,02allows a later retry, and03and21prohibit one; - Log the reason for every attempt, not just its outcome; otherwise no retry rule can ever be evaluated after the fact;
- Limit cascading to one retry. Beyond that, the added latency and authorization fees exceed the amount recovered;
- Separate technical retries from commercial dunning: a failed subscription payment calls for a re-presentment schedule, not an immediate retry.
Normalizing decline codes
Decline code normalization maps the failure reasons each rail returns onto a single reference set inside the orchestration layer. It exists because the vocabularies in use have nothing in common. Cards respond with two characters in the DE39 field of an ISO 8583 message, while European direct debits are rejected with a four-character ISO 20022 code such as AM04 or AC06. UPI returns NPCI response codes; Pix returns reason codes carried in ISO 20022 messages overseen by the Banco Central do Brasil; an Asian wallet returns its own proprietary labels. None of these vocabularies maps exactly onto another.
Normalization maps all these codes onto a small and stable set of statuses. Small, because product teams will never apply a 40-status taxonomy correctly. Stable, because retry rules, dashboards, and service-level commitments will depend on it for years. The original code is always kept.
| Rail | Where the status comes from | Sample codes | What the code doesn't tell you |
|---|---|---|---|
| Card | ISO 8583, field DE39, supplemented by the network's private fields | 00, 05, 51, 54, 41, 65, 1A | The real reason for the decline: it stays inside the issuer's decision engine |
| SEPA direct debits and credit transfers | ISO 20022, reject or return reason code | AC04 account closed, AC06 account blocked, AG01 transaction forbidden, AM04 insufficient funds, MD01 no mandate | When the return will arrive: refund rights last up to eight weeks on an SDD Core |
| Domestic instant transfer | ISO 20022 messages defined by the national operator | Reason codes published by the central bank or the rail operator | Why the payer gave up, when the request expires without ever being scanned |
| UPI | NPCI response codes | Codes specific to the switch and participating banks | Which link failed: the third-party app, the payer's bank, or the payee's bank |
| Wallet | The operator's proprietary API | Labels defined unilaterally, with no public reference | Internal risk rules, which account for most declines |
| Deferred cash | No event, then expiry | No decline code, just a timeout | Whether the customer gave up or plans to pay after the deadline |
RETRY_NOW The route failed, not the transaction.
card 91 / 96 | ISO 20022 technical reject | UPI switch failure
-> immediate failover to another route, with idempotency
RETRY_LATER Economic decline that can clear over time.
card 51 | SEPA AM04 | insufficient mobile money balance
-> re-presentment schedule, never an immediate retry
NEED_AUTH Authentication required; transaction can be replayed as is.
card 65 (Mastercard) / 1A (Visa) | A2A PIN or biometrics redone
-> replay with 3-D Secure, on the same route
NEED_CREDENTIAL The instrument must change before any new attempt.
card 54 | SEPA AC04 | wallet account deactivated
-> Account Updater, or ask the customer for a new payment method
TERMINAL Never replay, whatever the route.
card 41 / 43 / 59 | Merchant Advice Code 03 and 21 | SEPA AC06, AG01
-> flag the credential, stop the retry cycle
UNKNOWN Unmapped. Technical debt, not a working category.
-> measure its share per connector; above a few percent,
the layer is flattening codes, not normalizing them
Fields kept alongside the canonical status, without exception:
raw_code, raw_network, raw_message, raw_advice_code, route_id, attempt_idAlongside the decline code, the card networks send two pieces of information that carry an instruction rather than a verdict. Mastercard's Merchant Advice Code says whether to retry the transaction now, retry it later, or never try again. Credential update services (Visa Account Updater, Mastercard Automatic Billing Updater) answer a different question: does the cardholder now have a different card? An orchestration layer that ignores these two channels bases its retry rules on the decline code alone, even though the networks have already supplied the answer.
UNKNOWN status. A wrong canonical status is riskier than an unreadable raw code, because it triggers an automatic retry rule on a false premise. The check works in reverse: draw a sample of canonical statuses, match each line back to its original code, and verify the mapping on that sample. Without this check, no one knows what share of statuses is mapped correctly, and the taxonomy stops reflecting why payments are actually declined.Orchestration vendors
The orchestration market took shape in four waves, the first of which predates the term itself. The 2000s saw the rise of gateways of gateways, before the category had a name. In the late 2010s, PSPs absorbed the first independents. A third generation was founded between 2020 and 2022 by former executives of PayPal, Braintree, Rappi, and Delivery Hero. The fourth wave, now under way, comes from the PSPs themselves, which are building optimization into their own stacks.
| Model | Who decides the route | What you gain | What you pay | Companies cited |
|---|---|---|---|---|
| Single provider | The PSP, under its own rules | One contract, one reconciliation, no extra integration | Routing imposed on you; coverage limited to the provider's catalog | – |
| Independent orchestrator | The merchant, in a rules console | Claimed neutrality, provider-agnostic token vault, provider comparison on real data | Subscription plus per-transaction fees; one more intermediary on the critical path | Primer, Gr4vy, Payrails, Yuno, IXOPAY, Spreedly, CellPoint Digital |
| PSP-native orchestration | The PSP, with rules exposed to the merchant | Nothing to integrate, and a model trained on the provider's own volume | The router belongs to a party with a stake in the routing outcome | Adyen Uplift (announced January 9, 2025) |
| Open-source or in-house stack | The merchant, entirely | Full control, no per-transaction fees, auditable code | A dedicated team, PCI scope to carry, technical debt to own | Hyperswitch (Juspay, Apache 2.0 license) |
Inside or outside the flow of funds: the question that decides regulatory status
An orchestration layer's legal status depends on a single question: does it take possession of the funds? A layer outside the flow routes messages without ever holding money. Funds move from the payer to the acquirer, then from the acquirer to the merchant, and the layer remains a technical service provider. A layer inside the flow collects on the merchant's behalf, holds the funds temporarily, and then pays them out. From that point on, it falls under the rules for regulated institutions, country by country.
The switch from one regime to the other often happens without an explicit decision. A merchant asks for a consolidated payout, a marketplace asks for funds to be split among sellers, a jurisdiction requires collection through a local entity. A layer that agrees to any of these requests ends up receiving funds, and that changes its business: it goes from technical service provider to regulated institution. In several markets, exchange controls force this shift. Payments there are collected in local currency by an in-country entity, while repatriation in hard currency falls under an entirely separate contract.
| Obligation | Trigger | What it requires | Reference |
|---|---|---|---|
| Payment license in the European Union | The layer collects or holds funds on behalf of merchants | Payment institution or e-money institution license, safeguarding of funds, capital requirements, passporting to operate outside the home country | Directive (EU) 2015/2366 (PSD2) |
| Payment aggregator authorization in India | The layer collects on behalf of Indian merchants | Application filed on the PRAVAAH portal, net worth of ₹15 crore at application and ₹25 crore within three years, escrow account with a scheduled commercial bank | Reserve Bank of India, Regulation of Payment Aggregators Directions, 2025, issued September 15, 2025 |
| ICT third-party risk management | The layer sits on the critical path of a European financial entity | Entry in the register of information, mandatory contract clauses, documented exit strategy, resilience testing | Regulation (EU) 2022/2554 (DORA); 19 critical third-party providers designated by the European Supervisory Authorities on November 18, 2025 |
| PCI DSS compliance | The layer sees, transmits, or stores card data | Compliance scope shifts to the layer; an attestation to obtain and renew; management of scripts on the payment page | PCI DSS, PCI Security Standards Council |
- Credential ownership: under whose Token Requestor ID are the network tokens provisioned, the merchant's or the provider's? The answer determines how reversible the setup really is;
- Export: the format, turnaround time, and cost of a full vault extraction, tested once before you need it;
- Bypass route: the contractual and technical ability to call a provider directly, without going through the layer;
- Logs: access to raw codes, discarded routes, and their scores, exportable without relying on the vendor's interface;
- Subcontracting: a list of hosting providers and sub-processors, with a right to object and advance notice;
- Exit: length of the transition period, migration support, and what happens to the data after termination, tokens included.
The hidden cost of one more layer
An orchestration layer costs you in three distinct ways. The first is money, charged on every transaction presented, whatever the outcome. The second is latency added to the payment's critical path, since every extra call lengthens authorization time. The third is a new dependency: the layer becomes the single gateway to the very providers the merchant added to reduce its exposure.
The first cost is in the contract and can be negotiated. The other two only show up in use, and fixing them means redoing the integration. An orchestration program that only measures invoiced amounts is managing the least important of the three variables, and leaves latency and dependency out of view.
| Item | What it really costs | How to measure it |
|---|---|---|
| Per-transaction fees | Charged on top of the provider's fees, including on declined transactions | Divide total cost by authorized revenue, never by the number of API calls |
| Latency | An extra network round trip on the critical path, on every attempt and every cascade | Compare 95th-percentile authorization time before and after go-live, by hosting region |
| Reconciliation | As many settlement formats as there are providers, plus the layer's own format | Auto-reconciliation rate and time to monthly close |
| Token vault | Non-portable tokens rebuild the lock-in orchestration was supposed to remove | Share of credentials held under a Token Requestor ID owned by the merchant |
| Correlated outage | When the layer is down, every provider becomes unreachable at once | Whether a built-in bypass route exists, and when it was last tested in production |
| Know-how | Routing rules become an asset no one can justify anymore | Number of active rules, date of last review, named owner |
- Measure the layer's cost against authorized revenue, not call volume: declined transactions incur fees too;
- Measure added latency at the 95th percentile, by hosting region, before and after go-live;
- Track the share of credentials held under the merchant's own Token Requestor ID: it is the direct measure of reversibility;
- Name an owner for the routing rules, date every review, and delete rules no one can explain anymore;
- Run the bypass in production at least twice a year, log the result, and fix whatever failed.