What fails when payments fail
A payment outage is an interruption in payment acceptance at any point in the chain that links a merchant’s checkout to the payer’s account. The term describes the visible effect without identifying which layer failed, which is why diagnoses so often conflict. The merchant sees that payments are no longer going through, while its provider sees systems running normally at the same moment. Both are right. They are looking at different segments of the same chain. A payment acceptance chain stacks at least six independent layers, run by different companies, under different contracts, and under supervisory regimes that do not overlap. Each fails for its own reasons, under the responsibility of a party that answers only for its own layer. Nothing alerts one layer when another fails.
The nationwide blackout of April 28, 2025, left Spain and Portugal without power for several hours. The Banco de España and the Banco de Portugal measured its impact, one payment instrument at a time. Clearing and settlement infrastructure kept running. Points of sale, cut off from both power and their local network, went dark. The losses were therefore concentrated at the edge of the chain, where payments are initiated.
The Banco de España’s sector breakdown reveals gaps that the national average hides. Large retailers lost only 35% of their card payments, while small shops saw drops of more than 80% at the worst points, and many chose to close. Rail transport fell 73% and restaurants 63%. The recovery was immediate, and spending even overshot as purchases blocked on the day shifted to the following days. On Tuesday, April 29, card spending was 14% above comparable Tuesdays in April 2024, and on Wednesday, April 30, it was 37% higher (Revista de Estabilidad Financiera No. 49). A provider’s availability rate, measured over a full year and only on that provider’s own systems, captures none of these swings.
| Layer | What makes it fail | What the merchant sees | Workable fallback |
|---|---|---|---|
| Point-of-sale power and network | Power outage, internet service provider failure, cellular network congestion | Unresponsive terminal, frozen register, no response code | Battery backup (UPS), a second SIM from another mobile operator, the terminal’s offline mode |
| Terminal and POS software | Faulty update, expired certificate, uncertified terminals | Every transaction declined on one terminal model, but not on the others | Mixed terminal fleet; backup terminal not updated along with the main fleet |
| Payment service provider and gateway | Application incident, overload, hosting dependency | Response times lengthen, then technical errors pile up | A second provider reachable directly, not through the same orchestration layer |
| Acquirer and national switch | Switch failure, denial-of-service attack, migration | The whole country stops at once, for every merchant | Another payment method; rarely another acquirer |
| Card network or instant payment rail | Failure of a switching component, botched failover to the backup site | One brand declines, the others go through | Routing to the second brand on a co-badged card, or to a separate rail |
| Issuer and identity provider | Core banking system down, national authentication outage | Declines concentrated at one bank, authentication failures on remote purchases | Stand-in processing by the network, postponing the sale |
Documented outages and their real causes
Publicly documented payment incidents form a small body of cases: the episodes that an operator, a regulator, or a parliamentary committee described after the fact. The causes that recur are mundane: a planned change, a component that degrades without failing outright, a certificate that expired because nobody tracked it, or a supplier outside the financial sector. Almost none of these episodes stem from a spectacular attack. The ones that do are denial-of-service attacks, which flood the access channel rather than break into systems. What these episodes share lies in the response rather than the root cause: in almost every case a backup existed, but it had never been exercised under the conditions of the incident.
Five mechanisms come up again and again. Each one is addressed before the incident, through configuration, scheduling, or a failover drill, and none can be fixed while the outage is under way.
- A partial failure defeats failover. Backup architectures are designed for components that die. A component that degrades without stopping keeps responding, keeps control, and blocks automatic recovery. What needs testing is the detection of degradation, not redundancy itself.
- Planned change is the leading cause. Migrations, hardware upgrades, security updates. The maintenance window concentrates the risk, and a change freeze around peak periods protects better than a recovery plan.
- Silent expiry. Certificates, keys, licenses, and software versions expire without any business-side alert. The DCash case cost a central bank two months of downtime.
- The out-of-scope supplier. Telecom operators, hosting providers, security vendors, identity providers. None is supervised as a payment system, yet any of them can stop a country’s payment acceptance.
- The national single point of failure. Where one switch carries all card acceptance, merchant-level redundancy achieves nothing. The fallback has to be another payment method, not another provider.
DORA: resilience becomes an enforceable obligation
The European Union turned operational resilience into a verifiable obligation with Regulation (EU) 2022/2554, known as DORA, which has applied directly since January 17, 2025, without national transposition. Its scope covers banks, payment institutions and e-money institutions, crypto-asset service providers, insurers, asset managers, and market infrastructures. A non-financial merchant falls outside that scope. Its acquirer falls within it and passes the requirements on by contract, although the regulation does not address the merchant directly.
Incident reporting is the most visible part of the regulation. The entity first classifies the incident against harmonized criteria: number of clients affected, duration, data loss, economic impact, and geographic reach. The resulting classification drives everything that follows, and a major incident sets off a cascade of deadlines that the regulation first counts in hours.
The four-hour window opens in the middle of the crisis, while technical teams are busy restoring service. Meeting it requires a reporting channel set up in advance, classification criteria written down internally, and an on-call team able to classify the incident before it is resolved. Under-classifying exposes the entity to a regulatory breach. Over-classifying floods the authority with needless reports, and repeated false alarms dull attention to incidents that really are major.
- Register of information: a complete map of the arrangements with IT providers, submitted to the authorities in a mandated format. It often reveals second- and third-tier subcontracting for the first time.
- Mandatory contract clauses: audit rights, data location, cooperation during incidents, notice periods, and a documented exit strategy. A hosting contract without proven reversibility no longer passes muster.
- Advanced testing (TLPT): threat-led penetration testing on production systems, modeled on the TIBER-EU framework, at least every three years for the most significant entities.
- Direct oversight of critical providers: the European Supervisory Authorities designated 19 critical third-party providers on November 18, 2025. The regulator no longer deals only with the bank; it now deals with the bank’s hosting provider as well.
What central banks and supervisors require
Payment system oversight is the central bank function of assessing the safety and continuity of the arrangements that move money. It applies to the system itself, separately from the prudential supervision of each institution. DORA covers only Europe, and only financial entities. The rest of the world handles the same risk through this oversight and through national prudential regimes. The vocabulary varies widely from one market to another, and the numerical thresholds do not line up across regimes. The subject matter does converge, however: recovery time, testing, and oversight of providers appear in every regime.
The common foundation is the PFMI, the Principles for Financial Market Infrastructures published by CPMI and IOSCO in 2012. Principle 17 covers operational risk. It sets two objectives that most national regimes adopt. The first is recovery of critical systems within two hours of a disruption. The second is completion of settlement by the end of the day, even in extreme circumstances. Secondary sites, drills, and default plans all follow from these two objectives.
| Regime | Reach | What it requires in practice |
|---|---|---|
| PFMI, Principle 17 (CPMI-IOSCO, 2012) | Market infrastructures, including systemically important payment systems | Critical systems recovered within two hours, settlement completed the same day, secondary site, regular testing |
| SIPS Regulation (ECB/2014/28) | Systemically important payment systems in the euro area | Transposes the PFMI into directly binding EU law, under Eurosystem oversight |
| PISA framework (Eurosystem, 2021) | Electronic payment schemes and arrangements, including wallets | Extends oversight beyond systems to card schemes and payment arrangements |
| DORA, Regulation (EU) 2022/2554 | EU financial entities, since January 17, 2025 | Major incident reporting, register of providers, TLPT, oversight of critical providers |
| Operational resilience (Bank of England, PRA, and FCA) | UK firms | Identify important business services, set a quantified impact tolerance for each, and stay within it (required since March 31, 2025) |
| CPS 230 (APRA) | Australian banks, insurers, and pension funds, since July 1, 2025 | Operational risk management, disruption tolerances, oversight of material service providers |
| Notice on Technology Risk Management (Monetary Authority of Singapore) | Singapore banks and major payment institutions | No more than four hours of unplanned downtime per critical system over 12 months, a four-hour recovery objective, notification to the regulator within one hour of discovery |
Under the UK regime, impact tolerance is the point beyond which disruption to a service becomes intolerable for customers or the market. It turns the usual logic around: the supervisor does not ask how long a system can hold up, but at what point its failure becomes unacceptable to the outside world. The firm first identifies its important business services, then sets that threshold for each one; the supervisor does not impose a duration. The chosen value is a commitment. The firm must then stay within it under severe but plausible scenarios, not just in the normal course of business.
These regimes produce public data, released on a fixed schedule by the authorities and by the operators they oversee, and practitioners have a direct interest in it. The Banco de Portugal counted 51 severe incidents reported in 2025, up from 35 the year before, affecting 1.8 million transactions and 1.3 million users (Relatório dos Sistemas de Pagamentos 2025). NPCI publishes the technical decline rates observed on UPI every month, bank by bank. These series apply the same measure to every listed player, so providers can be compared with one another. A contractual availability commitment, negotiated one institution at a time, cannot offer that.
Degraded mode and offline payments
Degraded mode is a point of sale’s ability to accept a payment without getting the issuer’s approval in real time. It depends on technical configuration, set at the acquirer as much as in the terminal, well before any incident. A merchant who goes looking for this setting during an outage almost always finds that it isn’t there. Turning it on requires the acquirer to update the terminal fleet, which cannot happen within the time frame of an incident.
How a card pays without a network
Chip cards have supported offline payments from the start, and the terminal relies on four levers to do it, all configured by the acquirer. The floor limit sets the amount below which a transaction goes through without querying the issuer. Offline data authentication (SDA, DDA, CDA) cryptographically verifies that the chip is genuine, without contacting a server. With offline PIN, the chip itself checks the PIN. Terminal risk management makes the final decision, combining cumulative limits, counters of consecutive offline transactions, and random selection. Approved transactions are held in a store-and-forward queue and submitted once the network is back.
Offline acceptance shifts risk between the parties in the chain. A transaction accepted offline has not been approved by the issuer. If the account is empty, or the card was blocked before the transaction, the loss is allocated under the scheme rules. Those rules set limits and assign liability in ways that vary by country and industry. A merchant that turns on offline acceptance without reading them is buying continuity without knowing either the limit or who bears the risk.
| Framework | Country and operator | Operating limit | Who bears the risk |
|---|---|---|---|
| BankAxept offline | Norway, Stø AS (1991) | Six hours by default, up to seven days as an option for essential-goods retailers | Risk shared among issuers |
| Dankort offline | Denmark, Nets, Nexi group (1983) | Up to DKK 20,000 cumulative | Issuers, under the industry agreement |
| Betalingsrådet (Danish Payments Council) arrangement | Denmark, a body run by Danmarks Nationalbank | At least seven days, on Dankort, Visa, Mastercard, Apple Pay, and Google Pay, in the main grocery chains; extension to all pharmacies announced for September 2026 | Set by the industry agreement, under central bank oversight |
| Low-value offline payments | India, Reserve Bank of India framework of January 3, 2022, updated December 4, 2024 | ₹500 per transaction and ₹2,000 per instrument; top-ups online only | Instrument issuer, within a regulatory cap |
| UPI Lite and UPI Lite X | India, NPCI (2022) | ₹1,000 per transaction and ₹5,000 on-device balance; Lite X adds device-to-device NFC transfers | Pre-funded wallet: the amount has already been debited |
| Qi Card | Iraq, International Smart Card (2007) | Offline capability designed for unreliable telecoms, on the channel used for public-sector salaries and pensions | Issuer |
The most extensive arrangements are in the Nordic countries, where cash has declined fastest and has taken away the fallback that other markets still have. On October 6, 2025, Danmarks Nationalbank published explicit recommendations. It urges merchants to accept cards and credit transfers in addition to cash, to turn on offline payments, and to train staff in emergency procedures. It advises households to keep at least two physical cards from different brands, with PINs memorized, and about DKK 250 per person in cash. The central bank notes that 80% of Danes have a card that works offline. The Riksbank takes the same line, recommends around SEK 1,000 per adult, and is working with Swish on an offline version of the wallet.
Cash remains the only instrument that works without a network or electricity, and several countries have written it back into law. Norway amended § 2-1 of the finansavtaleloven (Financial Contracts Act) through an act of June 7, 2024, which took effect on October 1, 2024. On premises where a business regularly sells to consumers, customers must be able to pay in legal tender, although the merchant may refuse amounts above NOK 20,000. Since January 1, 2025, the consumer protection authority has penalized violations. The regulation on calculating fines, in force since May 1, 2025, caps them at 4% of annual revenue or NOK 25 million.
Dependence on a single operator
A single point of failure is a component whose unavailability breaks the entire chain, no matter how many parties are involved. Apparent diversification can hide one. A merchant with two acquirers, three payment methods, and a written backup plan still depends on a shared component if all those paths converge on it. Four types of dependency produce this outcome, and each calls for a different fix, at a different level of the contractual chain.
A public authority can tackle this concentration head-on, as Uzbekistan did. After a software failure at the country’s only processing center, the central bank created a second card scheme in 2018, Humo, alongside Uzcard, which is operated by the Common Republican Processing Centre. The stated goal targeted the single point of failure as much as the monopoly. Few markets have followed suit. Most strengthen the incumbent operator rather than fund a second one.
The cost of a single point of failure is measured by how long restoration takes rather than by the length of the initial incident. Mozambique lost its main banking switch for more than two years. Restoring service in 2025 required an entirely new processing platform and the replacement of millions of cards and thousands of terminals and ATMs. A continuity plan that assumes the national switch is available therefore has no answer for the kind of prolonged outage Mozambique went through.
- Identify the switch before writing anything else. In a single-switch market, the question is not who acquires the payment but which route the authorization takes.
- Plan a fallback to another payment method, never to another acquirer: domestic instant transfers, a wallet run by a different operator, or cash under a set procedure.
- Document each provider’s second-tier dependencies: hosting provider, network operator, identity provider. DORA’s register of information offers a proven template.
- Check for co-location. Two providers hosted in the same region of the same cloud provider go down together, whatever the contracts say.
What a merchant can plan for
A payment continuity plan sets out how a business keeps taking payments when one layer of the chain becomes unavailable. It fits on one page and starts with a duration: how long the business can go without taking payments before the losses become unacceptable. That duration depends on the business, roughly an hour for a highway gas station and a day for a downtown boutique. It drives every decision that follows, because it determines what equipment is bought, what tests are run, and which sales the operator is willing to lose.
- Name the service to protect. In-store and online payment acceptance do not fail together and do not fall back the same way.
- Set a quantified tolerance for each service, approved by senior management, not by IT. UK impact tolerances rest on this logic, and it works for a merchant even without a regulator requiring it.
- Map dependencies down to the third tier: the provider, its acquirer, its hosting provider, its network operator.
- Set up a truly independent second path. Another acquirer is not enough if it uses the same national switch or is hosted in the same place. Ask the question in writing, and keep the answer.
- Configure, then test, offline mode: per-transaction limit, cumulative limit, maximum duration, POS software behavior. Test by unplugging, not by reading the documentation.
- Write the checkout procedure on a single page, post it in the back room, and have it run once a year by staff who have never been through an outage.
- Plan for cash, including float, security, and bank deposits. In several markets, accepting cash is now a legal requirement.
- Keep a usable log: timestamp, terminal ID, amount, raw response code. Without these four fields, no claim will succeed.
| Scenario | What still works | What doesn’t help |
|---|---|---|
| Widespread power outage | Cash, battery-powered terminal in offline mode, register on a UPS | Switching acquirers, calling the provider |
| Store loses connectivity | Backup SIM from another mobile operator, offline mode, cash | Rebooting the terminal over and over |
| Payment service provider incident | Second provider called directly, terminal connected to another acquirer | Waiting out the incident without failing over |
| National switch down | Cash, instant transfer if it runs on a separate rail, wallet run by another operator | A second acquirer, a second terminal, repeated retries |
| Single-issuer outage | The customer’s other card, another payment method, stand-in processing by the network | Replaying the transaction repeatedly, which risks scheme penalties |
| Identity provider outage | Postponing the sale, taking payment in store, a direct debit under an existing mandate | Retrying authentication, which will fail the same way |
After the outage: reconciliation and liability
The cost of a payment incident has two parts, separated in time. The first is lost sales, recorded the same day and immediately visible. The second shows up later, in reconciliation work, in unpaid offline transactions, and in disputes. Its size depends on the quality of the log kept during the incident, since transactions recorded without a timestamp or response code cannot be reconstructed once the entries are posted.
The most common mechanism is known as a ghost authorization. The terminal sends a request, gets no response within the timeout, and so has no idea what the issuer decided. The rule is not to assume a decline. The system must send a reversal (reversal advice) and repeat it until it is acknowledged. Otherwise the funds stay on hold in the cardholder’s account while a second attempt goes through. The customer then sees two holds for a single purchase, only one of which will actually be charged. After an outage of several hours, these orphaned holds run into the thousands, and clearing them continues long after service is restored.
- Reconcile three sources: transactions recorded at the register, transactions submitted for clearing, and payouts received from the acquirer. An incident opens gaps in all three directions.
- Track down unacknowledged reversals. Every request that got no response must have produced a confirmed reversal, and the terminal’s queue must be fully cleared.
- Isolate transactions accepted offline and track their reject rate. They were never authorized, and some will not be paid.
- Prepare for a wave of disputes: apparent double charges, unexpected amounts, missing receipts. A well-logged incident can be defended; an unlogged one is lost.
- Report when required. A European financial entity follows the DORA cascade. A payment service provider must also inform its users when an incident affects their financial interests, under Directive (EU) 2015/2366 (PSD2).
- Base claims on the contract, not on your estimate of the loss: only the signed clause has any force.
Post-incident follow-up is a governance matter. A payment outage should be investigated like any operational incident, with a written report, an identified cause, and a dated corrective action. Keeping that log over several years reveals patterns that no single incident report shows. The same causes keep coming back: a change made without a freeze, a forgotten certificate, a failover never tested, and a second-tier dependency no one had mapped. Each is fixed by a simple measure, and each comes back as soon as the tracking stops.