Reference🧭 Global overviewsAdvanced⏱ 18 min read

🔁 The payment state machine

The statuses a merchant system has to track on its own: idempotent submissions and replays, notifications that arrive out of order or not at all, recovery by polling, and payments stuck between authorization, capture, settlement, and refund

Payment states and their valid transitions

A payment's state is its position in the sequence of operations that leads from an initial request to a final movement of funds. Every party in the chain keeps its own version of that position, and the versions never quite match. The issuer sees a hold on funds, the acquirer a clearing presentment, the provider an object in its API with a vocabulary of its own. The merchant sees an order. That object exists in none of the other systems. The merchant system itself must therefore keep the order's state machine. No other party matches it against the operations executed on the rails.

On card rails, valid transitions are carried by standardized messages whose catalog is defined by ISO 8583. An authorization request goes out as a 0100 and comes back as a 0110. An authorization reversal uses the 0400 when it expects a synchronous response, and the 0420, repeated as a 0421, when a 0430 acknowledgment is enough. In a dual-message scheme, capture does not travel in these messages. It goes through clearing presentment, which is a file-based process. Settlement of funds comes later still. These four steps run in four separate systems, each on its own processing calendar. Single-message debit networks, by contrast, combine authorization and clearing, which closes the capture window and eliminates the states it creates.

EMVCo specifications add another state machine inside the card's chip, beneath the messages that travel over the network. The card generates a request cryptogram, the ARQC, which the issuer verifies and answers with an ARPC. The card then makes its decision at the second GENERATE AC, returning a TC when it accepts the transaction and an AAC when it declines it. An issuer approval is therefore not enough to complete the transaction. The final decision belongs to the chip. The terminal must read that last cryptogram before reporting a result to the POS system. A POS log that records the issuer approval without reading the final cryptogram reports as approved some transactions the card has declined, and the discrepancy only surfaces at clearing.

TransitionWho decides itWhat carries itReversible by
Request → authorizedThe issuer, with input from the network and the chip0100 / 0110, ARQC and ARPC cryptogramsAn authorization reversal, 0400 or 0420
Authorized → capturedThe merchantClearing presentment, as a file-based processA cancellation before presentment, a refund after
Captured → settledThe acquirer and the networkClearing files, then settlement filesNothing; a refund is a new transaction
Settled → refundedThe merchantAn offsetting financial transactionNothing; the credit is final once settled
Settled → disputedThe cardholder, through the issuerThe network's dispute cycleRepresentment, then arbitration
Authorized → expiredNo one; only timeNo messageNot applicable; the state exists only on the merchant side
What decides each transition in a card payment, and what it leaves reversible. The messages cited belong to ISO 8583 and the EMVCo specifications.

Rails with immediate settlement have few intermediate states. An instant credit transfer is either accepted or rejected, with no prior hold and no capture window. A card authorization is a revocable promise, whereas a completed instant credit transfer is a final movement of funds. The difference is reversibility. That single property drives the entire design of the state machine. A system that applies the same diagram to both families of rails gets this wrong. When a customer has paid for an order over an instant rail, the only correction is a refund, which is a new and deliberate operation, never a technical cancellation.

🔑
Three properties per state, not a label
A useful state carries three properties that a label alone leaves undefined. The first says whether money has moved. The second names the party that can trigger the next transition. The third sets the time after which remaining in that state becomes abnormal. A provider status named pending sometimes conveys all three, and more often none. The internal model is therefore written before the integration, then mapped to the provider's statuses through an explicit, versioned mapping table. That table is application code: it is tested and reviewed at every API version upgrade.
⚠️
The same word means different things at different providers
The status tables that major payment platforms publish look alike without matching. A payment labeled authorized sometimes means funds on hold at the issuer, and sometimes a recorded intent that has not yet reached any network. The pending status covers waiting for authentication, waiting for a slow rail, and waiting for a fraud review: three situations with nothing in common in timing or outcome. Copying the provider's status into the merchant's own database imports a vocabulary whose evolution the merchant does not control. The day the provider adds a value to its table, the merchant's code hits an unknown state in production. These incidents usually break out on a peak-traffic day, on a status that QA testing never encountered.

Idempotency: key, scope, and retention

Idempotency is the property of an operation that produces the same result whether it runs once or several times. Payment APIs rely on it for a common situation: a request that went out but whose response never came back. The caller then cannot tell whether the operation took place. The replay must produce the same result as the first attempt, without creating a second payment. The mechanism relies on a key supplied by the caller. The server associates the key with the result of the first execution and returns that result unchanged to later attempts. That is all it does. It determines what a second submission carrying the same key produces, and has no bearing on whether the payment is approved or declined.

The key derives from the business intent, never from the technical attempt. A key generated at random inside the retry loop lets duplicates through, since every attempt carries a new one. The key is created when the intent arises, persisted with the order, and reused for every attempt on that order. Stripe accepts keys of up to 255 characters, recommends a version 4 UUID, and explicitly rules out personal data as key material (Stripe, API documentation, consulted in August 2026). Adyen caps its keys at 64 characters and also recommends a UUID (Adyen, developer documentation, consulted in August 2026). A multi-provider integration that built its key once, to the length allowed by the more permissive of the two, would have its calls rejected by the other.

The scope of an idempotency key is the set within which its value must stay unique. It is the least documented parameter of the mechanism, and the most treacherous. A key is unique within a given API credential and endpoint, not globally. Documentation rarely says so. Scope follows from the authentication model: each set of credentials defines a separate key space at the provider. The same value sent to another provider, or to another endpoint at the same provider, identifies a different operation. Failing over to another route after an error therefore creates a second payment, whatever key it carries, because the second acquirer never saw the first attempt. Duplicate protection across a cascade has to be built into the merchant system, since no provider knows about attempts sent to its competitors.

The retention period is how long the provider keeps a key and the result associated with it. It determines how the system behaves after a long outage. Stripe says keys may be deleted after 24 hours, and that a key reused after that purge produces a new request. Adyen states that a key stays valid for at least 7 days after the first submission. PayPal uses the PayPal-Request-Id header, returns the state of the previous request, and refers to each API's documentation for the retention period. A retry queue restarted after 48 hours therefore falls outside the first window and inside the second. Check this period before writing the retry policy, and check it again whenever the provider changes. It marks the point beyond which a new submission creates an additional payment instead of returning the first result.

255 characters
maximum length of an idempotency key; version 4 UUID recommended
Stripe, API documentation, consulted in August 2026
24 hours
time after which a key may be deleted from the system, so that a replay creates a new request
Stripe, API documentation, consulted in August 2026
64 characters
maximum key length at another provider; UUID recommended
Adyen, developer documentation, consulted in August 2026
7 days
minimum validity of a key after its first submission at that same provider
Adyen, developer documentation, consulted in August 2026

Conflict handling is a matter of practice rather than of any binding standard. Stripe compares incoming parameters with those of the original request and returns an error if they differ. Adyen responds with a 409 or 422 and error code 704 when two requests with the same key arrive at the same time. The IETF HTTPAPI working group carried a draft meant to standardize these conventions, draft-ietf-httpapi-idempotency-key-header. Its version 07 is dated October 15, 2025, and expired on April 18, 2026. The text never became an RFC, and no provider is required to follow it. It defines three responses: a 400 when the required header is missing, a 422 when a key is replayed with a different payload, and a 409 when a replay arrives before the original request has completed. Stripe, for its part, documents when the idempotent result is recorded. No result is saved until execution has started, so a request rejected at parameter validation can be resent as is.

Sequencing the write and the call: the intent is persisted before any network call
1. write the payment intent to the database, THEN commit the transaction
   (order_id, amount, currency, idempotency_key, state = SUBMISSION_IN_PROGRESS)
   -> the commit comes before any network call, no exceptions

2. call the provider with this key
   2xx response     -> state = AUTHORIZED, store the provider's ID
   4xx response     -> state = DECLINED, store the code received
   timeout or 5xx   -> state UNCHANGED, the row stays SUBMISSION_IN_PROGRESS

3. a background job picks up every row still in SUBMISSION_IN_PROGRESS
   within the retention window -> replay with THE SAME key
   beyond the window           -> query the state, never a new call

4. the key combines the order ID and a LOGICAL attempt
   number (a deliberate resubmission = new attempt = new key),
   never a hash of the amount and date, which collides

Idempotency protects the submission of the request. The payment itself is outside its scope. Four limits bound what it covers, and each one should be verified before go-live.

  • It does not reach the network. A replayed key prevents a second API call. It does not cancel an authorization the first attempt already obtained; that takes a 0420 authorization reversal, repeated until the 0430 acknowledgment.
  • It does not cross provider boundaries. A cascade to a second acquirer falls outside the key's uniqueness scope, so two authorizations can coexist on the cardholder's account.
  • It does not outlive its retention window. Once the period the provider states has passed, a replay is no longer a replay but a new creation.
  • It does not protect against a business collision. Two distinct intents that produce the same key make a payment vanish, with no error and no alert, and the gap only shows up at reconciliation.
⚠️
The key is persisted before the call, never after
The costliest mistake is to generate the key in memory, call the provider, and only then write whatever the response returned to the database. A process that dies between the call and the write loses the key along with its memory. The retry then starts over with a new key, on an operation that may well have succeeded, and the second payment is created by the very mechanism meant to prevent it. The correct order has three steps: persist the intent and its key, commit the transaction, and only then make the network call. Otherwise, the resulting duplicates are rare, impossible to reproduce in testing, and concentrated on peak-traffic days when response times stretch out.

Asynchronous notifications: ordering, replay, signatures, and acknowledgment

An asynchronous notification is a message the provider sends to a merchant endpoint when a state changes on its side. The message signals the change; it is not the record of it. The provider retries it if no acknowledgment arrives, within a window it sets itself, and guarantees neither delivery order nor uniqueness. Stripe says so in its documentation: events are not delivered in the order they were generated, and the same endpoint can receive the same message more than once. A merchant system that rebuilds its state solely from the sequence it receives therefore ends up with the wrong state, without raising any error or leaving any trace in the logs. The discrepancy comes to light through a customer complaint or the next day's reconciliation.

Out-of-order delivery means receiving an older notification after a more recent one. It happens often: a single operation generates several events a few milliseconds apart, delivered over independent network paths. The failure notification for one attempt sometimes arrives after the success notification for the next attempt. Protection rests on a single rule. On every receipt, a notification advances the state only if the transition is valid from the current state; any invalid transition is logged and then discarded. When the provider timestamps its events, the timestamp serves as a second guard, since a message older than the current state can no longer roll it back.

Authenticating a notification distinguishes a message sent by the provider from one forged by a third party, and no other check makes that distinction. Stripe signs every delivery in a Stripe-Signature header, which carries a timestamp prefixed with t= and one or more signatures prefixed with v1=. The signature is an HMAC-SHA256 computed over the concatenation of the timestamp, a period, and the raw request body. The body must be captured before any web framework processes it, because even rewriting whitespace breaks verification. By default, the official libraries reject a timestamp more than 5 minutes old, which limits the replay window for an intercepted message. During a secret rotation, two secrets stay active for up to 24 hours, and each delivery then carries one signature per secret. Stripe also publishes the list of IP addresses it sends from, and recommends combining network filtering with signature verification rather than choosing one over the other.

The acknowledgment is the status code the merchant endpoint returns to the provider. Its value determines what happens next. A 2xx code stops retries, anything else triggers them, and Stripe counts a 3xx redirect as a failure. The recommended sequence is to verify the signature, write the raw event to a queue, return 2xx, and then process the event outside the request cycle. Processing before responding risks a timeout, which triggers a retry, which creates a duplicate. Stripe retries for 3 days in production, with increasing intervals. Manual resends remain possible for 15 days from the web interface and 30 days from the command-line tool. Adyen asks merchants to accept the message with a 2xx, store it, and only then process its contents.

3 days
how long a provider retries delivery of an event in production, with increasing intervals
Stripe, webhooks documentation, consulted in August 2026
5 minutes
default tolerance in the official libraries between the signed timestamp and the current time
Stripe, webhooks documentation, consulted in August 2026
15 and 30 days
windows for manually resending an event, from the web interface and from the command-line tool
Stripe, webhooks documentation, consulted in August 2026
TLS 1.2
minimum version the merchant endpoint must support to receive notifications
Stripe, webhooks documentation, consulted in August 2026

A dead-letter queue holds the messages that processing could not handle. It serves two purposes: it keeps a poison message from blocking the main queue, and it makes failures visible. A dead-letter queue with no named owner and no alert on message age piles up messages nobody picks up, and the payments involved stay in a stale state. Replays from this queue go through the same transition check as the normal path; otherwise they could advance a state that the normal path had deliberately rejected.

FailureSymptomCountermeasure
Notification never deliveredThe state stays frozen at the merchant after it has changed at the providerPolling, then event reconciliation
Notification delivered twiceThe same change processed twice, two accounting entriesDeduplication on the event ID, enforced by a unique constraint in the database
Notifications out of orderA terminal state overwritten by an earlier stateValid transitions only, with the event timestamp as a guard
Forged notificationA state advanced when no payment took placeSignature verification on the raw body, timestamp window, constant-time comparison
Processing failed after acknowledgmentThe provider counts the delivery as successful; the merchant recorded nothingDead-letter queue with an owner, an age alert, and controlled replay
Five failure modes of a notification channel, and what covers them
⚠️
The notification payload is not the source for the amount
A notification announces that an object has changed. It does not prove what the object contains now, even when signed, because its payload reflects the moment it was sent and retries can deliver it much later. The amount actually collected, the current state, and the fees are read from the provider, using the ID the notification carries. Stripe formalized the distinction with so-called thin events, which carry only an ID and a type and which the receiver completes with an explicit API call. The same precaution applies whatever the format, including when the notification carries the full object. Deduplication follows the same logic. Stripe recommends logging the IDs of events already processed, and notes that two distinct events sometimes describe the same change. The deduplication key must therefore also include the event type and the ID of the object concerned.

Recovering by polling: the notification is never the source of truth

Recovery by polling means periodically asking the provider for the current state of payments whose outcome is still unknown. It complements the notification channel, which speeds up awareness of a change without guaranteeing delivery. The three providers cited above all put a time limit on their retries, and none promises delivery under all circumstances. An endpoint that is down for 3 days in a row falls outside the automatic retry window of a provider that retries for 3 days. Events from that period then come back only through a manual resend, while that option remains open, or through polling. The source of truth is the state held by the party that executes the payment. The only way to get it is to ask that party.

Polling targets payments left in a non-terminal state, ordered by age, with increasing intervals between attempts, rather than sweeping the entire database at a fixed interval. The polling rate follows the expected lifetime of each state, which ranges from a few seconds to several days depending on the rail. An instant payment that has not completed past its rail's hard deadline is abnormal. A card authorization awaiting capture is not abnormal for several days. A single polling rate therefore causes two problems. It generates pointless calls for states whose normal lifetime is long, and it detects incidents too late for states whose lifetime is measured in seconds.

The protocol linking the POS system to the terminal formalized the same need long ago. In the Retailer protocol from nexo Standards, every request-response pair carries a ServiceID, a short identifier that the specification recommends limiting to 10 characters and whose uniqueness is the responsibility of the POS system that issues it. A web API idempotency key and this identifier are therefore built differently: one accommodates a UUID that the other cannot hold. When the payment response does not come back, the POS system does not send a second payment request. It sends a transaction status request carrying the original ServiceID, and the terminal returns the last response it produced for that POS-terminal pair. If the transaction is still in progress, the response carries a failure result with the Busy condition, which calls for another status request and does not mean a decline. A POS system that treats this condition as a decline sends a second payment request and charges the customer twice.

On instant rails, no response within the time limit amounts to a rejection. The European Payments Council's SCT Inst scheme sets a maximum execution time of 10 seconds. A hard deadline of 20 seconds runs from the timestamp applied by the originator's bank, and past that point the participants in the chain must reject the transaction. A pacs.002 that has not arrived within the window therefore means rejection, never waiting. In the US, the Federal Reserve's FedNow Service applies a 20-second clock. Each receiving bank reserves 1 to 5 seconds of it, depending on its capacity to process the message (Federal Reserve, Understanding the payment timeout clock). An application timeout shorter than these limits leads the merchant to treat a completed operation as uncertain.

The card rail has its own recovery mechanism, older than web APIs and built on the same principle. A terminal that gets no 0110 within the time limit does not assume a decline. It sends a 0420 authorization reversal, repeated as a 0421 until the 0430 acknowledgment arrives, usually from a local store-and-forward queue. The difference from polling lies in the direction of the operation. The terminal does not ask for the state. It forces the cancellation of a state it does not know, which has no effect if the issuer never authorized anything.

🔑
Three questions a recovery loop must answer
A recovery loop answers three questions as of a given date, and produces a figure for each. The first concerns payments still in a non-terminal state past the time limit for that state. The second concerns states that changed at the provider without the merchant system processing any matching notification. The third concerns payments the merchant holds in a state the provider does not recognize. That last question is the one least often asked, and the only one that detects a state advanced in error on the merchant side. The provider has no way to answer it.
  • Query by state, not by table. The recovery query covers non-terminal states and nothing else, with a dedicated index; a full scan becomes impractical once the table reaches a few million rows.
  • Space out attempts. Increasing intervals, with a cap, keep an availability incident at the provider from turning into a denial-of-service attack launched by the merchant.
  • Treat the response like a notification. The polling result goes through the same transition check; otherwise, two write paths apply two different rules to the same state.
  • Log the origin of every transition. Notification, polling, reconciliation, or human action; without that trail, no status incident can be explained after the fact.

Payments stuck between authorization, capture, settlement, and refund

A stuck payment is one held in a non-terminal state beyond that state's normal lifetime. Such payments are the routine residue of a chain in which several systems move at different speeds, and their volume measures the quality of an integration better than a payment success rate does. Each type has a recognizable signature, a query that detects it, and a corrective action of its own. Bulk handling, in a single sweep that makes no distinction between cases, creates the very duplicates that recovery was supposed to eliminate.

The most common case is the authorization that is never captured. The funds stay on hold at the issuer, invisible in the merchant's tools but clearly visible to the cardholder, who takes the hold for a charge. Letting the authorization expire has a cost: the networks charge fees for authorizations with no follow-up. Clearing presentment deadlines vary by merchant segment and appear in both the Visa Core Rules and the Mastercard Rules, which are published and revised periodically. Read these deadlines in the current edition, never in an internal memo copied from year to year, which those periodic revisions render worthless. The corrective action is to send an authorization reversal as soon as the order is abandoned, without waiting for the authorization to expire.

Capture without settlement refers to a capture the provider accepted that never resulted in a payout. It is detected in files rather than through the API. An accepted capture does not automatically land in a payout batch: a clearing reject, a blocked batch, or a reserve hold can keep it out. The symptom is a payment marked as settled on the merchant side but missing from the settlement report for the corresponding day. Detection therefore cross-checks the internal state against the acquirer's file. Partial capture complicates matters further, since it leaves a remaining authorization amount that must be reversed separately to release the rest of the held funds.

A refund with no matching capture is the most dangerous case, because it moves real money in the wrong direction. It stems from an operator refunding a canceled order whose authorization was never presented, or from a poorly scoped recovery script that replays an entire queue. The credit goes out with no sale behind it. The networks restrict credits without an original transaction precisely because they are also used to drain compromised merchant accounts. The rule that settles the matter requires every refund to reference a settled capture, and caps total refunds on an order at the captured amount.

Authorization expiry is the end of the hold on funds at the issuer. No message accompanies this state. No network notifies the merchant that the hold has lapsed. The duration depends on the issuer, on the type of authorization (an estimated authorization is not a final authorization), and on the merchant segment. It is never hard-coded into the merchant system. Instead, the system sets its own deadline, shorter than the shortest duration observed across its portfolio, and treats an overrun as an order to reauthorize rather than capture.

CaseDetection queryMain causeAction
Authorized, never capturedAUTHORIZED state older than the internal capture deadlineOrder abandoned without a reversal, or capture job failedImmediate authorization reversal, then an alert if the volume persists
Captured, never settledInternal captures with no matching line in the settlement reportClearing reject, blocked batch, reserve holdReconciliation against the acquirer file, then a dated claim
Refunded without captureRefunds whose capture reference is missing or not settledManual action on a canceled order, or replay of an unbounded queueApplication-level block on unreferenced credits, access rights review
Authorization expiredAUTHORIZED state past the shortest validity period in the portfolioNone, only time; no message announces the expiryMove to EXPIRED state, reauthorize before any capture
Submission pending indefinitelySUBMISSION_IN_PROGRESS state older than the key retention windowResponse lost, then retry abandoned or never scheduledQuery the state at the provider, never a new creation call
Five kinds of stuck payment, the query that reveals each one, and the corrective action
⚠️
The “expired” state exists only on the merchant side
No message announces it, no report lists it, and the provider's API rarely exposes it as such. The state is something the merchant system infers, not information it receives. That system is the only place the state can originate, based on a clock the merchant sets itself. A data model with no deadline per state will never produce this state, and the affected orders will remain pending indefinitely until an audit turns them up. The consequence shows up first in customer support, over holds that cardholders see on their accounts and that nobody in the company can explain.

Event reconciliation as a safety net

Event reconciliation is the periodic comparison of a payment's state across the systems that track it. It comes before amount reconciliation, which compares the payout received at the bank with the net amount the provider calculated, and which belongs to a separate accounting process. Event reconciliation covers three data sets that should say the same thing at the same moment. It catches what the notification channel lost, before the discrepancy reaches the general ledger and becomes a period-end close issue. The two exercises complement each other without overlapping, one checking amounts and the other states, and a chain that runs only one of them discovers its gaps through the accounting.

Three data sets are compared: the state held by the merchant system, the list of payments as the provider exposes it through its API or a dated export, and the acquirer's settlement file. A payment must appear in all three, in states consistent with one another. Any line present in one set and missing from another is a discrepancy, and each type of discrepancy points to a different main cause.

Discrepancies fall into four categories, and their breakdown says more than an overall rate. A notification never received leaves the state frozen at the merchant while the provider has moved on to the next state. A notification received and then lost during processing produces the same symptom, but leaves a trace in the acknowledgment logs, which distinguishes it from the first case. A state change with no matching notification mostly involves disputes and issuer-initiated operations. A state advanced with no counterpart at the provider reveals a bug in the merchant's code or a manual intervention, a situation no provider is in a position to flag.

The reconciliation window is aligned with the provider's retry period rather than with the accounting day. A provider that retries for 3 days will deliver a Saturday event on Tuesday, and a strictly daily reconciliation will count the corresponding discrepancy twice. The usual practice is a daily run over the last 24 hours, followed by a weekly run over a wider window, with a resolution status and an owner attached to each discrepancy. The export itself must be bounded by dates and safe to rerun, or it will create the very discrepancies it was meant to measure. A discrepancy still open beyond its window is handled as an incident, with a named owner and a resolution deadline.

  • Age of the oldest non-terminal payment, by state and by provider. This indicator rises before all the others when a notification channel goes down, and it rises even when the number of discrepancies stays low.
  • Number of merchant states with no matching event at the provider, over the long window. A persistent nonzero value points to a code defect, never to a network glitch.
  • Depth and age of the dead-letter queue. A queue that empties on its own proves nothing; only its maximum age tells you what was actually processed.
  • Rate of events rejected for an invalid transition. An increase signals a vocabulary change at the provider, often announced in release notes nobody read.
  • Gap between the number of internal captures and the number of lines in the settlement report, by day. It links event reconciliation to amount reconciliation and deserves the same monitoring.
🔑
The metric that judges a payment status chain
A payment chain is judged by the number of payments whose merchant-side state and actual state diverge at any given moment. That number never falls to zero, and it should not, since transient states exist and propagation takes time. It must stay stable and bounded, with few old cases. The four mechanisms described in this guide divide the work and never substitute for one another. Idempotency prevents duplicates at submission, notifications speed up awareness of state changes, polling catches what was not delivered, and reconciliation proves that no payment slipped through unnoticed.