Skip to main content
Reliability, async & real-time
信頼性
アーキテクチャ設計
決済
AWS
Python

What Exactly-Once Actually Guarantees: Separating Delivery, Processing and Effect, and Converging External Side Effects to One with Idempotency Keys and Reconciliation

Exactly-once neither 'doesn't exist' nor 'works everywhere': its meaning depends on the boundary and on what it covers — delivery, processing or effect. This sorts out, from primary sources, what Kafka transactions, SQS FIFO and database transactions guarantee and what is not guaranteed for a business effect spanning PostgreSQL, Stripe and Lambda, then shows tested Python code that converges an external API call whose response was lost to exactly one effect, using idempotency keys and reconciliation.

Published
Reading time
20 min read
Author
友田 陽大
Share

The short answer. "Exactly-once" has no fixed meaning until you say what is once (delivery, processing or effect) and within which boundary. Within a limited boundary — a single database transaction, Kafka transactions, SQS FIFO de-duplication — exactly-once guarantees really exist. But a business effect that spans independent systems such as PostgreSQL, Stripe, an e-mail provider and Lambda can't simply inherit any one product's guarantee, because when an external API succeeds and only its response is lost, the caller can't know whether it succeeded. What you actually do is assume at-least-once delivery, converge the effect to one with an operation ID persisted before the side effect and an idempotency key, and fill the gap idempotency keys can't cover with reconciliation.

This article first separates the terms precisely, then checks against primary sources what the main technologies officially guarantee. It then implements convergence to a single business effect, using a Stripe API call whose response was lost as the example. The code was verified on PostgreSQL 18.6 / psycopg 3.3, with tests that use a fake mimicking Stripe's idempotency layer (October 2026).

Position: this article says neither "exactly-once doesn't exist" nor "use this product and you get exactly-once". Each product's guarantee is cited within the scope its official documentation states, and never stretched beyond it.


1. Separate five terms

Arguments about "exactly-once" go in circles because different things are called by the same name.

TermWhat happens onceExample
Exactly-once deliveryHow many times a message reaches the receiver"This message is handed to the consumer once"
Exactly-once processingHow many times the result of processing a message is reflected in stateCommitting the read position and the output together in a Kafka transaction
Exactly-once effect / effectively-onceHow many times a business side effect happens (a charge, a ledger entry, a stock reservation)However often you retry, one charge per order
Atomic local transactionWithin one system, several writes all happen or none doAn UPDATE and an INSERT in one PostgreSQL transaction
Distributed side effectAn effect spanning several independent systemsUpdate the DB, charge in Stripe, send an e-mail

This distinction resolves a lot of arguments:

  • At-least-once delivery can still produce a single effect. If the receiver is idempotent, two deliveries produce one effect. Most real systems work this way.
  • Exactly-once delivery doesn't imply a single effect. If the receiver crashes mid-processing, leaving the side effect but not the "processed" record, reprocessing causes the effect twice.
  • Local atomicity doesn't make distributed side effects atomic. A PostgreSQL transaction can't roll back a Stripe API call.

Don't confuse delivery guarantees with business guarantees: what messaging infrastructure guarantees are properties of delivery or processing, not business guarantees such as "charge the customer exactly once". Business guarantees are built by the application, with idempotency and reconciliation.


Double charges, missed webhooks, or churn from failed payments?

Payment-reliability diagnosis and billing architecture, as a technical advisor

2. Within a limited boundary, exactly-once exists

Saying "exactly-once doesn't exist" and leaving it there is inaccurate. Here are cases where official documentation claims exactly-once, and their scope.

2.1 A single database transaction

Within a PostgreSQL transaction, you can commit the business update together with a record that says "this message was processed". If it commits, both remain; if it fails, both disappear. Within that scope, the effect of processing happens exactly once. Applying this to webhook receipt is the transactional inbox; applying it on the sending side is the transactional outbox.

2.2 Kafka transactions

Apache Kafka's documentation (Design, Message Delivery Semantics) first cautions that claims of exactly-once require reading the fine print, then explains that for processing that reads from Kafka topics and writes to other topics, the transactional producer and a consumer using the read-committed isolation level can provide exactly-once delivery. The key is putting the output writes and the update of the consumer's position (the committed offset) in one transaction.

The same documentation also says that when writing to an external system, you need to coordinate the consumer's position with what is actually stored as output. Kafka's exactly-once covers "reading, processing and writing between Kafka topics"; it doesn't extend to calling an external payment API from Kafka.

2.3 SQS FIFO de-duplication

The AWS documentation page for SQS FIFO queues is titled "Exactly-once processing" and explains that if you retry SendMessage within the five-minute de-duplication interval, SQS doesn't introduce duplicates into the queue. You either pass a de-duplication ID explicitly or enable content-based de-duplication (a SHA-256 of the body).

The scope of this guarantee is "no duplicate messages enter the queue." If a consumer receives a message, processes it, and crashes before deleting it, the same message is received again after the visibility timeout. Consumer-side side effects don't become exactly-once.

For standard queues, AWS explains that because copies of messages are stored on multiple servers, a copy of a deleted message can occasionally be delivered again, and it states explicitly that you should design your applications to be idempotent (at-least-once delivery).

2.4 Lambda asynchronous invocation

With asynchronous invocation, when a function returns an error, Lambda by default retries two more times (one minute between the first and second attempts, two minutes between the second and third). For throttling and system errors it returns the event to its queue and retries for up to six hours by default (official documentation, as of October 2026). If a function times out after causing a side effect, the same event causes the side effect again. The documentation also states that even if your function doesn't return an error, it can receive the same event more than once, because the queue itself is eventually consistent. Lambda functions need to be written idempotently.


3. What happens across boundaries: an API call whose response is lost

Here is the most important failure, as a concrete timeline: charging a saved card when an order is confirmed.

Caller (your service)                         Stripe
t1  POST /v1/payment_intents  ─────────────▶ executes the charge (succeeds)
t2                             ◀───── ✕ ─── 200 OK (lost to a network fault or timeout)
t3  ReadTimeout exception
    → was the customer charged or not? You can't tell

At t3 the caller has only three choices:

  1. Treat it as a failure and retry → if it actually succeeded, a double charge.
  2. Treat it as a success → if it actually failed, an uncollected payment.
  3. Record it as unknown and find out later → this is the right answer, but you have to design how "find out later" works.

Stripe's documentation (Advanced error handling) also describes states where, because of network problems, the client can't know whether the server received the request, and advises retrying such requests with an idempotency key. It also says to treat the result of a 500 as indeterminate.

No amount of communication engineering removes that indeterminacy. As long as messages between the caller and the other side can be lost, the caller can't know for certain whether the other side executed the request. So the design goal isn't to eliminate the unknown, but to converge from an unknown state to one where the business effect happened exactly once.


4. Four parts that converge a business effect to one

PartRole
Durable operation identityRecord "we will perform this operation" in the database before the side effect, so retries and restarts refer to the same operation
Stable idempotency keyDerived deterministically from the operation ID, so retries of the same operation look like the same request to the receiver
RetryRe-run unknown or transient failures with the same key
ReconciliationSettle operations that outlived the idempotency key's lifetime, or that retries can't settle, by checking the external state

No single part is enough. An idempotency key alone can't be regenerated after a restart if it was never stored. An operation ID alone doesn't let the receiver recognise duplicates. Retries alone cause double execution, and reconciliation alone is too slow.

4.1 The trap in SDK-generated idempotency keys

The stripe-python README explains that with max_network_retries set, the library automatically retries connection errors, timeouts, 409 Conflict and so on, and generates and attaches an idempotency key if you didn't provide one. That's useful, but its protection is limited:

  • An auto-generated key is shared only across retries within that single SDK call.
  • If the call ends in an exception and your application retries, a new key is generated.
  • If the process crashes and restarts, the original key exists nowhere.

So for application-level retries, or retries across processes, auto-generated keys don't prevent double execution. Derive idempotency keys from an operation ID persisted in the database.


5. Implementation: record the operation first, retry with the same key, reconcile past the key's lifetime

Assumption in the code: the psycopg code assumes connections opened with psycopg.connect(dsn, autocommit=True) and expresses transaction boundaries only with with conn.transaction():. Every statement, reads included, runs inside transaction(), so even on a connection without autocommit no transaction is left open while the external API is called (verified by a test). The snippets are shown in reading order, so put from __future__ import annotations at the top of the module (annotations such as op: Operation refer to classes defined further down), and gather the imports they use: uuid, dataclass, timedelta, Protocol, psycopg, stripe and from psycopg.rows import class_row.

5.1 Schema: record the operation before the side effect

CREATE TABLE payment_operations (
    id                 uuid        PRIMARY KEY DEFAULT gen_random_uuid(),
    order_id           text        NOT NULL,
    kind               text        NOT NULL CHECK (kind IN ('charge')),
    amount             bigint      NOT NULL CHECK (amount > 0),
    currency           text        NOT NULL,
    customer_id        text        NOT NULL,
    payment_method     text        NOT NULL,
    status             text        NOT NULL DEFAULT 'pending'
                       CHECK (status IN ('pending', 'unknown', 'succeeded', 'failed')),
    external_id        text,                    -- pi_... (filled in once we know it exists externally)
    attempts           integer     NOT NULL DEFAULT 0,
    first_attempted_at timestamptz,             -- when we first tried to send it (start of the idempotency key's lifetime)
    last_error         text,
    created_at         timestamptz NOT NULL DEFAULT now(),
    updated_at         timestamptz NOT NULL DEFAULT now(),
    UNIQUE (order_id, kind)                     -- business identity: one charge operation per order
);

The key is having an explicit unknown status. first_attempted_at exists so that the idempotency key's lifetime is measured from "when we first tried to send it", not "when the operation was created" (section 5.3). Rounding a timeout to failed treats a successful charge as a failure and asks the customer to pay again.

UNIQUE (order_id, kind) makes the database enforce business identity. Even if the order's confirm button is pressed twice, only one operation is created.

def request_charge(
    conn: psycopg.Connection,
    *,
    order_id: str,
    amount: int,
    currency: str,
    customer_id: str,
    payment_method: str,
) -> uuid.UUID:
    """Persist the operation before the side effect. However often it's called for an order, there is one operation."""
    with conn.transaction():
        row = conn.execute(
            """
            INSERT INTO payment_operations (order_id, kind, amount, currency, customer_id, payment_method)
            VALUES (%s, 'charge', %s, %s, %s, %s)
            ON CONFLICT (order_id, kind) DO NOTHING
            RETURNING id
            """,
            (order_id, amount, currency, customer_id, payment_method),
        ).fetchone()
        if row is None:
            row = conn.execute(
                "SELECT id FROM payment_operations WHERE order_id = %s AND kind = 'charge'",
                (order_id,),
            ).fetchone()
        assert row is not None  # the conflicting row is committed, so it's always visible
        op_id: uuid.UUID = row[0]
        return op_id

In production, create this operation in the same transaction that confirms the order, and leave its execution to a dispatcher via the outbox (leases and fencing tokens in an outbox dispatcher). That avoids the two-write problem of "order confirmed, charge operation never recorded".

5.2 Classify external results by "was the outcome settled?"

class AmbiguousOutcome(Exception):
    """We can't tell whether the request reached the other side and ran (timeout, connection loss, 500)."""


class RetryLater(Exception):
    """A transient error where the request certainly did not run (429 and the like)."""


class Declined(Exception):
    """A settled failure (card declined and the like). Don't retry."""


class PaymentGateway(Protocol):
    def create_charge(self, op: Operation) -> ExternalCharge: ...
    def retrieve_charge(self, external_id: str) -> ExternalCharge: ...
    def find_charges(self, op: Operation) -> list[ExternalCharge]: ...


def outcome_for(external_status: str) -> str:
    """Map a PaymentIntent status to the operation's state. Only succeeded counts as success."""
    if external_status == "succeeded":
        return "succeeded"
    if external_status in ("canceled", "requires_payment_method"):
        return "failed"
    # processing / requires_action / requires_capture and so on: not settled yet
    return "unknown"


class StripeGateway:
    """The boundary that sorts Stripe SDK errors into three groups by whether the outcome is settled."""

    def __init__(self, client: stripe.StripeClient) -> None:
        self._client = client

    def create_charge(self, op: Operation) -> ExternalCharge:
        try:
            pi = self._client.v1.payment_intents.create(
                params={
                    "amount": op.amount,
                    "currency": op.currency,
                    "customer": op.customer_id,
                    "payment_method": op.payment_method,
                    "confirm": True,
                    "off_session": True,
                    "metadata": {"operation_id": str(op.id)},
                },
                options={"idempotency_key": op.idempotency_key},
            )
        except stripe.CardError as exc:
            raise Declined(exc.code or "card_error") from exc
        except stripe.RateLimitError as exc:
            raise RetryLater(str(exc)) from exc
        # stripe.IdempotencyError (same key, different parameters) is not caught: it's a key-design bug, so surface it
        except (stripe.APIConnectionError, stripe.APIError) as exc:
            raise AmbiguousOutcome(type(exc).__name__) from exc
        return ExternalCharge(pi.id, pi.status)

    def retrieve_charge(self, external_id: str) -> ExternalCharge:
        pi = self._client.v1.payment_intents.retrieve(external_id)
        return ExternalCharge(pi.id, pi.status)

    def find_charges(self, op: Operation) -> list[ExternalCharge]:
        # Search is not real-time (normally under a minute). Don't use it right after creating something
        result = self._client.v1.payment_intents.search(
            params={"query": f"metadata['operation_id']:'{op.id}'"}
        )
        return [ExternalCharge(pi.id, pi.status) for pi in result.data]

Why each is classified the way it is:

  • Card declined (CardError): Stripe processed the request and returned "declined", so the outcome is settled. Don't retry.
  • Rate limited (RateLimitError, 429): the request wasn't processed, so it's settled (not executed). Retry later with the same key.
  • Connection errors (APIConnectionError) and server errors (APIError, such as a 500) are unknown. Stripe's documentation also says to treat a 500's result as indeterminate.
  • IdempotencyError (different parameters under the same key) is a bug in how keys are built. Don't swallow it; surface it.

Also, no exception doesn't mean the charge succeeded. Even an off-session confirm (confirm: True, off_session: True) can return a PaymentIntent in processing or requires_action. outcome_for() maps the PaymentIntent's status to the operation's state: only succeeded is success, canceled and requires_payment_method are failure, and everything else is unknown (check later).

The operation ID goes into metadata so that, once the idempotency key's lifetime has passed, the operation can be identified from external state (section 5.4).

5.3 Derive the idempotency key deterministically from the operation ID

# Stripe may remove idempotency keys once they're at least 24 hours old. Only resend with the same key before that
IDEMPOTENCY_SAFE_WINDOW = timedelta(hours=23)


@dataclass(frozen=True)
class Operation:
    id: uuid.UUID
    order_id: str
    amount: int
    currency: str
    customer_id: str
    payment_method: str
    status: str
    external_id: str | None
    within_key_window: bool

    @property
    def idempotency_key(self) -> str:
        # Derived deterministically from the operation's durable ID: the same key even after a restart
        return f"payment-op:{self.id}"


@dataclass(frozen=True)
class ExternalCharge:
    id: str
    status: str  # Stripe PaymentIntent.status

According to Stripe's documentation, subsequent requests with the same idempotency key return the result of the first request, whether it succeeded or failed (including 500 errors). The idempotency layer also compares incoming parameters with the original request and errors if they differ. Keys can be up to 255 characters, and Stripe advises against including personal data.

The documentation also explains that keys can be removed once they are at least 24 hours old, and a key reused after removal creates a new request. This article uses 23 hours from the first send attempt as the "safe to resend with the same key" window, leaving a margin. If the clock started at the operation's creation, an operation that merely waited in a queue and was never sent would be treated as "the key may be gone" and needlessly sent for human review, so first_attempted_at is recorded just before sending. The 23 hours is my conservative judgement, not something Stripe guarantees.

5.4 Execute: a function that's safe to call any number of times

def _load(conn: psycopg.Connection, op_id: uuid.UUID) -> Operation:
    with conn.transaction(), conn.cursor(row_factory=class_row(Operation)) as cur:
        cur.execute(
            """
            SELECT id, order_id, amount, currency, customer_id, payment_method, status, external_id,
                   (first_attempted_at IS NULL OR first_attempted_at > now() - %s) AS within_key_window
              FROM payment_operations WHERE id = %s
            """,
            (IDEMPOTENCY_SAFE_WINDOW, op_id),
        )
        op = cur.fetchone()
    if op is None:
        raise LookupError(op_id)
    return op


def _mark_attempt_started(conn: psycopg.Connection, op_id: uuid.UUID) -> None:
    # Record this BEFORE sending, so the key's lifetime has a start point even if we die right after sending
    with conn.transaction():
        conn.execute(
            "UPDATE payment_operations SET first_attempted_at = coalesce(first_attempted_at, now()) WHERE id = %s",
            (op_id,),
        )


def _record(conn: psycopg.Connection, op_id: uuid.UUID, status: str, *, external_id: str | None = None,
            error: str | None = None) -> None:
    with conn.transaction():
        conn.execute(
            """
            UPDATE payment_operations
               SET status = %s, external_id = COALESCE(%s, external_id), last_error = %s,
                   attempts = attempts + 1, updated_at = now()
             WHERE id = %s AND status IN ('pending', 'unknown')   -- never overwrite a settled state
            """,
            (status, external_id, error, op_id),
        )


def execute_charge(conn: psycopg.Connection, gateway: PaymentGateway, op_id: uuid.UUID) -> str:
    """Safe to call any number of times. The external API call happens outside any transaction."""
    op = _load(conn, op_id)
    if op.status in ("succeeded", "failed"):
        return op.status

    if op.external_id is not None:
        # We know it was created externally (it stopped at processing or similar).
        # Resending with the same key only replays the first response, so read the latest state
        charge = gateway.retrieve_charge(op.external_id)
        _record(conn, op.id, outcome_for(charge.status), external_id=charge.id)
        return outcome_for(charge.status)

    if not op.within_key_window:
        # The key may have been removed. Resending with it could become a "new request",
        # so settle the outcome by reconciliation instead of creating again
        found = gateway.find_charges(op)
        if len(found) != 1:
            _record(conn, op.id, "unknown", error=f"key window expired; {len(found)} match(es); needs review")
            return "unknown"
        _record(conn, op.id, outcome_for(found[0].status), external_id=found[0].id)
        return outcome_for(found[0].status)

    _mark_attempt_started(conn, op.id)
    try:
        charge = gateway.create_charge(op)
    except AmbiguousOutcome as exc:
        _record(conn, op.id, "unknown", error=str(exc))  # not a failure; retry with the same key
        return "unknown"
    except RetryLater as exc:
        _record(conn, op.id, op.status, error=str(exc))
        return op.status
    except Declined as exc:
        _record(conn, op.id, "failed", error=str(exc))
        return "failed"
    _record(conn, op.id, outcome_for(charge.status), external_id=charge.id)
    return outcome_for(charge.status)

What this code guarantees:

  1. The same operation is always sent with the same idempotency key, however often it runs. If the response is lost, the next run sends the same key and Stripe returns the first result.
  2. Unknown stays unknown. An unknown operation with no known external ID is resent with the same key on the next run. If only the response was lost (a connection error), the resend settles it. A 500 can be replayed under the same key (section 5.3), so it may stay unknown — harmless, never a second charge — until the key window passes and the lookup in the third branch settles it or sends it to human review. For one whose external ID is known (it came back processing, say), resending with the same key only replays the first response, so the latest state is read with retrieve instead.
  3. A settled operation never calls the external API again. _record never overwrites a settled state either.
  4. Past the key's lifetime, it never creates. It settles by reconciliation, judging a found PaymentIntent by its status too (found does not mean succeeded), and sends zero or multiple matches for human review.

5.5 Why reconcile past the key's lifetime, and why a human if nothing is found

For an unknown operation past the key's lifetime that Stripe Search can't find, concluding "not charged" and recreating it with a new key is dangerous. Stripe's documentation says data is normally searchable in under a minute but can be delayed during an outage. Not finding something doesn't prove it doesn't exist.

In my judgement, operations that get this far are rare, and the amount is known, so having a person check the Stripe dashboard is safer than accepting the risk of an automatic double charge. The whole mechanism of "periodically compare with the external system and hand drift you can't fix automatically to a person" is covered in Designing reconciliation.


6. Tests: reproduce a lost response

Mimic Stripe's idempotency layer with a fake that returns the first result for the same key. "Succeeded on the other side, but only the response was lost" is reproduced by having the fake record the charge and then raise.

class FakeStripe:
    """A fake of Stripe's idempotency layer: returns the first result for the same key."""

    def __init__(self, first_status="succeeded"):
        self.by_key = {}
        self.charges = {}            # PaymentIntents actually created (the external side effect): id -> [status, op_id]
        self.lose_response_once = False
        self.pruned = False
        self.first_status = first_status

    def create_charge(self, op):
        key = op.idempotency_key
        if key in self.by_key and not self.pruned:
            return self.by_key[key]  # replay the first response as-is (including its status at the time)
        pi = f"pi_{len(self.charges) + 1}"
        self.charges[pi] = [self.first_status, str(op.id)]
        res = ExternalCharge(pi, self.first_status)
        self.by_key[key] = res
        if self.lose_response_once:
            self.lose_response_once = False
            raise AmbiguousOutcome("ReadTimeout")  # it was created on the other side
        return res

    def retrieve_charge(self, external_id):
        return ExternalCharge(external_id, self.charges[external_id][0])

    def find_charges(self, op):
        return [ExternalCharge(pi, st) for pi, (st, opid) in self.charges.items() if opid == str(op.id)]


def test_response_lost_then_retry_with_same_key(conn):
    gw = FakeStripe()
    gw.lose_response_once = True
    op = new_op(conn)
    assert execute_charge(conn, gw, op) == "unknown"
    assert execute_charge(conn, gw, op) == "succeeded"
    assert len(gw.charges) == 1  # one external side effect


def test_processing_is_not_success_and_is_settled_by_retrieve(conn):
    gw = FakeStripe(first_status="processing")
    op = new_op(conn)
    assert execute_charge(conn, gw, op) == "unknown"
    gw.charges["pi_1"][0] = "succeeded"                  # the external side settles later
    assert execute_charge(conn, gw, op) == "succeeded"   # read the latest via retrieve, not a resend
    assert len(gw.charges) == 1

Cases run against a PostgreSQL 18 container and confirmed passing:

TestExpected result
Request an operation twice for the same orderThe same operation ID is returned
Response lost → rerun with the same key1st unknown, 2nd succeeded, one external charge
❌ A new key on every retryTwo external charges (failure reproduced)
unknown past the key's lifetime (key already removed)No create call; settled by lookup as succeeded; still one external charge
Card declinedSettles as failed and never calls the external API again
Returned processing → settles externally later1st unknown, 2nd succeeded via retrieve; created once
Lookup past the key's lifetime finds requires_payment_methodfailed, not success
Lookup past the key's lifetime finds nothingStays unknown, goes to human review
Created three days ago but never sentSent normally rather than reconciled
Connection without autocommitThe connection is outside any transaction (IDLE) during the external call

A fake verifies the branches in your own code. Verify the behaviour of Stripe's idempotency layer itself in Stripe's test environment by sending the same key twice and confirming the same id comes back.


7. Side effects without idempotency keys

Not every external API accepts an idempotency key. Some e-mail sending APIs, for instance, don't support them (others, such as Resend, do: Resend idempotency keys, retries and error classification).

Without an idempotency key, you can't build a guarantee that the business effect happens once. Choose explicitly, per side effect, which failure you'll accept.

Side effectHarm from double executionHarm from a missed executionLean towards
Payment-completed notification e-mailSmall (two e-mails)ModerateAt-least-once (retry when unknown)
ChargeLarge (double charge)Large (uncollected payment)Use an API with idempotency keys; if none exists, make reconciliation mandatory
Password-reset e-mailSmallLarge (can't log in)At-least-once
Reserving stock in an external systemLargeModerateCheck by reconciliation before re-running

The design isn't "give up because exactly-once is impossible"; it's deciding per side effect which way to lean, and expressing that decision in code and monitoring.


8. Common misconceptions

  • "An outbox gives you exactly-once": an outbox guarantees that the business update and the event record are atomic, and that publishing is at-least-once. Duplicate publishing is possible, so consumers must be idempotent.
  • "SQS FIFO means no double processing": FIFO de-duplication keeps duplicate sends within five minutes out of the queue. A consumer that crashes after processing and before deleting receives the message again.
  • "Kafka's exactly-once makes external API calls happen once": Kafka transactions cover reading, processing and writing between Kafka topics.
  • "I added an idempotency key, so it's safe days later": Stripe's keys can be removed once they're at least 24 hours old.
  • "A timeout is a failure": a timeout is unknown. Treat it as a failure and you re-run an operation that actually succeeded.
  • "Idempotency and de-duplication are the same thing": de-duplication means not processing the same message ID twice; idempotency means applying the same operation any number of times doesn't change the result. When the same business fact arrives under a different event ID, de-duplication can't stop it, and you need idempotency on business keys.

Summary

  • Exactly-once has no fixed meaning until you say what (delivery, processing, effect) and within which boundary.
  • Exactly-once guarantees within limited boundaries really exist: a single DB transaction, Kafka transactions (between Kafka topics), SQS FIFO's five-minute send de-duplication. Don't stretch them to side effects in independent external systems.
  • When an external API's response is lost, the outcome is unknown. Don't round it to failed; keep it as unknown.
  • What converges a business effect to one is an operation ID persisted before the side effect, an idempotency key derived from it deterministically, retries with the same key, and reconciliation once the key's lifetime has passed.
  • SDK-generated idempotency keys only protect retries within a single call.
  • For side effects without idempotency keys, decide per side effect whether to accept double execution or a missed execution.

Applying this thinking to the receiving side is implementing webhook idempotency with a transactional inbox; atomicity on the sending side is the transactional outbox; and idempotent consumption from AWS queues is idempotent async processing with SQS + Lambda + EventBridge.

Frequently asked questions

What is the difference between exactly-once and at-least-once?
At-least-once means a message is never lost but may arrive more than once. Exactly-once means the delivery or the effect of processing happens exactly once within some boundary. What matters is what is once and within which boundary; in practice, the most common design combines at-least-once delivery with idempotent processing so that the business effect happens once.
Is exactly-once impossible?
Not across the board. Within a limited boundary — a single database transaction, processing between Kafka topics with Kafka transactions, sends within SQS FIFO's de-duplication interval — there are mechanisms that provide exactly-once semantics. What you can't do is stretch that guarantee to side effects in independent external systems.
If I use Kafka's exactly-once, will external API calls also happen once?
No. Kafka's documentation describes transactional exactly-once in terms of reading, processing and writing Kafka topics, and says that when writing to an external system you need to coordinate the consumer's position with what is actually stored as output. Side effects in external APIs must be protected separately, with idempotency keys or de-duplication on the output side.
With an idempotency key, can I retry days later without double execution?
That isn't guaranteed. With Stripe, keys can be removed after they are at least 24 hours old, and a key reused after removal creates a new request (per the documentation as of October 2026). For retries past the key's lifetime, don't re-run the creation; reconcile against the external state to settle the outcome.

References

友田

友田 陽大

Developer of a METI Minister's Award–winning product. With TypeScript + Python + AWS, I deliver SaaS, industry DX, and production-grade generative AI (RAG) end to end — from requirements to infrastructure and operations — single-handedly.

Double charges, missed webhooks, or churn from failed payments?

Payment-reliability diagnosis and billing architecture, as a technical advisor

"We get a double charge now and then." "We find out about dropped webhooks after the fact." "Failed payments keep turning into cancellations." The cause is rarely one bug — it is usually a structural gap in idempotency, ordering, retry policy, or the state machine. From reading your existing implementation to a receiver that survives redelivery, a recovery flow for failed payments, and where to draw the build-vs-buy line, we decide it together.

Available for both project-based (contract) and advisory engagements. Start with a free 30-minute consult.

Also worth reading