Skip to content
Search lessons, topics, tests…
Esc

    ↑ ↓ moveEnter openEsc close

    Guided course · Testing and Production Engineering

    Production-Grade .NET Backend Interview Path: back to the course

    Module 2 · Data Access and Production Reliability · Lesson 4 of 8

    Idempotency, Retries, and Transaction Boundaries

    Why retries create correctness problems

    Distributed systems fail in ambiguous ways. A client can time out after the server commits, a broker can redeliver a message, or a process can crash after charging a card but before updating its database. Retrying improves availability only when repeated execution is safe.

    Idempotency definition

    An operation is idempotent when applying the same logical request multiple times produces the same intended effect as applying it once. HTTP GET, PUT, and DELETE are defined with idempotent semantics, but an implementation can still violate them. POST is not inherently non-idempotent; it can be protected with an idempotency key.

    Choosing the key

    The client should generate a stable key for one business command, such as CreatePayment for one checkout attempt. A random key generated on every retry defeats deduplication. Scope the key by tenant and operation to prevent collisions. Store a hash of the canonical request payload so that reusing a key with different input returns a conflict instead of silently replaying the wrong response.

    Persistence model

    A robust idempotency record can contain tenant ID, operation, key, request hash, status, response code, response body or resource ID, created time, and expiry. Enforce uniqueness in the database.

    SQL
    CREATE UNIQUE INDEX UX_Idempotency
    ON IdempotencyRecords(TenantId, Operation, IdempotencyKey);

    Do not rely on "check then insert" without a unique constraint. Two concurrent requests can both see no record and both execute. The database constraint is the final concurrency guard.

    Atomic local transaction

    When the business change and idempotency record share one database, write them in one transaction. The following service-method fragment assumes your EF Core context and domain types. CreateValidatedPayment constructs a local payment record; it must not call a payment provider. Generate stablePaymentId before any retry delegate and retain it across retries.

    C#
    await using var transaction = await db.Database.BeginTransactionAsync(token);
    
    var payment = CreateValidatedPayment(request, stablePaymentId);
    db.Payments.Add(payment);
    db.IdempotencyRecords.Add(IdempotencyRecord.Completed(
        tenantId, operation, key, requestHash, payment.Id));
    
    await db.SaveChangesAsync(token);
    await transaction.CommitAsync(token);

    Handle the unique-constraint race by loading the existing record and returning its recorded result. Distinguish a duplicate request from an unrelated database error.

    In-progress records and leases

    Long operations may first create an InProgress record. Concurrent duplicates can receive a documented 409 Conflict, or 202 Accepted with a status URL when the operation is still processing. Specify a polling/backoff contract. 425 Too Early concerns replay risk from HTTP early data; it is not a general-purpose “job still running” response. If a worker can crash, the record needs a lease or recoverable state; otherwise it may remain in progress forever. A takeover policy must be conservative because the original external side effect may have happened.

    External side effects

    A local transaction cannot atomically commit your database and a third-party payment. Use the provider's idempotency key where available. Record the provider operation ID. Reconcile uncertain outcomes by querying the provider rather than issuing another charge with a new key.

    For publishing events after a database change, use the transactional outbox. In the same database transaction, save the aggregate change and an outbox message. A separate dispatcher sends the message and marks it sent. Delivery is at least once, so consumers still need idempotent handling.

    Retry classification

    Retry only transient failures. Good candidates include a connection reset, a 429 response with Retry-After, or a short service-unavailable period. Do not retry validation failures, authentication failures, payload errors, or deterministic constraint violations.

    Use exponential backoff with jitter so many clients do not retry at the same instant. Bound attempts by the end-to-end deadline. Honor server retry hints. Log the attempt number and delay, but avoid one error log per retry plus another final error that inflates alert volume.

    A conceptual retry policy:

    1. Attempt 1 immediately.
    2. Attempt 2 after a small randomized delay.
    3. Attempt 3 after a larger randomized delay.
    4. Stop when the deadline would be exceeded.
    5. Return or route the command for controlled recovery.

    Retries and transactions

    A database execution strategy may replay an entire transaction when a transient error occurs. Every action inside the retry delegate must be replay-safe. Never call a non-idempotent external API inside a database retry block. Generate stable IDs outside the delegate when repeated inserts must represent the same entity.

    Isolation and concurrency

    Idempotency is not a substitute for domain concurrency rules. Two different keys can still try to reserve the last item. Use optimistic concurrency tokens, conditional updates, or appropriate isolation. For example:

    SQL
    UPDATE Inventory
    SET Available = Available - @quantity
    WHERE ProductId = @productId
    AND Available >= @quantity;

    Check the affected row count. This prevents overselling without globally serializing every request.

    Message consumers

    A broker commonly provides at-least-once delivery. Store a processed-message ID with the business update in one transaction. Acknowledge only after commit. If processing fails before commit, redelivery is safe. If acknowledgement fails after commit, the deduplication record makes the next delivery a no-op or replays the recorded result.

    Ordering complicates the model. Deduplication stops repeated messages but does not stop an older distinct message from arriving later. Use aggregate sequence numbers, version checks, or partitioning when order matters.

    Cancellation and ambiguous completion

    If the caller cancels while commit is in progress, do not assume the transaction failed. The server should finish or reconcile the critical section according to its design. The client should retry with the same idempotency key. The API can then return the completed response or a known in-progress state.

    Testing the hard cases

    1. Send two concurrent requests with the same key; assert one business effect.
    2. Reuse the key with a different payload; assert conflict.
    3. Simulate a timeout after commit; retry and assert the original response.
    4. Crash the worker after external success but before local completion; reconcile.
    5. Redeliver the same message; assert no duplicate state.
    6. Inject transient failures and verify bounded backoff.
    7. Verify non-transient failures are not retried.
    8. Expire old idempotency records according to business replay windows.

    Interview scenario

    Design a CreateOrder endpoint that reserves inventory, charges payment, and emits OrderCreated. A credible answer refuses to pretend one distributed ACID transaction exists. It uses a stable order ID and idempotency key, an inventory concurrency check, provider-side payment idempotency, durable state transitions, and an outbox. It explains compensation or reconciliation for partial progress and exposes a status endpoint for uncertain client outcomes.

    Operational signals

    Track duplicate-hit rate, in-progress age, conflict count, retry attempts, exhausted retries, outbox lag, reconciliation backlog, and provider mismatch count. Alert on stuck in-progress records and growing outbox age rather than on expected duplicate requests.

    Answer pattern

    Begin with the failure window. Name the stable command identity, the database uniqueness guarantee, transaction boundary, external idempotency mechanism, retry classification, and reconciliation path. This demonstrates that reliability is designed around ambiguity, not added later with a retry loop.

    Practice

    Sign in to mark lessons done and keep your place in the course.Sign in