Skip to content
Search lessons, topics, tests…
Esc

    ↑ ↓ moveEnter openEsc close

    Guided course · Testing and Production Engineering

    Production-Grade .NET Backend Interview Path: back to the course

    Module 3 · API Resilience, Backpressure, and Observability · Lesson 6 of 8

    Timeouts, Retries, and Circuit Breaker Composition

    Learning outcome

    By the end of this lesson, you can compose timeouts, retries, and circuit breakers without multiplying traffic or exceeding the caller's deadline. You will also be able to explain idempotency, cancellation, observability, and policy ownership in a senior .NET interview.

    Start with one end-to-end budget

    A resilient request begins with a deadline, not with a retry count. If the public API has a 3-second objective, reserve time for routing, serialization, and the final response. Give each dependency a smaller attempt timeout and ensure the entire retry sequence fits inside the total request timeout.

    Timeouts bound waiting; they do not make an operation succeed. Always propagate the caller's CancellationToken so abandoned work releases sockets, database connections, and CPU.

    Compose policies deliberately

    A practical outbound HTTP order is:

    1. A total-request timeout bounds the complete operation.
    2. A retry strategy handles a small set of transient failures.
    3. A circuit breaker stops sending work to a dependency that is persistently unhealthy.
    4. A per-attempt timeout prevents any single try from consuming the whole budget.

    The precise nesting depends on the library, but the invariants do not: retries must be bounded by the outer deadline, and every attempt must be cancellable.

    C#
    builder.Services
        .AddHttpClient<InventoryClient>(client =>
        {
            client.BaseAddress = new Uri(
                builder.Configuration["Inventory:BaseUrl"]!);
        })
        .AddStandardResilienceHandler(options =>
        {
            options.TotalRequestTimeout.Timeout = TimeSpan.FromSeconds(3);
            options.AttemptTimeout.Timeout = TimeSpan.FromSeconds(1);
            options.Retry.MaxRetryAttempts = 2;
            options.CircuitBreaker.SamplingDuration = TimeSpan.FromSeconds(30);
            options.CircuitBreaker.MinimumThroughput = 20;
            options.CircuitBreaker.FailureRatio = 0.5;
        });

    Retry only safe operations

    Retry temporary connection failures, selected 5xx responses, HTTP 408, and HTTP 429 when the server provides retry guidance. Do not retry validation errors, authentication failures, or deterministic business conflicts.

    For writes, a timeout means the outcome is unknown—not that the server did nothing. Use an idempotency key or operation identifier persisted with the result before retrying a payment, email, inventory decrement, or other side effect.

    C#
    public async Task<InventoryItem?> GetAsync(
        string sku,
        CancellationToken cancellationToken)
    {
        using var response = await httpClient.GetAsync(
            $"inventory/{Uri.EscapeDataString(sku)}",
            cancellationToken);
    
        if (response.StatusCode == HttpStatusCode.NotFound)
            return null;
    
        response.EnsureSuccessStatusCode();
        return await response.Content.ReadFromJsonAsync<InventoryItem>(
            cancellationToken: cancellationToken);
    }

    Scope the circuit breaker correctly

    A breaker should represent one dependency boundary. Avoid one global circuit for unrelated services: a failing recommendations endpoint must not block inventory or payments. For multi-tenant systems, consider whether one tenant can poison shared breaker statistics and whether partitions are needed.

    The open state fails fast. After the break duration, a limited probe enters the half-open state. Recovery must be observable; do not silently return stale data unless the product explicitly permits it.

    Avoid retry amplification

    If the client, gateway, application, and SDK each retry twice, a single user action can fan out into many dependency calls. Assign one owner for retries, document downstream behavior, and expose the final attempt count in telemetry.

    Use exponential backoff with jitter. A small random delay prevents every instance from retrying at the same moment after a shared outage.

    Production telemetry

    Record dependency name, operation, attempt number, final outcome, duration, timeout source, breaker state transition, and whether a fallback was served. Keep recovered retries separate from final failures: a healthy success rate can hide a dependency that succeeds only after repeated attempts.

    Alert on sustained breaker openings, rising retry volume, tail latency, and request cancellation. Never log access tokens, connection strings, or complete sensitive payloads.

    Failure-injection exercise

    Test four cases against a controlled dependency stub:

    • slow success just inside the attempt timeout;
    • a timeout followed by success on a safe read;
    • repeated 503 responses that open the breaker;
    • caller cancellation while a retry delay is pending.

    For each case, verify the total duration, number of attempts, emitted telemetry, and final HTTP response. The test passes only if the total work stays inside the caller's budget.

    Interview drill

    • Why is a timeout not a retry policy?
    • When is retrying a POST safe?
    • How can retries at multiple layers amplify an outage?
    • What should determine a circuit breaker's partition key?
    • Why must cancellation flow through retry delays as well as HTTP calls?
    • Which metrics prove that a fallback is helping rather than hiding failure?

    Revision checklist

    • I can draw the total deadline and per-attempt budget.
    • I can classify transient versus deterministic failures.
    • I require idempotency before retrying side effects.
    • I keep circuit breakers scoped to meaningful dependency boundaries.
    • I prevent retry multiplication across layers.
    • I test recovery, cancellation, and observability under injected faults.

    Practice

    Sign in to mark lessons done and keep your place in the course.Sign in