Module 3 · API Resilience, Backpressure, and Observability · Lesson 5 of 8
Rate Limiting, Backpressure, and Overload Protection
Why overload protection is a design feature
A healthy service does not accept unlimited work. It admits requests at a rate the downstream database, remote APIs, and worker queues can sustain. Without admission control, latency rises, timeouts trigger retries, and retries amplify load until the service collapses.
Rate limiting answers how many requests may enter. Backpressure answers what producers should do when consumers cannot keep up.
Choose the limiter that matches the constraint
- Concurrency limiter: caps simultaneous in-flight operations. Use it when memory, database connections, or another scarce concurrent resource is the bottleneck.
- Fixed window: simple request quota per interval. It can allow a burst at a window boundary.
- Sliding window: smooths the boundary effect by dividing time into segments.
- Token bucket: permits controlled bursts while enforcing an average refill rate.
- Partitioned limiter: applies independent limits per tenant, API key, user, route, or workload class.
A queue is not free capacity. Every queued request consumes time and memory and may expire before execution. Keep queue limits small, propagate cancellation, and reject excess work explicitly.
ASP.NET Core fixed-window policy
The following complete minimal API protects a write endpoint and returns HTTP 429 when capacity is exhausted.
using System.Threading.RateLimiting;
using Microsoft.AspNetCore.RateLimiting;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddRateLimiter(options =>
{
options.RejectionStatusCode = StatusCodes.Status429TooManyRequests;
options.AddFixedWindowLimiter("write-api", limiter =>
{
limiter.PermitLimit = 50;
limiter.Window = TimeSpan.FromSeconds(10);
limiter.QueueProcessingOrder = QueueProcessingOrder.OldestFirst;
limiter.QueueLimit = 10;
limiter.AutoReplenishment = true;
});
options.OnRejected = async (context, cancellationToken) =>
{
context.HttpContext.Response.Headers.RetryAfter = "10";
await context.HttpContext.Response.WriteAsJsonAsync(
new { error = "capacity_exhausted" },
cancellationToken);
};
});
var app = builder.Build();
app.UseRateLimiter();
app.MapPost("/orders", (CreateOrder command) =>
Results.Accepted(value: new { command.OrderId }))
.RequireRateLimiting("write-api");
app.Run();
public sealed record CreateOrder(Guid OrderId, decimal Amount);The numbers above are examples, not universal recommendations. Derive limits from load tests, downstream capacity, and service-level objectives.
Backpressure with a bounded channel
A bounded channel makes the overload decision explicit. FullMode = Wait causes producers to wait for capacity, and the request cancellation token prevents abandoned requests from occupying the queue.
using System.Threading.Channels;
var queue = Channel.CreateBounded<WorkItem>(
new BoundedChannelOptions(capacity: 200)
{
FullMode = BoundedChannelFullMode.Wait,
SingleReader = true,
SingleWriter = false
});
app.MapPost("/reports", async (
WorkItem work,
CancellationToken cancellationToken) =>
{
await queue.Writer.WriteAsync(work, cancellationToken);
return Results.Accepted($"/reports/{work.Id}");
});
public sealed record WorkItem(Guid Id, string ReportType);In a real service, register the channel as a singleton and consume it from a BackgroundService. Define shutdown behavior: stop accepting new work, complete the writer, drain within a budget, and persist or requeue anything that cannot finish safely.
Operational signals
Monitor these together:
- accepted and rejected requests by policy and partition;
- current concurrency and queue depth;
- queue wait duration, not only processing duration;
- end-to-end latency percentiles and timeout count;
- downstream saturation such as database pool usage;
- retry volume and duplicate-work detection;
- background consumer throughput and oldest-item age.
A falling error rate is not enough if queue age is still growing.
Failure modes and design decisions
- An unbounded queue converts traffic spikes into memory pressure and stale work.
- A large queue hides overload and pushes failures beyond the caller's timeout.
- A global tenant limit lets one noisy tenant consume all capacity.
- Returning 429 without a stable retry contract can create synchronized retry storms.
- Retrying non-idempotent writes can duplicate side effects.
- Limiting only HTTP traffic misses scheduled jobs and message consumers that share the same database.
Senior interview checkpoint
Be ready to explain:
- why concurrency and request-rate limits solve different problems;
- how you would partition capacity across free and paid tenants;
- when to reject immediately versus wait briefly;
- how cancellation travels from HTTP to a bounded queue;
- which metrics prove that the selected limit is protecting the dependency;
- how idempotency keys make a retryable write safer.
Hands-on exercise
Design an order API with a 100-connection database pool and three workload classes: interactive reads, writes, and exports. Define separate admission policies, queue limits, cancellation behavior, and dashboard metrics. Explain how you would tune them from load-test evidence rather than guesswork.