Module 2 · Concurrency, Failure Semantics, and Deterministic Tests · Lesson 3 of 4
Bounded Fan-Out and Failure Semantics
Concurrency improves latency only until a dependency saturates. Beyond that point, more parallel work increases queueing, timeouts, memory use, and rate-limit failures.
A bounded fan-out helper
public static async Task<TResult[]> SelectBoundedAsync<T, TResult>(
IEnumerable<T> source,
int maxConcurrency,
Func<T, CancellationToken, Task<TResult>> operation,
CancellationToken cancellationToken)
{
ArgumentOutOfRangeException.ThrowIfLessThan(maxConcurrency, 1);
using var gate = new SemaphoreSlim(maxConcurrency);
var tasks = source.Select(async item =>
{
await gate.WaitAsync(cancellationToken);
try
{
return await operation(item, cancellationToken);
}
finally
{
gate.Release();
}
});
return await Task.WhenAll(tasks);
}This limits in-flight operations, but it still creates one task per source item. For very large or streaming inputs, use a bounded channel, worker pool, or Parallel.ForEachAsync so queued work is bounded too.
Choose the failure contract first
Task.WhenAll is appropriate for all-or-nothing aggregation. Awaiting it throws when any task fails, but all supplied tasks are allowed to reach a terminal state. Inspect individual tasks or preserve structured results when callers need every failure.
For partial success, make failure explicit:
public sealed record ItemResult<T>(T? Value, Exception? Error)
{
public bool Succeeded => Error is null;
}Catch per item, return ItemResult<T>, and let the boundary decide whether a partial response is useful. Do not silently swallow an exception and return a default value that looks successful.
Select a defensible limit
Start with dependency limits and measurements:
- connection-pool size and expected sharing;
- documented remote API quota;
- p95 latency and timeout budget;
- CPU versus I/O workload;
- memory held per in-flight item;
- retry amplification during incidents.
A per-request limit may still overload a dependency when many requests arrive together. Combine local limits with a service-wide bulkhead, queue, or rate limiter where necessary.
Cancellation and cleanup
Pass the cancellation token both while waiting for a permit and during the operation. Always release the permit in finally. Dispose the semaphore only after all tasks have finished using it.
Interview design exercise
Given 10,000 customer IDs and an API quota of 50 requests per second, ask what must be bounded: simultaneous calls, request rate, queued IDs, retries, and total operation time. “Use Task.WhenAll” answers none of those policy questions.
The mature answer separates throughput, concurrency, rate, and failure semantics, then makes each limit observable.