Module 7 · 7. Async, Concurrency, Cancellation, and Performance · Lesson 21 of 24
Concurrency, Race Conditions, and Measured Performance
What you will learn
Force a lost update into a reproducible schedule, repair it with two synchronization strategies, and prove that an asynchronous gate returns exactly the permits it acquires. Then collect measurements without mistaking one fast run for a production decision. All examples use in-memory data and short, bounded work.
1. Concurrency, parallelism, and the invariant
Concurrency concerns overlapping operation lifetimes. Parallel execution means work actually executes simultaneously, for example on different processor cores. Concurrent work can be interleaved on one thread; starting tasks does not prove simultaneous CPU execution.
Our invariant is simple: after two completed increments from zero, the count should be two. The broken experiment deliberately violates that invariant. The repair must work regardless of the scheduler's chosen order, even if the machine never runs the workers in parallel.
Minimize shared mutable state and keep synchronization local. Thread safety of individual operations does not automatically protect a larger sequence. Choose a synchronization boundary around the invariant you need to preserve, rather than adding locks wherever a value is read.
Analogy: two edited copies
Two editors both copy the same sheet showing zero. Each changes its own copy to one, then files it as the latest sheet. There were two edits but the final sheet says one. An atomic increment is a single controlled adjustment of the shared number; a lock lets an editor own the whole read-and-update transaction. Neither analogy means all unrelated editing should stop.
2. Make the lost update deterministic
Create a ConcurrencyWalkthrough folder with the four C# files and project file below. Run dotnet run --configuration Release using .NET 10. No package installation is needed.
ConcurrencyWalkthrough/Counters.cs
internal static class Counters
{
public static async Task<int> LostUpdateAsync()
{
int count = 0;
int readers = 0;
var bothRead = new TaskCompletionSource<bool>(TaskCreationOptions.RunContinuationsAsynchronously);
async Task IncrementBrokenAsync()
{
int snapshot = count;
if (Interlocked.Increment(ref readers) == 2) bothRead.SetResult(true);
await bothRead.Task;
count = snapshot + 1;
}
await Task.WhenAll(IncrementBrokenAsync(), IncrementBrokenAsync());
return count;
}
public static async Task<int> CorrectAsync(bool useLock, int workers = 4, int increments = 1000)
{
if (workers <= 0 || increments < 0 || (long)workers * increments > int.MaxValue)
throw new ArgumentOutOfRangeException(nameof(workers));
int count = 0;
var sync = new System.Threading.Lock();
var tasks = new Task[workers];
for (int worker = 0; worker < workers; worker++)
{
tasks[worker] = Task.Run(() =>
{
for (int n = 0; n < increments; n++)
{
if (useLock) { lock (sync) { count++; } }
else Interlocked.Increment(ref count);
}
});
}
await Task.WhenAll(tasks);
return count;
}
}Read the forced schedule
- The first IncrementBrokenAsync saves snapshot = 0, records one reader, and awaits bothRead.
- The second invocation also saves zero before completing bothRead. Therefore both snapshots are fixed at zero before either write is permitted.
- Each invocation writes its own snapshot plus one. Both write one. Awaiting both completions proves the final value is one, with no guess about which write happened last.
This explicitly split read/write is a teaching model of the lost-update mechanism. A free-running counter++ race might happen to produce the correct total in some executions; that would not prove it safe. Our controlled schedule is useful for explaining the defect, not for estimating its production frequency.
Choose the repair by the required operation
Interlocked.Increment performs one atomic increment on the target variable and returns the updated value. Use it when that single-variable operation expresses the required invariant.
A lock excludes competing entrants using the same lock instance for its protected block. Keep that synchronous block short. C# does not permit await inside a lock body. A dedicated System.Threading.Lock is supported by .NET 10.
CorrectAsync uses four worker tasks, each adding 1,000. Every shared update follows the selected strategy, and no final result is read until all workers finish. Both strategies must yield 4,000. This is CPU work, so Task.Run intentionally schedules worker delegates; it does not wrap naturally asynchronous I/O. The check validates totals, not the number of physical threads.
If your rule is “reserve an item only when stock remains, then decrement stock,” protecting just the decrement is insufficient. The check and decrement must participate in one valid atomic operation or one shared critical section. A thread-safe dictionary does not turn an arbitrary check-then-update sequence into one transaction.
3. Bound asynchronous work without over-releasing
SemaphoreSlim counts permits. WaitAsync can wait asynchronously and observe cancellation; Release returns permits. Release only after successful acquisition. Use the same gate for all work sharing that limit. Do not assume FIFO ordering, and dispose the gate only after its users finish.
ConcurrencyWalkthrough/AsyncGate.cs
internal sealed class AsyncGate : IDisposable
{
private readonly SemaphoreSlim _slots;
public AsyncGate(int capacity)
{
if (capacity <= 0) throw new ArgumentOutOfRangeException(nameof(capacity));
_slots = new SemaphoreSlim(capacity, capacity);
}
public int AvailableSlots => _slots.CurrentCount;
public async Task<int> RunAsync(Func<CancellationToken, Task<int>> work, CancellationToken token)
{
await _slots.WaitAsync(token);
try
{
return await work(token);
}
finally
{
_slots.Release();
}
}
public void Dispose() => _slots.Dispose();
}Why the try starts after the wait
There are two different paths. If the wait is canceled before acquiring a permit, this invocation owns nothing and must release nothing. If work is reached, exactly one permit was acquired, so the finally returns it whether work succeeds, fails synchronously, or is canceled. Moving the wait inside this try would allow a canceled waiter to release a permit it never owned.
Capacity one can protect one asynchronous critical section. Capacity two permits two owners at once; it bounds resource use but does not make their shared data mutually exclusive. AvailableSlots is a diagnostic observation in our tests, not an admission test: checking it and then acting would create another check-then-act race.
ConcurrencyWalkthrough/GateDemo.cs
internal static class GateDemo
{
public static async Task ShowAsync(TextWriter output)
{
using var gate = new AsyncGate(2);
var firstReply = new TaskCompletionSource<int>(TaskCreationOptions.RunContinuationsAsynchronously);
var secondReply = new TaskCompletionSource<int>(TaskCreationOptions.RunContinuationsAsynchronously);
int started = 0;
Task<int> first = gate.RunAsync(token =>
{
Interlocked.Increment(ref started);
return firstReply.Task.WaitAsync(token);
}, CancellationToken.None);
Task<int> second = gate.RunAsync(token =>
{
Interlocked.Increment(ref started);
return secondReply.Task.WaitAsync(token);
}, CancellationToken.None);
Task<int> third = gate.RunAsync(_ =>
{
Interlocked.Increment(ref started);
return Task.FromResult(30);
}, CancellationToken.None);
output.WriteLine($"started before release: {started}");
firstReply.SetResult(10);
secondReply.SetResult(20);
int[] results = await Task.WhenAll(first, second, third);
output.WriteLine($"results: {string.Join(", ", results)}");
output.WriteLine($"slots after completion: {gate.AvailableSlots}");
}
}Before either reply is supplied, first and second are inside the gate, and third is waiting for a permit. That is why started before release is two. Only the controller can complete the two reply signals, and it does so after printing that count. After awaiting all three results, the gate again has two permits. There is no assertion about which waiting request would win among several waiters.
A request deadline should cover the time spent waiting for admission as well as the work inside the gate. Pass the same request token through both steps, and measure end-to-end latency before the gate if queueing is part of the user experience. The earlier budget lesson supplies the shared deadline policy.
A semaphore does not bound the number of tasks waiting outside it. Creating a million waiting tasks can still consume excessive memory. Nor does a concurrency limit specify requests per second.
For an owned producer/consumer queue, a bounded channel can apply backpressure when full. Its configured full-mode policy determines whether producers wait or items are dropped.
The original payment gateway sketch is replaced with safe in-memory replies and capacity two so the entire behavior is visible. No money is moved. A production capacity is chosen from dependency limits and measurements, not copied from this classroom constant.
ConcurrencyWalkthrough/Program.cs
internal static class Program
{
public static async Task Main()
{
Console.WriteLine($"forced lost update: {await Counters.LostUpdateAsync()}");
Console.WriteLine($"interlocked total: {await Counters.CorrectAsync(useLock: false)}");
Console.WriteLine($"lock total: {await Counters.CorrectAsync(useLock: true)}");
await GateDemo.ShowAsync(Console.Out);
}
}ConcurrencyWalkthrough.csproj
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<OutputType>Exe</OutputType>
<TargetFramework>net10.0</TargetFramework>
<ImplicitUsings>enable</ImplicitUsings>
<Nullable>enable</Nullable>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
</PropertyGroup>
</Project>Expected output
forced lost update: 1 interlocked total: 4000 lock total: 4000 started before release: 2 results: 10, 20, 30 slots after completion: 2
4. Worked correctness exercises
Exercise A: cancel a queued caller
A capacity-one gate is occupied by an unfinished reply. A second invocation waits and is canceled before it enters. What should its work counter and the gate's available count show?
Solution: Its work must never have started, and available count remains zero while the original owner is active. After that owner finishes, the count becomes one. The semantic harness establishes entry using a signal, then tests precisely this sequence. It does not sleep and hope the first invocation entered.
Exercise B: fail after entry
The work delegate throws before returning a task. Is a permit leaked? What happens if a token-aware wait is canceled inside work instead?
Solution: Both paths are inside the try, so the finally returns the acquired permit. The harness checks the original synchronous exception identity, an in-work cancellation, and a successful subsequent call. Its readiness signal remains unfinished after a canceled wait and is explicitly completed afterward; canceling a wait did not abort its underlying producer.
Exercise C: remove shared writes
Each worker only needs to count its own 1,000 independent items. Must every iteration update the same shared counter?
Solution: No. Each worker can keep a private count and return it; combine the returned counts after all finish. That reduces coordination. It changes when a live global total becomes available, so it is suitable only if the requirement permits an end-of-batch total.
5. Measure a question, not a hunch
First preserve correctness. Then state the workload, hardware/runtime, concurrency level, warmup policy, repetitions and metric that will decide the change. Separate a deliberately contended counter microprobe from a service benchmark with queueing, network dependencies and failures. Their results answer different questions.
Stopwatch measures elapsed duration. Timing an operation does not identify its bottleneck or make one sample representative.
GetTotalAllocatedBytes counts process-wide managed allocations, excluding native allocations. Precise collection has overhead; its delta is not a per-operation or per-thread allocation attribution.
The optional CounterBenchmark project links the exact Counters.cs you just ran. Put it beside ConcurrencyWalkthrough, preserving the folder structure. It warms both strategies, alternates their measurement order across five rounds, verifies every total, and prints results after measurement. Run dotnet run --project CounterBenchmark --configuration Release from their parent folder. No debugger or package is needed.
CounterBenchmark/CounterBenchmark.csproj
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<OutputType>Exe</OutputType>
<TargetFramework>net10.0</TargetFramework>
<ImplicitUsings>enable</ImplicitUsings>
<Nullable>enable</Nullable>
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
</PropertyGroup>
<ItemGroup>
<Compile Include="../ConcurrencyWalkthrough/Counters.cs" Link="Learner/Counters.cs" />
</ItemGroup>
</Project>CounterBenchmark/Program.cs
using System.Diagnostics;
using System.Text.Json;
internal sealed record Trial(string Strategy, int Round, int Total, double ElapsedMs, long AllocatedBytes);
internal static class Program
{
public static async Task Main()
{
const int workers = 4;
const int increments = 10000;
const int expected = workers * increments;
for (int warmup = 0; warmup < 3; warmup++)
{
if (await Counters.CorrectAsync(false, workers, increments) != expected ||
await Counters.CorrectAsync(true, workers, increments) != expected)
throw new InvalidOperationException("Warmup correctness failure.");
}
var trials = new List<Trial>();
for (int round = 1; round <= 5; round++)
{
bool[] order = round % 2 == 0 ? [true, false] : [false, true];
foreach (bool useLock in order)
{
long before = GC.GetTotalAllocatedBytes(precise: true);
long start = Stopwatch.GetTimestamp();
int total = await Counters.CorrectAsync(useLock, workers, increments);
double elapsed = Stopwatch.GetElapsedTime(start).TotalMilliseconds;
long allocated = GC.GetTotalAllocatedBytes(precise: true) - before;
if (total != expected || !double.IsFinite(elapsed) || elapsed < 0 || allocated < 0)
throw new InvalidOperationException("Measurement invariant failed.");
trials.Add(new Trial(useLock ? "lock" : "interlocked", round, total, elapsed, allocated));
}
}
// Print only after measurement so console I/O is outside measured regions.
Console.WriteLine(JsonSerializer.Serialize(trials));
}
}What the measurement output proves
The output is a JSON array of ten trials. Each trial contains Strategy, Round, Total, ElapsedMs and AllocatedBytes. Total must be 40,000 every time; there must be five lock and five interlocked trials. ElapsedMs and AllocatedBytes vary with the machine and runtime. There is deliberately no exact expected duration and no required winner.
This probe measures task setup, scheduling, contention and joining as part of the batch. Warmup reduces some startup effects but does not remove every source of noise or later JIT changes. Process-wide allocation deltas include task/runtime infrastructure and can include unrelated in-process work; they are not isolated payload allocation per strategy. The tiny workload and five rounds are a learning scaffold; a production decision requires enough representative samples and measurement of the actual workload. Retain the full distribution rather than selecting the fastest run.
One observed run of the measurement probe
The ten trials below were observed in a Release build on Windows x64 with .NET SDK 10.0.401 and runtime 10.0.12. Each batch used four worker tasks with 10,000 increments each. There were three warmup batches per strategy, followed by five measured rounds with strategy order alternating. Rows retain the program's emitted order and values.
| Trial | Strategy | Round | Total | Elapsed ms | Process-wide managed bytes |
|---|---|---|---|---|---|
| 1 | interlocked | 1 | 40000 | 0.8753 | 648 |
| 2 | lock | 1 | 40000 | 4.8007 | 704 |
| 3 | lock | 2 | 40000 | 4.2618 | 704 |
| 4 | interlocked | 2 | 40000 | 0.7585 | 648 |
| 5 | interlocked | 3 | 40000 | 0.694 | 648 |
| 6 | lock | 3 | 40000 | 4.2032 | 704 |
| 7 | lock | 4 | 40000 | 3.2987 | 816 |
| 8 | interlocked | 4 | 40000 | 0.5905 | 648 |
| 9 | interlocked | 5 | 40000 | 0.3673 | 648 |
| 10 | lock | 5 | 40000 | 2.3043 | 704 |
Every trial reached the required 40,000 total. These are observed batch measurements, including scheduling and synchronization, not a promise for another machine or workload. CPU model, core count and background-load details were not recorded with this example. Interpret the allocation column using the process-wide measurement limits described above. Retain these samples as evidence of this run; they do not establish a universal strategy choice.
Worked performance decision
Consider illustrative service measurements, not results from our counter program. Baseline: 400 successful requests per second and 160 ms p95 latency. Candidate: 500 successful requests per second and 400 ms p95 latency. The product requires p95 at most 200 ms. Which would you ship?
Solution: The candidate fails the latency requirement despite higher throughput. Keep the baseline while investigating the tail. For 1,000 successful requests over two seconds, throughput is 500 per second; latency measures individual request duration and is not its reciprocal. A p95 marks the 95th percentile of the recorded latency distribution, not the maximum. Include failures, timeouts, queue time and allocation/GC behavior in the comparison instead of optimizing only successful fast requests.
Blocked thread-pool workers can leave queued work unable to proceed. Low CPU use alone does not rule out starvation; inspect thread-pool behavior and blocking stacks rather than assuming every slow request needs more CPU.
Interview answer framework
Name the invariant, show one bad interleaving, and explain how the selected boundary prevents it. Distinguish atomicity of one update from a transaction involving several values. For asynchronous capacity, say who owns each permit and what happens before acquisition, during work and at shutdown. For performance, present representative evidence and the acceptance criterion before proposing a larger concurrency limit.