Module 4 · Testing, Incident Diagnosis, and System Design Review · Lesson 7 of 8
Production Incident Diagnosis with Metrics, Traces, and Logs
The interview skill: turn symptoms into evidence
Senior backend interviews rarely reward a list of monitoring tools. They test whether you can reduce uncertainty during an incident without making the system less safe.
Use this loop:
- state the user-visible symptom and affected scope;
- check whether the change is traffic-related, dependency-related, or release-related;
- form one falsifiable hypothesis;
- choose the cheapest signal that can disprove it;
- mitigate customer impact;
- preserve evidence for the deeper fix.
Start with the service-level symptom
Translate “the API is slow” into measurable boundaries:
- which routes, tenants, regions, and versions are affected;
- when the change began;
- latency percentile, error rate, and throughput before and after;
- whether queued or background work is also delayed;
- whether the dependency time is included in the observed latency.
Do not start by searching logs for arbitrary errors. First decide what a healthy request should look like.
Metrics: detect and bound the problem
Metrics are best for trends and comparisons. A useful first dashboard includes:
- request rate, failures, and p50/p95/p99 latency;
- in-flight requests and rate-limit rejections;
- thread-pool queue length and process CPU;
- managed heap size, allocation rate, and garbage-collection pauses;
- database connection-pool usage and command duration;
- outbound dependency duration and error rate;
- worker queue depth, dequeue rate, and oldest-item age.
Look for correlated changes. High latency with low CPU and an exhausted connection pool points toward waiting, not computation.
Traces: follow one slow path
A distributed trace should answer:
- where the request spent its time;
- which dependency or internal stage dominated;
- whether retries created repeated spans;
- whether baggage or correlation identifiers crossed service boundaries;
- whether the same operation crossed an asynchronous handoff.
Compare a slow trace with a healthy trace for the same route and tenant. One trace is an example, not proof of frequency, so return to metrics to estimate impact.
Logs: explain the decision and failure
Structured logs should add facts that metrics and traces do not carry. Prefer stable fields such as operation name, tenant identifier, order identifier, attempt number, outcome, duration, and trace identifier.
Avoid:
- interpolated messages that cannot be grouped reliably;
- secrets, access tokens, or full personal data;
- logging the same exception at every layer;
- unbounded high-cardinality fields in metric labels;
- success logs for every hot-path operation when traces already provide that detail.
A worked incident: latency after a deployment
Symptom: p95 order-write latency rises from 300 ms to 4 s after a release, while throughput and CPU remain stable.
Hypothesis 1: the release introduced slower SQL.
Disproof signal: traces show SQL duration is unchanged, but requests wait before opening a database span.
Hypothesis 2: requests are waiting for a database connection.
Evidence: connection-pool usage is pinned at its limit; logs show timeouts acquiring connections; memory and CPU are normal.
Next question: why are connections held longer?
Compare code paths and traces. A new external HTTP call was placed inside a transaction, extending the lifetime of the connection.
Mitigation: roll back or feature-disable the new path. Do not increase the connection pool blindly; that may move saturation to the database.
Permanent fix: keep remote calls outside the transaction, shorten transaction scope, add an acquisition-time metric, and load-test the repaired flow.
Safe mitigation hierarchy
- Stop or roll back the triggering change when evidence is strong.
- Shed optional work and protect critical routes.
- Reduce expensive fan-out, retries, or background concurrency.
- Apply a temporary capacity increase only if the dependency can absorb it.
- Communicate customer impact and recovery criteria.
Every mitigation needs an owner, expiry condition, and rollback plan.
Common interview mistakes
- claiming root cause from temporal correlation alone;
- restarting before capturing useful state;
- increasing timeouts when the dependency is saturated;
- retrying all failures, including permanent or non-idempotent operations;
- averaging latency and hiding tail behavior;
- using trace IDs as metric labels and causing cardinality explosion;
- monitoring queue depth without oldest-item age.
Senior interview questions
- How do you distinguish thread-pool starvation from high CPU?
- Why can p50 stay healthy while customers still report severe slowness?
- Which signals reveal a retry storm?
- What evidence would justify increasing a database connection pool?
- How do you trace a request after it is handed to a background worker?
- Which mitigation would you choose if one tenant causes most saturation?
Practice scenario
A report-export queue grows continuously, consumer CPU is 20%, no errors are logged, and the oldest message is 45 minutes old. Build three competing hypotheses, name one signal that can disprove each, and propose a reversible mitigation. Your answer should discuss downstream throttling, consumer concurrency, poison messages, and deployment changes rather than assuming a single cause.
Review checklist
A strong diagnosis separates symptom, hypothesis, evidence, mitigation, and permanent fix. It also explains what would change your mind.