Which four metrics are worth alerting on
Alert fatigue is a design failure. Most teams instrument dozens of signals and still discover outages from customers.
Which four metrics are worth alerting on
Teams instrument extensively and still learn about outages from customers. The reason is that instrumentation and attention are different things, and dashboards do not allocate the second.
A useful monitoring setup has very few alerts, and every one of them is tied to a decision somebody is expected to make.
The four that consistently justify themselves
Availability, measured as user-visible success. Not CPU. Not pod restarts. The percentage of requests that completed with a result the caller could use. This is the number to page on and the number to show a board.
Latency at a chosen percentile, for the slowest meaningful journey. p99 for the operation that matters, not for every endpoint. Percentiles because averages hide the users having a bad time.
Saturation of the resource most likely to constrain you first. Usually a connection pool, a queue depth or a queue consumer count — not CPU. Alerting on CPU in a horizontally scaled service mostly generates alerts that resolve themselves before anyone reads them.
Error rate, split by cause. A single error percentage mixes your bugs with a downstream dependency being unavailable. Separate them, because the first needs a fix and the second needs a decision.
What to remove
Anything where the response is "wait and see" or "restart the service". Those are noise, and noise has a specific cost: it trains the on-call engineer to triage before reading, which means real signals get triaged too.
Also remove alerts that fire on a threshold nobody has ever acted on. If the last three occurrences were all resolved without human intervention, that is a dashboard, not an alert.
A practical test: for each alert, ask what the responder does. If there is no specific action, delete it or demote it to a dashboard.
Dashboards are for humans in meetings
Three views carry most of the value:
- user-visible success rate over time, by journey
- latency percentiles for the top journeys
- current open incidents with elapsed time
Everything else belongs in exploration. Build dashboards for people who will actually look at them, not for completeness.
Instrument the journey, not the request
The failure mode is rich infrastructure telemetry and no end-to-end view. You can see that every service is healthy and still have a checkout flow failing at the third step.
One synthetic check per critical journey, running every few minutes from outside your network, will tell you more about real availability than a hundred endpoint metrics. Make the failure message name the step.
Ownership is part of the signal
Every alert needs a named team and a runbook link. An alert that arrives with no context forces the responder to work out whether it matters before they can decide whether to act — which is the exact behaviour that produces missed incidents.
If the same alert fires weekly for a known reason and nobody can fix it, either fix the cause or remove the alert and put the signal on a dashboard. An alert everybody ignores is worse than no alert, because it reduces trust in all of them.
Retention and access matter more than dashboards
For an incident investigation, access logs over a long window are usually more useful than metrics over a short one. A customer reports a problem from last Tuesday; if you can see what that device did, you have an answer in minutes.
Log enough to reconstruct a request path — correlation identifiers propagated end to end — and keep it long enough to cover the delay between a problem and its discovery. For a bank, that is typically longer than teams assume.
The summary
Four paged metrics, a synthetic check per critical journey, correlation identifiers on every request, and a runbook behind every alert.
That is less instrumentation than most estates have and more than most estates use.
In this article
- observability
- SRE
- on-call
Working on something similar?
These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.
Start a conversationRelated reading
Continue from here
Articles connected to the same delivery problems.
Have a related problem in front of you?
Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.