Skip to main content
InfromatinTechnologies
Engineering6 min read

Which four metrics are worth alerting on

Alert fatigue is a design failure. Most teams instrument dozens of signals and still discover outages from customers.

Infromatin Technologies

Which four metrics are worth alerting on

Teams instrument extensively and still learn about outages from customers. The reason is that instrumentation and attention are different things, and dashboards do not allocate the second.

A useful monitoring setup has very few alerts, and every one of them is tied to a decision somebody is expected to make.

The four that consistently justify themselves

Availability, measured as user-visible success. Not CPU. Not pod restarts. The percentage of requests that completed with a result the caller could use. This is the number to page on and the number to show a board.

Latency at a chosen percentile, for the slowest meaningful journey. p99 for the operation that matters, not for every endpoint. Percentiles because averages hide the users having a bad time.

Saturation of the resource most likely to constrain you first. Usually a connection pool, a queue depth or a queue consumer count — not CPU. Alerting on CPU in a horizontally scaled service mostly generates alerts that resolve themselves before anyone reads them.

Error rate, split by cause. A single error percentage mixes your bugs with a downstream dependency being unavailable. Separate them, because the first needs a fix and the second needs a decision.

What to remove

Anything where the response is "wait and see" or "restart the service". Those are noise, and noise has a specific cost: it trains the on-call engineer to triage before reading, which means real signals get triaged too.

Also remove alerts that fire on a threshold nobody has ever acted on. If the last three occurrences were all resolved without human intervention, that is a dashboard, not an alert.

A practical test: for each alert, ask what the responder does. If there is no specific action, delete it or demote it to a dashboard.

Dashboards are for humans in meetings

Three views carry most of the value:

  • user-visible success rate over time, by journey
  • latency percentiles for the top journeys
  • current open incidents with elapsed time

Everything else belongs in exploration. Build dashboards for people who will actually look at them, not for completeness.

Instrument the journey, not the request

The failure mode is rich infrastructure telemetry and no end-to-end view. You can see that every service is healthy and still have a checkout flow failing at the third step.

One synthetic check per critical journey, running every few minutes from outside your network, will tell you more about real availability than a hundred endpoint metrics. Make the failure message name the step.

Ownership is part of the signal

Every alert needs a named team and a runbook link. An alert that arrives with no context forces the responder to work out whether it matters before they can decide whether to act — which is the exact behaviour that produces missed incidents.

If the same alert fires weekly for a known reason and nobody can fix it, either fix the cause or remove the alert and put the signal on a dashboard. An alert everybody ignores is worse than no alert, because it reduces trust in all of them.

Retention and access matter more than dashboards

For an incident investigation, access logs over a long window are usually more useful than metrics over a short one. A customer reports a problem from last Tuesday; if you can see what that device did, you have an answer in minutes.

Log enough to reconstruct a request path — correlation identifiers propagated end to end — and keep it long enough to cover the delay between a problem and its discovery. For a bank, that is typically longer than teams assume.

The summary

Four paged metrics, a synthetic check per critical journey, correlation identifiers on every request, and a runbook behind every alert.

That is less instrumentation than most estates have and more than most estates use.

In this article

  • observability
  • SRE
  • on-call

Working on something similar?

These articles come from real engagements. If the problem here sounds familiar, a 30-minute call is usually enough to tell you whether we can help.

Start a conversation

Related reading

Continue from here

Articles connected to the same delivery problems.

Have a related problem in front of you?

Send us the problem in whatever detail you have. A senior engineer replies within one business day, and you will get an honest read on whether we are the right partner for it.

We would like to use Google Analytics to understand how this website is used. No analytics are loaded unless you accept. Your choice is stored for six months.

See our Privacy Policy for details.