Most service level objectives are picked from the metric the team already had, and never change anyone's behaviour. Here is how to choose the indicator from the user's experience and wire an error budget to a real decision.
Cardinality is behind almost every Prometheus problem you will have, and it is usually one label added by one well-meaning engineer. How to find it, how to survive retention and HA, and when Thanos or Mimir is actually justified.
OpenTelemetry is not equally mature across traces, metrics and logs, and the adoption plans that fail treat it as one migration. Here is the order that works and where the double-billing months come from.
Log volume rises with traffic, with every new service and with every debugging session somebody forgot to turn off. Structure decides whether they are useful, and six rules decide whether they are affordable.
Most tracing deployments produce disconnected single spans and a bill. Context propagation, tail sampling and a handful of useful attributes are what turn a trace store into the tool you open during an incident.
An on-call rota that gets 40 pages a week is not monitoring, it is a filter that trains humans to dismiss things. Alert on symptoms, alert on burn rate, and delete everything nobody has ever acted on.