Problem
Systems are black boxes. Dashboards show symptoms, not causes. Alerts are noisy. Debugging takes hours. Need systematic observability.
Also known as: observability, monitoring, logging, metrics, tracing, sli-slo
Build comprehensive observability: three pillars + SLI/SLO, alerting, debugging workflows, cost control.
Systems are black boxes. Dashboards show symptoms, not causes. Alerts are noisy. Debugging takes hours. Need systematic observability.
Medium-High — storage, query, retention
High — instrumentation, dashboards, alerts
Medium — three pillars, correlation
High cardinality → storage explosion, query timeout
Sampling misses rare errors (tail-based mitigates)
Clock skew → trace ordering wrong
Missing context propagation → fragmented traces
Alert fatigue: too many, not actionable