Reliability & Operations
Observability, service reliability, incident readiness, and production operations.
Showing 1 resource
Resources
Engineering for Reliability
PlaylistTwenty-three substantive episodes on SLOs, telemetry, alerting, investigation, and operations
Twenty-three reviewed episodes that connect user-centered SLIs and SLOs to metrics, logs, traces, alerting, Prometheus, OpenTelemetry, quota headroom, GKE diagnosis, and observability cost. The learning sequence excludes the non-substantive trailer and treats all 2021–2022 interfaces, schemas, defaults, and pricing as historical.
Google Cloud TechLatest summary: May 25, 2022
Google CloudReliabilityObservability