Beyond Your Bill — selected guide
Autoscaling with GKE: Overview and pods
Google Cloud Tech 8 of 11
In this collection Browse 11 summaries 8 of 11
An unnamed presenter explains pod autoscaling as a cost-and-reliability control loop shaped by demand, startup latency, and headroom. The GKE and Kubernetes modes, metrics, probes, and compatibility advice are from 2020; check current docs and test with representative workloads.
Key Points Covered
- Autoscaling can add units horizontally or resize them vertically across workloads and infrastructure. [00:01:02]
- Horizontal Pod Autoscaler uses a demand metric and target; safe headroom depends on spikes and startup time. [00:02:05]
- The historical Vertical Pod Autoscaler modes resize CPU and memory, potentially through pod recreation. [00:03:08]-[00:04:11]
- Startup, probes, restart tolerance, bounds, and disruption management constrain safe scaling. [00:03:08]-[00:05:15]
- Horizontal and vertical loops can interfere when both react to the same resource signal. [00:05:15]-[00:06:17]
Full video: https://www.youtube.com/watch?v=7naCIxIaV1M(opens in a new tab)