Autoscale from Zero
Feature state: betaLet an autoscaling node pool drop to zero nodes and grow again when a pod needs it
An autoscaling node pool normally holds at least one node. With autoscale from zero a pool drops to zero nodes and still grows when a pod needs it, so you pay for a pool only while it does work.
This runs in the opposite direction from scale node pools to zero, which covers emptying a cluster you no longer want running. There you take the cluster down and it stays down. Here the pool comes back on its own.
When to use it
A pool that sits idle most of the time and carries a distinct machine type pays for itself here. Batch jobs, a nightly build pool, a pool with a taint that only a few workloads tolerate.
A pool that always has work does not benefit. Autoscaling from a minimum of one already covers it, and it avoids the cold start.
Pair it with a system node pool
An autoscaling pool reaches zero only when nothing permanent runs on it. Give the cluster a small fixed pool for CoreDNS, ingress controllers, operators and anything else that always runs, and let the autoscaling pools carry the work that comes and goes.
Without that split the always-on pods land on the autoscaling pool, and the autoscaler will not remove a node that runs them. The pool then stops at one node instead of zero.
Enabling it
Set the minimum size of an autoscaling node pool to zero and leave the maximum size above zero through the REST API, acloud or Terraform. The pool then shrinks to nothing once its nodes are empty, and the autoscaler adds a node again as soon as a pod needs one that no other pool can take.
With acloud:
acloud node-pools scale <node-pool-id> \
--auto-scaling \
--node-count 0 \
--max-size 3Leave --auto-scaling off and acloud sets the maximum size to zero as well, which pins the pool shut instead. AME then reads it as a pool you finished with and overrides your Pod Disruption Budgets while it empties.
Cold start
A pod that triggers a scale-up from zero waits while AME creates a machine, boots it and joins it to the cluster. Expect a few minutes, depending on the cloud and the machine type. Keep the minimum size at one for anything that cannot wait.
When a pool does not come back
A pod only wakes a pool it can run on. Read the pending pod's events first.
- The pod asks for more CPU or memory than the node type offers. Leave room for the kubelet and the system pods, because a pod never gets all of a node's memory.
- The pod misses the toleration for a taint on the pool.
- The pod misses the node selector or affinity that points at the pool.
Keeping a pool at zero
A pool empties only when every pod on its last node can move off it. Two things hold a node open.
- A pod that always runs, such as a controller or an operator you pinned to the pool with a node selector. Put always-on workloads on a fixed pool instead.
- A pod using local storage, an
emptyDirfor instance. The autoscaler leaves that node alone while the pod exists. A short-lived pod is no problem, because the node comes free once the pod is gone. A pod that stays around holds the node until you remove it.
The annotation cluster-autoscaler.kubernetes.io/safe-to-evict: "true" tells the autoscaler to ignore both rules for a pod. Use it only on a pod that can be interrupted and restarted without losing anything. Never put it on a build or a batch job, because the autoscaler will take the node away mid-run and the work is lost.
A taint on the pool is the simplest way to keep unrelated pods off it, so nothing lands there that you did not intend.
Limits
- The last node in a cluster stays. The autoscaler does not remove a node that runs system pods, so a cluster never reaches zero nodes on its own. Use scale node pools to zero for that.
- Keep a minimum size of one on a pool with GPUs.