Avisi cloud logo
Kubernetes

Scale Node Pools to Zero

Feature state: stable

Understand when AME overrides Pod Disruption Budgets to drain a cluster to zero

AME drains a node gracefully and respects your Pod Disruption Budgets (PDBs) and grace periods. When the nodes are meant to go and nothing will bring them back, AME enters a two-phase shutdown flow that overrides PDBs so the cluster can reach an absolute zero state.

This page covers emptying a cluster you no longer want running. A single autoscaling pool that sits at zero and grows again when a pod needs it is a different feature, described under autoscale from zero.

When force eviction applies

AME enables force eviction mode when the machines must go and nothing brings them back.

  • Every node pool in the cluster sits at zero because you asked for it. Setting all of them to zero counts, and so does deleting all of them.
  • The cluster is being deleted. Nodes whose node pool you already removed take this path.

In every other case AME performs the standard drain and your PDBs and grace periods keep applying. Two cases look similar and do not count.

  • A pool that the autoscaler scaled down to zero does not count. Zero is a normal resting point for an autoscaling pool and the next pending pod brings it back, so your PDBs stand. An autoscaling pool counts as emptied on purpose only when you set its maximum size to zero as well.
  • Deleting one node pool while another pool keeps running does not count. The cluster still has somewhere to reschedule your workload onto, so AME drains those nodes normally.

Flow overview

  1. AME scales a node pool down, or clears nodes whose node pool you removed.
  2. If neither condition above applies, AME follows the standard drain process and the workflow ends.
  3. Otherwise AME enables force eviction mode and enters Phase 1.
  4. Phase 1 performs graceful evictions for up to five minutes.
  5. After the timeout, AME checks if pods remain:
    • If no pods remain, the nodes and backing machines are deleted and the workflow completes.
    • If pods remain, AME moves to Phase 2 and force deletes the remaining pods before removing the nodes and machines.

Phase details

Phase 1: graceful eviction

AME attempts a normal drain for five minutes:

  • Uses the eviction API so PDBs and graceful termination periods continue to apply.
  • Retries failed pod evictions until the timeout is reached.

Phase 2: force cleanup

If pods remain after the five-minute window, AME:

  • Deletes remaining pods with a zero-second grace period.
  • Bypasses the eviction API to prevent PDBs from blocking termination.
  • Deletes the backing machines so the cluster reaches zero provisioned nodes.

Deleting a cluster

Deleting a cluster does not put its machines through the two-phase flow. AME removes machines that still belong to a node pool straight away. Their pods get a zero-second grace period, AME does not use the eviction API, and your PDBs never apply. Only nodes whose node pool you removed earlier take the two-phase path.

Plan for that if a workload needs to shut down in order. Scale its node pool to zero first, wait for the drain to finish, and delete the cluster after that.

Emptying a cluster on a schedule

Automation can empty a cluster this way without manual cleanup, on one condition. Every pool has to be pinned shut, which means a fixed pool at zero nodes, or an autoscaling pool with both its minimum and its maximum size at zero. Leave one pool with autoscaling headroom and AME reads that pool as idle rather than finished. It then keeps honouring your PDBs, and a budget that cannot be satisfied stalls the drain.

On this page