Avisi cloud logo
Kubernetes

Kubernetes Node Recycling

Feature state: stable

Manually request replacement or deletion of a node using annotations.

Applies to

This page describes how the platform behaves today on the AME versions offered in the update channels. A cluster that has not been upgraded in a while may still behave differently. Upgrade it before relying on what follows.

You drive the lifecycle of a single Kubernetes node with two annotations. Use them to replace a node you no longer trust because its kubelet is stuck or its hardware is degrading, or to pick which node goes the next time the pool shrinks.

Annotations

We currently support the following node annotations:

AnnotationAction
k8s.avisi.cloud/node-recycle-requestedReplace the node, value "true". A new node joins the cluster first. Once it is Ready and passes its health checks, the platform cordons, drains and deletes the annotated node.
k8s.avisi.cloud/node-delete-requestedMark the node for removal, value "true". The next time the pool scales down, the platform picks it ahead of every other node except ones the cluster-autoscaler has already marked for deletion. Nothing is deleted until the pool size drops, through the autoscaler or a configuration change.

Neither annotation takes effect the moment you set it. The platform reads both on its next node reconcile, which runs about every two minutes.

When to use recycle vs delete

Recycle replaces a node without changing the pool size. The replacement joins and passes its health checks before the platform drains the old one, so the pool never runs a node short. Reach for this one.

Delete only marks a node. The platform acts on it when the pool scales down, and not before. If no scale-down comes, the node keeps running.

Recycle when you want a specific node gone now. Delete when the pool is already about to shrink and you want to pick which node goes.

How to request a recycle

kubectl annotate node/<node-name> k8s.avisi.cloud/node-recycle-requested=true

This will:

  1. Provision a new node (surge) and wait until it is Ready.
  2. Cordon & drain the old node.
  3. Delete the old node.

Recycle several nodes at once

Every node in the cluster:

kubectl annotate node --all k8s.avisi.cloud/node-recycle-requested=true

Every node in one node pool, selected by the pool name:

kubectl annotate node -l k8s.avisi.cloud/node-pool=<pool-name> \
  k8s.avisi.cloud/node-recycle-requested=true

The platform recycles the annotated nodes one at a time, one node pool after another, so the cluster carries one extra machine at any moment and never twice its node count. Every node waits for its replacement to provision and stay healthy for 30 seconds before it drains, so a large cluster takes a while to work through.

A node that belongs to no node pool keeps the annotation and nothing acts on it.

How to request a delete

kubectl annotate node/<node-name> k8s.avisi.cloud/node-delete-requested=true

This will:

  1. Mark the node for prioritized removal during the next scale-down.
  2. When the pool scales down, the platform cordons, drains and deletes this node early in the candidate order.

The node keeps running and accepting new pods until a scale-down comes.

Draining a node instead

Draining is the recommended way to take a node out of service. kubectl drain cordons the node and evicts its pods, respecting PodDisruptionBudgets, so the workload has already moved by the time the machine goes.

kubectl drain <node-name> --ignore-daemonsets

Cordoning stops new pods landing on the node and leaves the running ones alone.

kubectl cordon <node-name>

Either command makes the node unschedulable, which puts it ahead of the healthy nodes in the candidate order and behind a node carrying the delete annotation. Combine the two when you want a node emptied now and picked early at the next scale-down.

Scale-down candidate order

You set the size of a node pool, either directly or through the cluster autoscaler. That number decides how many machines leave when the pool shrinks. The order below decides which ones.

#CandidateWhat it means
1Nodes the autoscaler marked for removalThe autoscaler found the node redundant and has already drained it. Removing anything else strands it in the pool, cordoned and empty.
2Nodes you marked for removalYou set the k8s.avisi.cloud/node-delete-requested annotation. See how to request a delete.
3Machines that never became a nodeThe machine has no node yet, or its node is already gone. It costs money and runs no workload.
4Unreachable nodesKubernetes lost contact with the node and is already moving its pods elsewhere.
5Not ready nodesThe kubelet still reports in and says the node is not ready.
6Cordoned nodesThe node keeps running its pods but accepts no new ones.
7Healthy nodesNothing is wrong with them.

A machine that matches several rows takes the lowest number. Within a row the oldest goes first. Row 1 counts age from the moment the autoscaler marked the node, row 3 from when the machine was created, and the rest from when the node joined.

A machine that never became a node is also reclaimed outside a scale-down, so it cannot consume a removal meant for a node someone asked to have removed. The platform gives the machine 15 minutes to join, which is how long its join token stays valid, then holds it 10 minutes longer before deleting it. Reclaiming drops the pool below target, so the platform sizes the pool again in the same pass and the replacement is on its way before that pass ends.

Cancelling a request

Removing a recycle annotation cancels the request only until the platform picks it up, which takes up to about two minutes. Once the replacement node starts provisioning, the recycle runs to the end.

Removing a delete annotation cancels the request at any point before the scale-down, because the annotation does nothing on its own.

kubectl annotate node/<node-name> k8s.avisi.cloud/node-recycle-requested-
# or
kubectl annotate node/<node-name> k8s.avisi.cloud/node-delete-requested-

Notes and limitations

  • Neither annotation works on Bring Your Own Node (BYON) nodes. Recycle and delete those yourself.
  • Draining respects PodDisruptionBudgets. A strict PDB can hold a recycle or a scale-down open for as long as it blocks eviction.
  • Give critical workloads more than one replica before you recycle or delete a node.

On this page