Autoscaling from zero

Beta. Autoscaling from zero is a beta feature. Behaviour and defaults can still change while we gather feedback. We plan to move it to stable after three months of production use, around December 2026.

We are happy to announce autoscaling from zero. An autoscaling node pool can now hold zero nodes and still grow when a pod needs it. Set the minimum size of the pool to zero, leave the maximum above zero, and the pool empties itself when the work is done.

Until now the minimum was one. A pool that ran a nightly build or a handful of CI jobs kept a node running the other twenty-three hours of the day, and you paid for all of it. We got tired of looking at that node too, which is most of why this feature exists.

What it saves

Count node-hours. A pool with a minimum of one bills 24 node-hours a day no matter how little it does. Move that minimum to zero and you bill the hours your workload asks for, plus the few minutes each node spends booting and draining.

A CI pool that runs 40 minutes of pipelines a day drops from 24 node-hours to roughly 2. A nightly batch job that finishes in an hour drops from 24 to about 1.5. The bigger the machine type, the more that gap is worth, which is why a burst pool is usually the one to move first.

Do not move your general workload pool. It has pods on it all day, so the minimum never comes into play, and all you buy yourself is a cold start on the day something does drain it.

Pair it with a system node pool

Across the clusters we run and the ones we help run, autoscaling pools work best when the cluster also has a small fixed pool for everything that never stops. CoreDNS, the ingress controller, operators, a runner manager.

Give those pods a home that is not going anywhere and the autoscaling pool is free to empty. Skip the split and they spread over whatever pool exists, the autoscaler refuses to remove a node that runs them, and the pool floors at one node instead of zero. You end up paying for the node you set out to remove.

Both examples below have that shape. One small fixed pool, one pool that comes and goes.

A dedicated pool for GitLab runners

CI is the clearest case. The load arrives in bursts, it can wait a few minutes, and it wants bigger machines than the rest of the cluster needs.

Split the cluster in two pools. A small fixed pool carries everything that always has to run, and a tainted pool carries the CI jobs and drops to nothing between pipelines.

# Always-on pool for CoreDNS, ingress and the runner manager
acloud node-pools create \
  --name=system \
  --cluster=example-cluster \
  --node-type=cpx21 \
  --node-count=2

# Burst pool for CI jobs, empty when no pipeline runs
acloud node-pools create \
  --name=runners \
  --cluster=example-cluster \
  --node-type=cpx41 \
  --auto-scaling \
  --node-count=0 \
  --max-size=10

Add the taint gitlab-runner=true:NoSchedule to the runners pool in the Avisi Cloud Console or through the REST API. The taint keeps everything else off those nodes, so one unrelated pod cannot hold the pool open.

Point the runner at the pool from your config.toml:

[[runners]]
  [runners.kubernetes]
    [runners.kubernetes.node_selector]
      "k8s.avisi.cloud/node-pool" = "runners"
    [runners.kubernetes.node_tolerations]
      "gitlab-runner=true" = "NoSchedule"

Three things decide whether the pool really reaches zero.

Keep the runner manager off the burst pool. The manager pod runs all the time. Put it on the runners pool and it holds a node open forever, which is the opposite of what you built the pool for. Give it the node selector your other always-on workloads use, or no selector at all, so it lands on system.

Do not mark job pods as safe to evict. A GitLab runner job pod mounts an emptyDir for its build directory, and the autoscaler leaves a node with local storage on it alone. That is what you want here. Annotate the pod with cluster-autoscaler.kubernetes.io/safe-to-evict: "true" and the autoscaler is free to take the node out from under a running build, which fails the pipeline. The runner deletes its job pods when they finish, so the node comes free on its own a moment later.

Size the requests to the machine type. A pod that asks for more CPU or memory than a node in the pool can offer never starts that pool, so it sits pending while the pool sits at zero. Set the runner's requests against the node type you picked, and leave room for the kubelet and the system pods on the node.

You can move an existing pool the same way, without recreating it:

acloud node-pools scale <node-pool-id> --auto-scaling --node-count 0 --max-size 10

A Playhouse between sessions

Our own Playhouse clusters have the same shape. acloud playhouse create provisions a small fixed system pool for Flux and the Tailscale operator, and an autoscaling playrooms pool that is labelled and tainted so only Playrooms land there.

A Playhouse is a sandbox. It is busy during a workshop and idle the rest of the week, and that playrooms pool used to hold a node open the whole time for nobody. With a minimum of zero it costs its system node between sessions, and the first Playroom somebody starts brings the pool back.

What to watch for

A cold start takes a few minutes. A pod that wakes a pool up waits while AME creates a machine, boots it and joins it to the cluster. How long depends on the cloud and the machine type. Keep the minimum at one for anything that cannot wait, and use this on work that already runs asynchronously.

A pod has to match the pool to wake it. A pod that misses the toleration for the pool taint, or asks for more than the node type offers, leaves the pool at zero and stays pending. Read the pod events first when a pool does not come back.

A cluster still keeps one node. The autoscaler does not remove a node that runs system pods, so a cluster never reaches zero nodes on its own. Emptying a whole cluster is a different feature, and scale node pools to zero covers it.

Advanced: tuning how fast a pool empties

The defaults suit a steady workload. A burst pool usually wants something sharper, and you set that per cluster with a PATCH:

curl -X PATCH \
  -H "Authorization: Bearer $ACLOUD_TOKEN" \
  -H "Content-Type: application/json" \
  https://api.avisi.cloud/api/v1/orgs/$ORG/clusters/$ENVIRONMENT/$CLUSTER \
  -d '{
    "clusterAutoscalerSettings": {
      "scale-down-unneeded-time": "2m",
      "scale-down-delay-after-add": "2m",
      "scale-down-utilization-threshold": 0.5,
      "max-node-provision-time": "15m"
    }
  }'

The four that matter most here:

  • scale-down-unneeded-time is how long a node has to sit unneeded before the autoscaler removes it. The default waits longer than a CI pipeline usually takes, so shortening it is what gets a burst pool back to zero soon after the last job.
  • scale-down-delay-after-add holds off a scale-down right after a scale-up. Set it too low on a pool that gets work in waves and you pay for the same node twice.
  • scale-down-utilization-threshold is the fraction of requested to allocatable resources below which a node counts as unneeded.
  • max-node-provision-time is how long the autoscaler waits for a new node to join before it calls the scale-up failed. Raise it if your cloud or machine type boots slowly, because a cold start from zero is exactly where that limit bites.

The durations take Go duration strings, so 90s, 2m and 1h30m all work. See the cluster API reference for the full list.

One catch. These settings are cluster-wide, because the cluster autoscaler runs one instance per cluster. A value you pick for the burst pool lands on every other pool as well, so weigh it against your steadiest pool before you shorten anything.

Try it

Set the minimum size of an autoscaling pool to zero through the REST API, acloud or Terraform.

You need a cluster on cluster-controller v2.40.0 or newer, so upgrade it if it is behind.

The reference is at autoscale from zero. Tell us how it behaves on your workload through our support desk. What you report in these three months decides the defaults we ship when it goes stable.