NewCompare CPU & GPU pricing across AWS, Azure & GCP
Scheduler

Troubleshooting

Common issues and solutions for the DevZero Scheduler (dz-scheduler), including the PriorityClass quota denial seen on GKE.

Troubleshooting

Scheduler Pod Never Starts: "insufficient quota to match these scopes"

The Deployment exists but no dz-scheduler pod appears. The DevZero dashboard flags the cluster under Issues Detected, and kubectl describe on the ReplicaSet shows a FailedCreate event:

Error creating: insufficient quota to match these scopes: [{PriorityClass In [system-node-critical system-cluster-critical]}]

Cause

dz-scheduler defaults to the Kubernetes system-cluster-critical PriorityClass, so the kubelet never preempts it under node pressure. Some clusters restrict that class -- and system-node-critical -- to the kube-system namespace with a ResourceQuota scoped to those PriorityClasses. GKE does this on every cluster; other managed and hardened distributions can too.

Any pod outside kube-system that requests one of those classes is denied at admission. The denial is a native ResourceQuota check, not an admission webhook, so it does not appear alongside Gatekeeper or Kyverno errors, and there is usually no ResourceQuota object in devzero-system to inspect.

Confirm it is the quota:

# The ReplicaSet carries the FailedCreate event
kubectl describe replicaset -n devzero-system -l component=dz-scheduler | grep -A3 Events

# The restricting quota lives in kube-system, not in your namespace
kubectl get resourcequota -A -o yaml | grep -B10 'system-cluster-critical'

Fix: use a PriorityClass the cluster permits

The dakr-operator chart (version 0.2.23 and later) exposes scheduler.priorityClassName. Point it at the chart's own critical class, which the operator and agent already use. The class is named <release>-dakr-operator-critical, so with the release name dakr that the dashboard's install command uses, it is dakr-dakr-operator-critical. (If your release name already contains dakr-operator, the chart doesn't repeat it and the class is <release>-critical.)

helm upgrade dakr \
  oci://registry-1.docker.io/devzeroinc/dakr-operator \
  --namespace devzero-system \
  --set priorityClass.enabled=true \
  --set scheduler.priorityClassName=dakr-dakr-operator-critical \
  --reuse-values

Or with the CLI:

dz install write --upgrade \
  --set priorityClass.enabled=true \
  --set scheduler.priorityClassName=dakr-dakr-operator-critical

If you named the release differently, find the class first:

kubectl get priorityclass | grep critical

Any other cluster-provided class works too. It only needs to exist and to be permitted in devzero-system.

priorityClass.enabled defaults to true, so the chart's class normally exists already. Setting it explicitly guards against an earlier install that disabled it.

Verify the pod comes up:

kubectl get pods -n devzero-system -l component=dz-scheduler
kubectl get pod -n devzero-system -l component=dz-scheduler \
  -o jsonpath='{.items[0].spec.priorityClassName}'

Standalone install

If you installed from the raw manifest rather than the Helm chart, priorityClassName: system-cluster-critical is hardcoded in the dz-scheduler Deployment. Patch it to a permitted class:

kubectl patch deployment dz-scheduler -n devzero-system --type merge \
  -p '{"spec":{"template":{"spec":{"priorityClassName":"dakr-dakr-operator-critical"}}}}'

Other DevZero components

The same denial can hit any pod outside kube-system that asks for a system-* class. Current versions of the Read Operator and Network Operator charts ship their own devzero-zxporter-devzero-zxporter-critical class for exactly this reason. If an older install of either shows this error, upgrade the chart:

dz install read --upgrade
dz install network --upgrade

Pods Stay Pending With schedulerName: dz-scheduler

If the scheduler pod is running but pods that opt in never get bound:

  1. Confirm the scheduler is healthy and holding its leader lease:

    kubectl logs deployment/dz-scheduler -n devzero-system
    kubectl get lease -n kube-system | grep dz-scheduler
  2. Check the pending pod's events for a filter or scoring failure:

    kubectl describe pod <name> -n <namespace> | grep -A10 Events
  3. A pod annotated checkpoint-shim.io/try-restore: "true" only lands on nodes labelled dakr.devzero.io/checkpoint-node: "true". If no such node exists, the CheckpointRestore filter rejects every node.

Scheduler Fails to Start: No Control Plane Token

The NodeCost plugin validates its token source at startup and exits if none resolves. Check the logs for a token error, then confirm the Secret or ConfigMap it references exists:

kubectl logs deployment/dz-scheduler -n devzero-system | grep -i token
kubectl get secret devzero-zxporter-token -n devzero-system

See Configuration for the token resolution order.

How to Check Scheduler Logs

# Current logs
kubectl logs deployment/dz-scheduler -n devzero-system

# Previous container logs (if restarted)
kubectl logs deployment/dz-scheduler -n devzero-system --previous

# Follow logs in real-time
kubectl logs deployment/dz-scheduler -n devzero-system -f

On this page