GPU scarcity is real. Waste is optional.

Optimize the resources and cost at the cluster, node, and workload level.

NamespaceCPUMemoryTotalStatus
keywest2.1 m / 0 m41.06 Mib / 0 Mib$0.0970Active
monitoring233.2 m / 320 m158 Mib / 291 Mib$5.1749Active
fluxcd1.86 m / 51.55 m18.96 Mib / 64 Mib$0.8279Active
lander8.93 m / 314.4 m0.24 Gib / 1.11 Mib$5.8474Active
ingress-nginx7.4 m / 20.01 m130 Mib / 171 Mib$0.5725Active
karpenter0.04 m / 1 cores0.31 Gib / 1 Gib$11.526Active

GPU requests over time

Capacity: 72 devices
Requests: 16.03 devices
Usage: 0 devices
020406080100
Current margin: Oct 2, 18:34Request/Usage
20 devices40 devices60 devices
Requests: 30 GPUs
Used: 5 GPUs

GPU optimization

DevZero continuously analyzes real-time GPU allocation and usage across your Kubernetes clusters, automatically identifying idle capacity and enforcing policy-driven controls — without disrupting active training or inference jobs.

Stop paying for idle GPUs

GPUs are expensive, scarce, and frequently over-provisioned for AI and ML workloads. Teams conservatively allocate resources, leaving capacity unused between jobs or during traffic lulls. The result? GPU spend driven by fear and guesswork, not utilization.

How it works

DevZero continuously monitors GPU allocation and actual usage across Kubernetes clusters. The system identifies three key waste patterns: ML training jobs that complete and leave GPUs idle, AI inference endpoints with warm pools consuming capacity during low traffic, and interactive notebooks left running after work ends.

Policy-driven management

You set the rules, DevZero executes them. Define allocation duration, cleanup triggers, and which workloads can access GPU resources at the cluster, namespace, or workload level.

GPU requests over time

Capacity: 72 devices
Requests: 16.03 devices
Usage: 0 devices
020406080100
Current margin: Oct 2, 18:34Request/Usage

Ready to get started?

How it works

DevZero continuously analyzes real-time CPU/GPU/RAM/Storage allocation and usage across your Kubernetes clusters, automatically identifying idle capacity, enforcing policy-driven controls, and reclaiming unused resources. By optimizing at the workload level and integrating with existing autoscalers, it ensures CPU/GPU/RAM/Storage resources are efficiently utilized without disrupting running workloads.

3 simple steps

  1. 1

    Install a read-only operator

    Deploy the lightweight DevZero operator to your cluster. It runs in read-only mode, gathering metrics without making any changes to your infrastructure. Takes only minutes to set up.

    Select your cloud provider:

    Curl

    $ curl -XPOST -H 'Authorization: Bearer ....' \
    -H "X-Kube-Context-Name: $(kubectl config current-context)" \
    "https://dakr.devzero.io/dakr/installer-manifest?cluster-provider=AWS" \
    | kubectl apply -f -

  2. 2

    Gather metrics and calculate waste

    DevZero immediately begins analyzing your cluster's real resource utilization across all workloads. By comparing requested resources to actual usage, we identify waste and calculate potential savings.

    Cost

    $272.12

    CPU

    13% request utilization

    Memory

    32% request utilization

    WorkloadCPU RequestMemory RequestTotal
    Keywest

    2.12 / 0 m

    41.06 MiB / 0 MiB

    $0.0970

    CPU: $0.0270 · Mem: $0.0700

    ActiveOptimize
    ETL

    233.29 m / 320.24 m

    158.94 MiB / 291.8 MiB

    $5.1749

    CPU: $4.6535 · Mem: $0.5214

    ActiveOptimize
    Event_Proces

    1.67 m / 50.06 m

    34.18 MiB / 500.57 MiB

    $1.5794

    CPU: $0.6777 · Mem: $0.9017

    ActiveOptimize
  3. 3

    Define policies and optimize

    Configure cost optimization policies tailored to your workloads. Choose between Conservative, Moderate, or Aggressive optimization strategies, then let DevZero automatically optimize your resources.

    Policy

    Moderate Deltas (VPA) (mutating w/h w/o...
    # Policy NameCustom Policy
    Moderate Deltas (VPA) (mutating w/h w/o...Attached

    General settings

    BalancedOn DetectionScheduledPod CreationPod Update
    Every 45 minutes, every 4 hoursLookback: 6 days 23 hours

    Advanced settings (Vertical scaling)

    CPU

    Min request: 15 m

    Target percentile: 75%

    Adjust Limits: Disabled

    Memory

    Min request: 25 MiB

    Target percentile: 75%

    Adjust Limits: Disabled

    GPU

    Adjust Limits: Disabled

    GPU VRAM

    Adjust Limits: Disabled

    Live migration● On

Frequently asked questions

What our customers say

Databahn logo

“We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.”

Mihir Nair

Mihir Nair

Head of Architecture, Databahn