“We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.”

Mihir Nair
Head of Architecture, Databahn
Optimize the resources and cost at the cluster, node, and workload level.
DevZero continuously analyzes real-time GPU allocation and usage across your Kubernetes clusters, automatically identifying idle capacity and enforcing policy-driven controls — without disrupting active training or inference jobs.
GPUs are expensive, scarce, and frequently over-provisioned for AI and ML workloads. Teams conservatively allocate resources, leaving capacity unused between jobs or during traffic lulls. The result? GPU spend driven by fear and guesswork, not utilization.
DevZero continuously monitors GPU allocation and actual usage across Kubernetes clusters. The system identifies three key waste patterns: ML training jobs that complete and leave GPUs idle, AI inference endpoints with warm pools consuming capacity during low traffic, and interactive notebooks left running after work ends.
You set the rules, DevZero executes them. Define allocation duration, cleanup triggers, and which workloads can access GPU resources at the cluster, namespace, or workload level.
Ready to get started?
DevZero continuously analyzes real-time CPU/GPU/RAM/Storage allocation and usage across your Kubernetes clusters, automatically identifying idle capacity, enforcing policy-driven controls, and reclaiming unused resources. By optimizing at the workload level and integrating with existing autoscalers, it ensures CPU/GPU/RAM/Storage resources are efficiently utilized without disrupting running workloads.
3 simple steps
Deploy the lightweight DevZero operator to your cluster. It runs in read-only mode, gathering metrics without making any changes to your infrastructure. Takes only minutes to set up.
Select your cloud provider:
$ curl -XPOST -H 'Authorization: Bearer ....' \
-H "X-Kube-Context-Name: $(kubectl config current-context)" \
"https://dakr.devzero.io/dakr/installer-manifest?cluster-provider=AWS" \
| kubectl apply -f -
DevZero immediately begins analyzing your cluster's real resource utilization across all workloads. By comparing requested resources to actual usage, we identify waste and calculate potential savings.
Cost
$272.12
CPU
13% request utilization
Memory
32% request utilization
| Workload | CPU Request | Memory Request | Total | |
|---|---|---|---|---|
Keywest | 2.12 / 0 m | 41.06 MiB / 0 MiB | $0.0970 CPU: $0.0270 · Mem: $0.0700 | ActiveOptimize |
ETL | 233.29 m / 320.24 m | 158.94 MiB / 291.8 MiB | $5.1749 CPU: $4.6535 · Mem: $0.5214 | ActiveOptimize |
Event_Proces | 1.67 m / 50.06 m | 34.18 MiB / 500.57 MiB | $1.5794 CPU: $0.6777 · Mem: $0.9017 | ActiveOptimize |
Configure cost optimization policies tailored to your workloads. Choose between Conservative, Moderate, or Aggressive optimization strategies, then let DevZero automatically optimize your resources.
Policy
General settings
Advanced settings (Vertical scaling)
Min request: 15 m
Target percentile: 75%
Adjust Limits: Disabled
Min request: 25 MiB
Target percentile: 75%
Adjust Limits: Disabled
Adjust Limits: Disabled
Adjust Limits: Disabled
“We were essentially able to reduce the cost of that cluster by about 75%. On AWS, DevZero demonstrated they could achieve significantly higher savings than we initially thought possible.”

Mihir Nair
Head of Architecture, Databahn