Most Kubernetes workloads have 30-65% waste from over-provisioning. Implementing Karpenter + spot instances + rightsizing typically saves 40-60% on compute within 3-6 months. Kubecost provides cost visibility per namespace/deployment. VPA recommendation mode drives rightsizing decisions based on actual usage. Spot instances are safe for stateless production with proper pod disruption budgets (Source: FinOps Foundation Kubernetes framework, 2026).
Last verified: Sep 14, 2026.
At a glance
- Typical waste in K8s: 30-65%
- Karpenter + spot + rightsizing: 40-60% savings
- Spot savings: 60-90% vs on-demand
- Kubecost free tier: up to 50 nodes
- VPA recommendation mode is the rightsizing tool
Where Kubernetes waste comes from
Kubernetes waste concentrates in three areas: over-provisioned requests, idle nodes, and missed spot opportunities. Identifying and addressing these three delivers 80 percent of optimization value.
| Waste Source | Typical Magnitude | Remediation |
|---|---|---|
| Over-provisioned CPU/memory requests | 30-50% of compute | VPA recommendation mode, then rightsizing |
| Idle nodes from slow autoscaling | 10-20% of compute | Karpenter or cloud-managed autoscaler |
| On-demand instead of spot/RI | 40-80% premium | Spot for stateless, RI/SP for steady state |
| Unattached PVs and orphaned load balancers | 5-10% of compute + storage | Cleanup automation |
| Inappropriate storage tiers | 10-30% of storage | Tiering policies |
| Cross-region egress | 5-15% of network | Co-locate compute and data |
Source: FinOps Foundation Kubernetes Framework and CloudZero K8s benchmarks, 2026.
Karpenter vs Cluster Autoscaler
Karpenter represents a significant improvement over the legacy Cluster Autoscaler for most workloads. Karpenter selects the optimal instance type per pod rather than relying on pre-configured node groups.
| Capability | Cluster Autoscaler | Karpenter |
|---|---|---|
| Provisioning model | Node groups (predefined) | Direct pod-to-instance matching |
| Bin-packing | Limited | Optimized consolidation |
| Instance type selection | Fixed per node group | Selects from instance type catalog |
| Spot integration | Manual configuration | Native consolidation + spot diversification |
| Consolidation | Limited | Active consolidation of underutilized nodes |
| Time to provision | Minutes (group scale-up) | Seconds (direct instance launch) |
| Cost savings (vs default) | 20-30% | 40-60% |
Source: AWS Karpenter benchmarks and FinOps case studies, 2026.
Spot instance strategy for production
Spot instances are safe for production when architected correctly. The 60 to 90 percent discount makes spot one of the largest single cost levers in Kubernetes FinOps.
| Workload Type | Spot Suitability | Mitigation Pattern |
|---|---|---|
| Stateless web/API | Excellent | Pod disruption budgets, multiple instance families |
| Batch/async jobs | Excellent | Job retry policies, checkpoint to durable storage |
| CI/CD runners | Excellent | Re-run capability, ephemeral data |
| Caches/ephemeral data | Good | Cache rebuilt on warm-up, no persistent state |
| Databases (with replicas) | Moderate | Primary on-demand, replicas on spot |
| Databases (single instance) | Avoid | Use on-demand or reserved |
Source: AWS spot best practices and FinOps Foundation Kubernetes guide, 2026.
Kubecost visibility
Cost visibility is the foundation of Kubernetes FinOps. Kubecost (now part of IBM Instana) provides real-time cost allocation by namespace, deployment, label, and team, plus recommendations for rightsizing.
| Kubecost Tier | Coverage | Cost |
|---|---|---|
| Free / Open Source | Single cluster up to 50 nodes | Free |
| Team | Multi-cluster, more than 50 nodes | $200-$1,000/mo (estimate) |
| Enterprise | Multi-cluster, SSO, governance, alerts | Custom quote |
Source: Kubecost pricing documentation, 2026.
Rightsizing with VPA recommendation mode
Vertical Pod Autoscaler (VPA) recommendation mode observes workloads without changing requests and recommends appropriate sizing. Deploy VPA in recommendation mode for 7 to 14 days, then apply recommendations with appropriate headroom.
| Step | Action | Outcome |
|---|---|---|
| 1. Deploy VPA recommenders | Run VPA in recommendation mode across all namespaces | Continuous collection of actual usage |
| 2. Aggregate recommendations | Export VPA recommendations to dashboard | Visibility into over-provisioning |
| 3. Calculate target requests | Use 95th percentile + 20-30% headroom | New request values per workload |
| 4. Apply via Helm/Kustomize | Update deployment manifests | Gradual rollout with monitoring |
| 5. Monitor for OOMKilled and throttling | Watch pod events and metrics | Validate savings without reliability regression |
Source: Kubernetes VPA documentation and FinOps Foundation guidance, 2026.
Internal links
See related cloud cost optimization guides: EKS vs AKS vs GKE, AWS cost optimization, and Cloud cost anomaly tools.
FAQs
See FAQ section above for K8s cost optimization savings, Karpenter vs Cluster Autoscaler, spot instance production patterns, Kubecost visibility, and VPA rightsizing workflow.
Kubernetes cost optimization roadmap
Kubernetes cost optimization follows a 4-phase roadmap. Most enterprises complete Phases 1-3 within 6-9 months.
| Phase | Duration | Activities | Outcomes |
|---|---|---|---|
| 1. Visibility | 1-2 months | Kubecost deployment, cost allocation setup | Cost visibility per workload |
| 2. Rightsizing | 2-4 months | VPA recommenders, manual tuning | 30-50% compute reduction |
| 3. Spot + Autoscaling | 2-4 months | Karpenter, spot diversification | 40-60% baseline reduction |
| 4. Continuous optimization | Ongoing | Policy enforcement, FinOps culture | Sustained 60-80% reduction |
Source: FinOps Foundation Kubernetes Framework, 2026.
FAQ expansion
Q: Is Karpenter production-ready? Yes. Karpenter graduated to GA in 2024 and is now broadly adopted. AWS, Azure (AKS Node Autoprovisioning), and Google Cloud (GKE Autopilot) have all adopted Karpenter-style approaches to just-in-time node provisioning.
Q: What is the difference between HPA and Karpenter? HPA (Horizontal Pod Autoscaler) scales pod replicas based on CPU/memory metrics. Karpenter scales nodes (the underlying VMs) based on unschedulable pods. They are complementary: HPA adjusts workload size, Karpenter adjusts cluster size.
Q: Can spot instances handle production stateful workloads? With proper design, yes. Run primary stateful workloads on on-demand or reserved instances, and replicas on spot. Use pod disruption budgets to ensure availability. Stateful workloads that cannot tolerate interruption (single-instance databases) should remain on on-demand.






