The EKS Checklist I Run Before Any Production Launch
Private subnets, IRSA, Karpenter limits, PodDisruptionBudgets, image signing, and the five Grafana dashboards I refuse to launch without.
Private subnets, IRSA, Karpenter limits, PodDisruptionBudgets, image signing, and the five Grafana dashboards I refuse to launch without.
Pre-flight configuration checks save engineering teams from midnight Sev-1 incidents.
Every quarter I get called in to rescue an Amazon EKS cluster that "just fell over" during a traffic surge or after a minor version upgrade. In 90% of cases, the failure had nothing to do with Kubernetes itself — it was caused by three or four fundamental configuration omissions that took five minutes to overlook during initial cluster creation.
First rule of production clusters: the EKS API endpoint must never have public 0.0.0.0/0 access enabled. Use private endpoint access within your Transit Gateway or front it with AWS Systems Manager / Tailscale bastion nodes.
Second, never give an EC2 Node Instance Role broad AWS permissions. Pods requiring AWS service access must use fine-grained IAM Roles for Service Accounts (IRSA) or the newer EKS Pod Identity feature. When every pod shares the node role, any remote code execution (RCE) inside a container compromises your entire AWS account.
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
name: default-workloads
spec:
template:
spec:
requirements:
- key: kubernetes.io/arch
operator: In
values: ["arm64", "amd64"]
- key: karpenter.sh/capacity-type
operator: In
values: ["spot", "on-demand"]
limits:
cpu: "250"
memory: "500Gi"
disruption:
consolidationPolicy: WhenUnderutilized
expireAfter: 720h If you're still relying on the legacy Kubernetes Cluster Autoscaler, you are burning money. Karpenter evaluates pending pods and launches exactly right-sized EC2 instances in under 45 seconds directly through the EC2 Fleet API.
However, running Karpenter without hard spec.limits is dangerous. A runaway deployment with recursive pod creation can scale your cluster to hundreds of nodes before your billing alert fires. Always set CPU and memory hard ceilings on every NodePool.
Ensure every critical workload has a PodDisruptionBudget (PDB) with minAvailable: 1. When AWS reclaims a Spot instance with a 2-minute warning, Karpenter will drain gracefully without dropping live requests.
A Kubernetes cluster without an admission controller is an accident waiting to happen. Use Kyverno or Gatekeeper to reject any deployment that fails these four non-negotiable rules:
readOnlyRootFilesystem: true.Before putting real customer traffic on a cluster, verify that these five Grafana / Datadog dashboards are active and monitored:
Click items to audit your own cluster readiness:
One monthly breakdown of real production failures, architectural fixes, and reusable Terraform/Kubernetes snippets. No fluff.