Back to Journal Kubernetes Platform

The EKS Checklist I Run Before Any Production Launch

Private subnets, IRSA, Karpenter limits, PodDisruptionBudgets, image signing, and the five Grafana dashboards I refuse to launch without.

Kubernetes monitoring dashboards tracking production cluster health

Pre-flight configuration checks save engineering teams from midnight Sev-1 incidents.

Every quarter I get called in to rescue an Amazon EKS cluster that "just fell over" during a traffic surge or after a minor version upgrade. In 90% of cases, the failure had nothing to do with Kubernetes itself — it was caused by three or four fundamental configuration omissions that took five minutes to overlook during initial cluster creation.

1. Network & IAM Roles for Service Accounts (IRSA)

First rule of production clusters: the EKS API endpoint must never have public 0.0.0.0/0 access enabled. Use private endpoint access within your Transit Gateway or front it with AWS Systems Manager / Tailscale bastion nodes.

Second, never give an EC2 Node Instance Role broad AWS permissions. Pods requiring AWS service access must use fine-grained IAM Roles for Service Accounts (IRSA) or the newer EKS Pod Identity feature. When every pod shares the node role, any remote code execution (RCE) inside a container compromises your entire AWS account.

karpenter-nodepool.yaml
apiVersion: karpenter.sh/v1beta1
kind: NodePool
metadata:
  name: default-workloads
spec:
  template:
    spec:
      requirements:
        - key: kubernetes.io/arch
          operator: In
          values: ["arm64", "amd64"]
        - key: karpenter.sh/capacity-type
          operator: In
          values: ["spot", "on-demand"]
      limits:
        cpu: "250"
        memory: "500Gi"
  disruption:
    consolidationPolicy: WhenUnderutilized
    expireAfter: 720h

2. Karpenter Limits & Disruption Budgets

If you're still relying on the legacy Kubernetes Cluster Autoscaler, you are burning money. Karpenter evaluates pending pods and launches exactly right-sized EC2 instances in under 45 seconds directly through the EC2 Fleet API.

However, running Karpenter without hard spec.limits is dangerous. A runaway deployment with recursive pod creation can scale your cluster to hundreds of nodes before your billing alert fires. Always set CPU and memory hard ceilings on every NodePool.

Pro-Tip on Spot Instances

Ensure every critical workload has a PodDisruptionBudget (PDB) with minAvailable: 1. When AWS reclaims a Spot instance with a 2-minute warning, Karpenter will drain gracefully without dropping live requests.

3. Pod Disruption Budgets & Admission Guardrails

A Kubernetes cluster without an admission controller is an accident waiting to happen. Use Kyverno or Gatekeeper to reject any deployment that fails these four non-negotiable rules:

  • CPU/Memory limits: Every container must specify explicit requests and limits.
  • Run as non-root: Containers must execute as UID 1000+ with readOnlyRootFilesystem: true.
  • Readiness & liveness probes: Prevent traffic routing before applications are initialized.
  • Anti-affinity: Spread replicas across different Availability Zones (AZs) automatically.

4. The 5 Dashboards I Refuse to Launch Without

Before putting real customer traffic on a cluster, verify that these five Grafana / Datadog dashboards are active and monitored:

  1. Node Memory Pressure & OOMKilled Rates: Catch memory leaks before pods cascade fail.
  2. Kube-State-Metrics Unschedulable Pods: Alerts if Karpenter hits cloud quota limits.
  3. CoreDNS Latency & Packet Drops: The silent killer of microservice response times.
  4. Ingress Controller 5xx vs 4xx Error Ratio: Direct reflection of user experience.
  5. AWS Cost Anomaly per Namespace: Identifies which squad's deploy created a spend spike.

Production Readiness Interactive Checklist

Click items to audit your own cluster readiness:

Ryan Cole portrait

Written by Ryan Cole

Cloud Architect with 9+ years helping SaaS and fintech teams build resilient, cost-efficient Kubernetes platforms on AWS and Azure.

Work With Me

Get Monthly Cloud Architecture Notes

One monthly breakdown of real production failures, architectural fixes, and reusable Terraform/Kubernetes snippets. No fluff.