Kubernetes (K8s)
Kubernetes (K8s)
Definition: An open-source container orchestration platform that automates the deployment, scaling, and management of containerized applications across clusters of hosts. Originally designed at Google, drawing on over a decade of internal experience running containers at scale with an internal system called Borg, it was open-sourced in 2014 and donated to the newly formed Cloud Native Computing Foundation (CNCF) in 2015. This note focuses on Kubernetes itself in depth — its architecture, objects, and day-to-day usage; see Container Orchestration and Kubernetes for the broader orchestration landscape it grew out of and competes within.
How It Works
- Control Plane: the master components that manage the cluster — the
kube-apiserver(the single entry point every client and component talks to), thescheduler(decides which node a new pod runs on based on resource requests, affinity rules, and taints/tolerations), thecontroller-manager(runs reconciliation loops that continuously push actual state toward desired state), andetcd(the distributed key-value store holding all cluster state) - Worker Nodes: run the actual workloads via the
kubelet(the agent that talks to the API server and manages pod lifecycle on that node), a container runtime (containerdor CRI-O, implementing the Container Runtime Interface), andkube-proxy(programs iptables/IPVS rules to implement Service networking) - Pods: the smallest deployable unit, one or more tightly coupled containers sharing a network namespace (same IP,
localhostbetween them) and optionally storage volumes - Core objects: a
Deploymentmanages aReplicaSetto keep N pod copies running and handles rolling updates; aServicegives a stable virtual IP/DNS name that load-balances across a dynamic set of pods (ClusterIP, NodePort, or LoadBalancer types); anIngresshandles HTTP(S) routing and TLS termination from outside the cluster - Declarative reconciliation: you write YAML manifests defining desired state (e.g., “3 replicas of my web app, image
v2.1”), and controllers continuously diff actual vs. desired state and act to close the gap — this is what makes a crashed pod get automatically replaced - Health checks:
livenessProberestarts a container that’s stuck;readinessProberemoves a pod from a Service’s endpoints until it’s actually ready to accept traffic, preventing traffic from hitting a pod that’s still starting up
Why It Matters
- It is the de-facto industry standard operating system for the cloud, letting companies run large microservice architectures reliably and self-heal from node/container failures without tying themselves to a single cloud provider’s proprietary orchestration
- The declarative model plus reconciliation loops means the same manifests describe the intended state on a laptop (minikube/kind), a self-managed cluster, or a managed offering like EKS/GKE/AKS
- Its extensibility (Custom Resource Definitions, Operators, admission webhooks) turned it from a scheduler into a platform other platforms are built on top of, rather than a fixed feature set
- Standardizing on Kubernetes means engineers, tooling, and hiring pipelines transfer across companies and clouds, an increasingly rare kind of portability in infrastructure
Under the Hood: The Reconciliation Loop
Every Kubernetes controller follows the same pattern, which is why the system is so predictable once you internalize it: watch the API server for changes to objects it cares about, compare the object’s current status to its spec (desired state), and issue whatever create/update/delete calls close the gap — then repeat, forever. This is a level-triggered system, not an edge-triggered one: a controller that misses an event still converges correctly on its next pass, because it’s always comparing full current state to full desired state rather than reacting to a diff. That property is what makes Kubernetes tolerant of a controller crashing and restarting mid-operation, a naive event-driven system would need to replay missed events precisely, but a reconciliation loop just re-evaluates from scratch.
Comparison: Kubernetes vs Docker Swarm vs Nomad
| Kubernetes | Docker Swarm | Nomad | |
|---|---|---|---|
| Complexity | High, steep learning curve | Low, simplest of the three | Moderate |
| Extensibility | Very high (CRDs, Operators, huge ecosystem) | Limited | Moderate, plugin-based |
| Scheduling scope | Containers only | Containers only | Containers, VMs, and raw binaries |
| Best fit | Large-scale, complex microservice systems | Small teams wanting simple container orchestration | Mixed workload types, simpler ops than K8s |
Common Pitfalls
- Utilizing Kubernetes for simple, monolithic applications where a basic Platform-as-a-Service (PaaS) like Heroku or Render would be significantly cheaper and easier to operate — K8s adds real operational overhead (control plane upgrades, RBAC, networking) that only pays off at a certain scale/complexity
- Not setting resource
requestsandlimitson containers: without requests, the scheduler can’t pack nodes efficiently; without limits, a single leaking pod can starve every other pod on its node (or trigger an OOM-kill at the worst moment) - Misconfigured liveness probes that restart a healthy-but-slow-to-respond container in a crash loop, making a temporary slowdown into a full outage
- Ignoring Pod Disruption Budgets, so a routine node drain or cluster upgrade takes down more replicas simultaneously than the application can tolerate
- Storing secrets as plain Kubernetes
Secretobjects (base64-encoded, not encrypted, by default) without enabling encryption at rest or using an external secrets manager - Running
kubectl applyby hand from a laptop against production, with no record in version control of what was actually applied or by whom
Code Example
# A minimal Deployment + Service: 3 replicas of an app, load-balanced internally
apiVersion: apps/v1
kind: Deployment
metadata:
name: web-app
spec:
replicas: 3
selector:
matchLabels: { app: web-app }
template:
metadata:
labels: { app: web-app }
spec:
containers:
- name: web-app
image: myregistry/web-app:v2.1
resources:
requests: { cpu: "250m", memory: "256Mi" }
limits: { cpu: "500m", memory: "512Mi" }
readinessProbe:
httpGet: { path: /healthz, port: 8080 }
---
apiVersion: v1
kind: Service
metadata:
name: web-app
spec:
selector: { app: web-app }
ports: [{ port: 80, targetPort: 8080 }]
Best Practices
- Always set resource
requestsandlimits, letting the scheduler and kubelet make informed packing and eviction decisions - Use readiness probes on every service that takes time to warm up, and separate them from liveness probes so a slow startup doesn’t trigger unnecessary restarts
- Define Pod Disruption Budgets for anything stateful or capacity-sensitive before the first cluster upgrade, not after an incident
- Prefer managed Kubernetes (EKS, GKE, AKS) over self-hosting the control plane unless you have a dedicated platform team, running
etcdand the API server reliably is its own specialty - Adopt GitOps (manifests in version control, applied via a controller like Argo CD or Flux) rather than manual
kubectl applyfrom a laptop, for auditability and repeatability
FAQ
Do I need Kubernetes for a small side project? Usually not — a PaaS like Render, Railway, or Fly.io gets a small app deployed with far less operational overhead, Kubernetes’ benefits show up at a scale or complexity most side projects never reach.
What’s the difference between a Deployment and a StatefulSet? A Deployment treats its pods as interchangeable and disposable; a StatefulSet gives each replica a stable identity and, typically, its own persistent volume, which is why databases and other stateful workloads use StatefulSets instead.
Why is etcd considered the most critical piece of a cluster?
It’s the sole source of truth for all cluster state — losing it without a backup means losing the record of every object in the cluster, even if the workloads themselves keep running temporarily on their existing nodes.
What happens to running pods if the control plane goes down? Already-scheduled pods keep running on their nodes, the kubelet doesn’t need the API server to keep an existing container alive, but no new scheduling, scaling, or self-healing can happen until the control plane recovers, which is why managed offerings run a highly-available multi-node control plane by default.
Is Helm part of Kubernetes itself? No — Helm is a separate, widely-adopted package manager for Kubernetes that templates and versions collections of manifests as a “chart,” it isn’t a core Kubernetes component but is close to a de-facto standard for distributing complex applications.
History
- Google open-sourced Kubernetes in 2014, drawing on lessons from its internal Borg and Omega cluster-management systems used for over a decade prior
- Donated to the newly formed Cloud Native Computing Foundation (CNCF) in 2015, establishing it as a vendor-neutral project rather than a Google-controlled one
- The name comes from the Greek for “helmsman” or “pilot,” and the seven-spoked wheel in its logo is a nod to the project’s original internal codename, “Project Seven of Nine”
- Became the de-facto standard by the early 2020s, with every major cloud provider offering a managed Kubernetes service (EKS, GKE, AKS) rather than competing with a proprietary orchestrator
- The CNCF ecosystem grew around it rapidly, projects like Prometheus, Envoy, and Helm graduated alongside Kubernetes, forming a broader “cloud native” toolchain that assumes Kubernetes as its foundation
Related Terms
- Container Orchestration and Kubernetes
- Service Mesh
- Auto-Scaling
- Infrastructure as Code (IaC)
- Microservices Architecture
Example
When a pod crashes due to an out-of-memory error, the kubelet reports the failure to the API server, the ReplicaSet controller notices actual replica count (2) no longer matches desired (3), and schedules a replacement pod on whichever node has capacity — typically within seconds, with zero human intervention.
Referenced by
- Apache Spark
- API Gateway
- Auto-Scaling
- AWS (Amazon Web Services)
- Cloud and Infrastructure Terms MOC
- Cloud Service Models
- Container Orchestration and Kubernetes
- DigitalOcean
- Docker Compose
- Google Cloud Platform (GCP)
- Grafana
- Helm
- Microservices Architecture
- Microsoft Azure
- Observability and Monitoring
- Prometheus
- Serverless Computing and Cold Starts
- Service Mesh
- Virtual Machines (VMs)
- Virtual Private Cloud (VPC) and Subnets