Infrastructure as Code (IaC)

Infrastructure as Code (IaC)

Definition: The practice of managing and provisioning cloud infrastructure — servers, networks, load balancers, DNS records — through machine-readable definition files rather than manual clicks in a web console (commonly called “ClickOps”). The underlying idea traces back to configuration-management tools like CFEngine (1993) and later Puppet and Chef, but it crystallized into its modern form once tools were built specifically around cloud provider APIs — AWS CloudFormation (2011) and, more influentially, HashiCorp’s cloud-agnostic Terraform (2014). Treating infrastructure definitions as code means they can be version-controlled, code-reviewed, diffed, and tested like application source rather than existing only as one engineer’s memory of what they clicked.

How It Works

  • Declarative IaC (Terraform, CloudFormation, Pulumi): you specify the desired end-state — “3 EC2 instances of type t3.medium behind a load balancer” — and the tool computes a diff against recorded state and applies only what’s needed to close the gap
  • Imperative IaC (Ansible, Chef, raw bash/CLI scripts): you specify explicit step-by-step commands to reach a state — “install nginx, copy this config, restart the service” — leaving the tool with no independent notion of “desired state” to reconcile against
  • State files: declarative tools track a mapping between resource blocks in code and real-world resource IDs in a state file (terraform.tfstate), which is what lets a plan compute an accurate diff instead of re-creating every resource on every run
  • Providers and modules: provider plugins translate a tool’s generic resource syntax into actual cloud API calls, and modules package reusable, parameterized bundles of resources (a “standard VPC” module every team imports) instead of copy-pasted configuration
  • Plan/apply workflow: a plan step renders a human-readable preview of additions, changes, and deletions before anything touches real infrastructure — a review checkpoint that a raw CLI script or console click never gives you
  • Drift detection: running plan against real infrastructure surfaces configuration drift — resources modified outside of IaC, like a manual console tweak during an incident — before the code’s model of reality silently diverges further from the truth

Why It Matters

  • Makes infrastructure changes repeatable and auditable — every change goes through the same review and version-control process as application code, instead of being an undocumented click only one departed engineer remembers
  • Enables real disaster recovery at the infrastructure level — an entire environment (VPC, subnets, clusters, DNS) can be torn down and rebuilt identically from code rather than reconstructed from tribal knowledge under pressure
  • Eliminates configuration drift between environments, since dev, staging, and production are generated from the same modules with different parameters instead of hand-built environments that quietly diverge over months
  • Collapses the line between “ops” and “software engineering” — infrastructure changes flow through the same CI/CD pipelines, pull requests, and automated checks as any other code change
  • Turns infrastructure knowledge into an organizational asset rather than an individual one — a new hire can read main.tf to understand exactly what exists and how it’s wired together, instead of reverse-engineering a live console

Under the Hood: The Dependency Graph and Parallel Apply

Terraform’s core engine builds a directed acyclic graph (DAG) of every resource in a configuration, inferring edges from implicit references — one resource’s attribute feeding another resource’s argument, like a subnet ID produced by aws_subnet.public.id being consumed by an EC2 instance block — rather than requiring the author to declare ordering explicitly. Resources with no dependency relationship to each other are created, updated, or destroyed in parallel up to a configurable concurrency limit, which is why applying 200 independent resources finishes in a fraction of the time a naive top-to-bottom script would take, and why destroying infrastructure walks the same graph in reverse so dependents are torn down before their dependencies. This graph is also what makes plan trustworthy as a review artifact: before touching anything real, Terraform refreshes its record of existing resources, walks the graph to compute exactly which nodes have changed, and prints that diff — the apply phase is then just a mechanical execution of a plan a human already reviewed, not a fresh interpretation of the configuration.

Comparison: IaC vs Manual Provisioning (ClickOps) vs Configuration Management

IaC (Terraform/CloudFormation)Manual Provisioning (ClickOps)Config Management (Ansible/Chef)
RepeatabilityHigh — same code, same result every timeLow — depends on human memory and consistencyModerate — scripts are repeatable, but often imperative
Audit trailFull history via version controlNone, beyond cloud provider audit logs (if enabled)Partial — playbooks are versioned, execution isn’t always
Speed to first resourceSlower upfront (write and review code)Fastest for a one-off resourceModerate
Best fitWhole-environment provisioning, teams, productionPrototyping, one-off experiments, learning a new serviceConfiguring software on already-provisioned servers
Failure modeDrift if changes bypass the toolUndocumented, irreproducible snowflake infrastructurePlaybooks drift from actual server state over time
RollbackRe-apply a previous commit’s stateManual, error-prone, relies on memoryRe-run a previous playbook version

Common Pitfalls

  • Committing unencrypted state files containing plaintext database passwords or API keys to version control, since state files often store resource attributes in full, including sensitive ones
  • Manual out-of-band changes (“just this once, I’ll click it in the console” during an incident) causing drift that later plan/apply runs either silently overwrite or refuse to reconcile cleanly
  • Not using remote state with locking (an S3 backend plus a DynamoDB lock table, or Terraform Cloud), so two engineers running apply simultaneously race against or corrupt the same state file
  • Building one monolithic state file for an entire organization’s infrastructure, so a single stuck resource or bad apply blocks every team instead of just the one that owns it
  • Hardcoding values instead of using variables, making a module unreusable across environments and forcing copy-paste drift between dev, staging, and production
  • Treating apply as a rubber stamp rather than actually reading the plan diff, which is the last checkpoint before a destructive change lands in production
  • Pinning no version constraints on providers or modules, so a routine terraform init months later silently pulls a breaking major-version upgrade into an otherwise unrelated change

Code Example

terraform {
  backend "s3" {
    bucket         = "acme-terraform-state"
    key            = "prod/network/terraform.tfstate"
    region         = "us-east-1"
    dynamodb_table = "terraform-locks"
  }
}

variable "instance_type" {
  default = "t3.medium"
}

resource "aws_instance" "web" {
  ami           = "ami-0abcdef1234567890"
  instance_type = var.instance_type
  subnet_id     = aws_subnet.public.id
  tags = { Name = "web-app", Environment = "production" }
}
terraform plan   # preview additions/changes/deletions before touching anything
terraform apply  # execute the reviewed plan

Best Practices

  • Use remote state with locking (S3+DynamoDB, Terraform Cloud, GCS) so concurrent applies never silently overwrite each other
  • Split state by environment and service boundary rather than maintaining one monolithic state file for the whole org
  • Always read the plan diff before every apply, treating it as a required review gate rather than a formality to click through
  • Pull secrets from a dedicated secrets manager (Vault, AWS Secrets Manager) at apply time instead of committing them into version control
  • Run infrastructure changes through the same CI/CD pipeline and pull-request review as application code, rather than applying ad hoc from a laptop
  • Pin provider and module versions explicitly, so a fresh init months from now reproduces the same behavior instead of quietly picking up a breaking upgrade

FAQ

Is IaC the same thing as configuration management? Related but distinct — IaC (Terraform, CloudFormation) typically provisions the infrastructure itself (VMs, networks, load balancers), while configuration management (Ansible, Chef) typically configures software once a server already exists; many teams use both together.

What happens if someone manually changes a resource IaC manages? The next plan detects the drift and shows it as a diff — depending on the change, the next apply will either revert the manual change back to what the code declares, or the code needs to be updated to match the new reality, but either way the drift doesn’t stay silent for long.

Do I need IaC for a single small server? Usually not essential at that scale, a manually configured server or a simple provisioning script is fine — IaC’s value compounds as the number of resources, environments, and people touching infrastructure grows.

Can Terraform manage infrastructure it didn’t originally create? Yes, via terraform import, which links an existing real-world resource to a resource block in code so future plans treat it like anything else IaC manages — a common step when adopting IaC on top of infrastructure that was originally built by hand.

History

  • CFEngine (1993) pioneered declarative, idempotent configuration management, predating the term “Infrastructure as Code” by nearly two decades
  • Puppet (2005) and Chef (2009) popularized configuration-as-code for servers and fleets, before public cloud APIs made whole-infrastructure provisioning a realistic target for the same philosophy
  • AWS CloudFormation (2011) brought declarative provisioning to cloud resources themselves, not just the software running on top of them
  • HashiCorp’s Terraform (2014) became the dominant cloud-agnostic IaC tool by supporting dozens of providers behind one syntax (HCL) rather than being locked to a single cloud, and its 2023 license change to BSL prompted the community fork OpenTofu

Example

Writing a main.tf file that declares a VPC, its public and private subnets, and a handful of EC2 instances, then running terraform plan to review exactly what will be created before terraform apply provisions all of it in one auditable step — and if that environment needs to be rebuilt in a different region six months later, the same configuration does it again with a one-line variable change.

Dig deeper