Terraform

Terraform

Definition: HashiCorp’s tool for defining cloud infrastructure as declarative configuration files, then applying that config to create, update, or destroy real infrastructure to match it. Released in 2014, it grew from a way to provision a handful of AWS resources into the de-facto standard for multi-cloud Infrastructure as Code (IaC), with providers for hundreds of platforms beyond the major clouds — SaaS tools, DNS registrars, monitoring platforms, even other DevOps tools like this vault’s own GitHub Actions and Datadog entries.

Core Services & Concepts

  • HCL (HashiCorp Configuration Language) — the declarative syntax used to describe resources, designed to be more structured than YAML but more readable than raw JSON, and it’s also valid JSON under the hood if a tool needs to generate config programmatically
  • State file — Infrastructure as Code (IaC), a JSON file (terraform.tfstate) tracking what infrastructure Terraform believes exists and mapping it to the resources declared in config; the source of most real-world Terraform pain when it drifts from reality
  • Providers — plugins that let Terraform manage AWS, GCP, Azure, Cloudflare, and hundreds of other platforms with the same core workflow, each provider translating HCL resource blocks into that platform’s own API calls
  • Modules — reusable, parameterized bundles of resources (e.g. “a standard VPC with public/private subnets”) that can be versioned and shared across teams or pulled from the public Terraform Registry
  • Plan / Apply workflow — terraform plan computes and displays a diff between desired and current state without changing anything; terraform apply executes that diff, the two-step separation being Terraform’s core safety mechanism
  • Workspaces — a way to maintain multiple distinct state files (e.g. dev/staging/prod) from the same configuration, without needing entirely separate directories per environment
  • Remote backends — storing the state file in a shared location (S3 + DynamoDB locking, Terraform Cloud, etc.) instead of locally, which is required for any team beyond a single operator to avoid state conflicts
  • Data sources — read-only blocks that query existing infrastructure not managed by this Terraform configuration (e.g. looking up a shared VPC’s ID by tag), letting new resources reference infrastructure created elsewhere without importing it into the current state
  • Provisioners — an escape hatch for running scripts on a resource after creation (e.g. remote-exec), officially discouraged by HashiCorp as a last resort since it reintroduces the imperative, order-dependent fragility Terraform’s declarative model is designed to avoid

How It Works: The Reconciliation Model

  • Terraform builds a dependency graph from the resources declared in .tf files, so it knows a subnet must exist before a VM can be placed in it, and applies changes in the correct order automatically
  • terraform refresh (now folded into plan/apply by default) queries the real infrastructure via each provider’s API to detect drift — changes made outside Terraform, by hand or by another tool
  • plan performs a three-way diff between the configuration (desired state), the state file (what Terraform last knew), and reality (what the provider’s API reports), surfacing any of the three disagreeing with the others
  • Resources are identified internally by a resource address (aws_instance.web), and state operations like terraform state mv let an operator relabel a resource in state without destroying and recreating the real infrastructure behind it
  • Destructive changes (replacing a resource that can’t be updated in place, like most AWS instance type changes) are shown explicitly in the plan output before apply, marked for destroy-then-create or create-then-destroy depending on the resource

How Pricing Works

  • Terraform’s CLI and core engine are open-source (Business Source License since 2023, a change that caused some controversy and the OpenTofu fork discussed in History below)
  • Terraform Cloud offers a free tier for small teams (remote state, limited runs) and paid tiers adding SSO, private module registries, policy-as-code (Sentinel), and more concurrent runs
  • Terraform Enterprise is self-hosted, custom-priced, aimed at large organizations needing Terraform Cloud’s features behind their own firewall
  • The infrastructure Terraform provisions is billed by the underlying cloud provider, not by HashiCorp, Terraform itself has no per-resource fee for the open-source CLI usage
  • Most cost surprises come from the infrastructure being provisioned, not the tool, though Terraform Cloud’s paid tiers do bill per “run” or per user seat depending on the plan

Pros

  • Cloud-agnostic, one tool and syntax across every major provider plus hundreds of smaller ones
  • Plan-before-apply workflow shows exactly what will change before it happens, a meaningful safety margin compared to imperative scripts
  • Huge community module ecosystem for common infrastructure patterns, rarely necessary to write a VPC or Kubernetes cluster module completely from scratch
  • Explicit dependency graph means correct ordering happens automatically instead of being hand-sequenced in a script
  • Strong ecosystem tooling has grown up around it (Terragrunt for DRY multi-environment configs, Atlantis for PR-based apply workflows, checkov/tfsec for security scanning)
  • Import functionality (terraform import, and newer config-driven import blocks) lets existing hand-created infrastructure be brought under management incrementally, rather than requiring a risky big-bang migration

Cons

  • State file management is a common source of real incidents (drift, corruption, merge conflicts on concurrent applies without locking)
  • HCL has a learning curve distinct from general-purpose programming languages, and its limited expressiveness (no real loops until for_each/count matured) frustrates engineers used to imperative code
  • Destructive changes can happen silently if state and reality drift too far apart, or if a plan is applied without being carefully read first
  • Large monolithic state files slow down every plan/apply as infrastructure grows, pushing teams toward splitting state, which adds its own coordination overhead
  • The 2023 licensing change to BUSL restricts certain competitive commercial uses, a real consideration for vendors building products directly on top of Terraform
  • Provisioners and other imperative escape hatches, when used, break the clean declarative model and make plans harder to reason about, teams that lean on them heavily often end up fighting Terraform rather than working with it

Comparison: Terraform vs Ansible vs Pulumi

TerraformAnsiblePulumi
Primary purposeDeclarative infrastructure provisioningConfiguration management + ad hoc automationDeclarative infrastructure, in general-purpose languages
LanguageHCL (declarative)YAML playbooks (declarative-ish)Real languages: TypeScript, Python, Go, etc.
State trackingExplicit state fileStateless — checks actual system each runExplicit state, similar model to Terraform
Best fitProvisioning cloud resources across providersConfiguring servers after they exist, app deploymentTeams wanting IaC in a language they already use

Best For

  • Teams managing infrastructure across multiple cloud providers with one consistent workflow and syntax
  • Organizations that want infrastructure changes reviewed as a diff (via plan output) before anything touches production

Real Examples

  • Widely used to provision infrastructure at companies of every size, from startups to Uber, Slack, and most cloud-native enterprises
  • The public Terraform Registry hosts thousands of community and verified provider/module combinations, evidence of how broadly it’s been adopted as the default IaC choice
  • Many managed Kubernetes and cloud platform onboarding guides now include a Terraform quickstart alongside or instead of console click-through instructions

Use Cases

  • Provisioning cloud infrastructure (VMs, networks, databases, load balancers) as version-controlled code
  • Reproducing identical environments across dev, staging, and production from the same module
  • Multi-cloud or hybrid-cloud provisioning where one tool needs to speak to several providers consistently
  • Bootstrapping the initial infrastructure a Kubernetes cluster or CI/CD pipeline runs on top of
  • Managing non-cloud resources too — DNS records, SaaS tool configuration, monitoring dashboards — anything with a Terraform provider
  • Disaster recovery — with infrastructure fully defined as code, rebuilding an entire environment in a new region after a catastrophic failure is a matter of running apply against a new backend rather than manual reconstruction

Integration Notes & Common Pitfalls

  • Never edit the state file by hand, use terraform state subcommands, a manually edited state file is a common source of unrecoverable drift
  • Enable remote state with locking (S3+DynamoDB, Terraform Cloud) the moment more than one person touches an environment, local state plus concurrent applies is a recipe for corruption
  • Pin provider and module versions explicitly, an unpinned provider upgrading itself mid-project can silently change resource behavior
  • Treat terraform apply in CI as a gated, reviewed step, never auto-apply plans nobody has read, especially for anything touching production
  • Split state by blast radius (e.g. networking vs application resources in separate state files) once a single state file starts taking minutes to plan

Code Example

# main.tf — a minimal VPC + EC2 instance
resource "aws_vpc" "main" {
  cidr_block = "10.0.0.0/16"
}

resource "aws_subnet" "public" {
  vpc_id     = aws_vpc.main.id
  cidr_block = "10.0.1.0/24"
}

resource "aws_instance" "web" {
  ami           = "ami-0c55b159cbfafe1f0"
  instance_type = "t3.micro"
  subnet_id     = aws_subnet.public.id

  tags = { Name = "web-server" }
}

Code Example: The Plan/Apply Workflow

terraform init      # downloads providers, sets up the backend
terraform plan       # shows: 3 to add, 0 to change, 0 to destroy
terraform apply      # prompts for confirmation, then executes the plan
# Later, after editing main.tf to change instance_type:
terraform plan        # shows: 1 to change (or replace, depending on the attribute)

Ecosystem

  • Terragrunt — a thin wrapper around Terraform for keeping multi-environment configurations DRY, widely adopted despite being a separate, community-maintained tool
  • Atlantis — a self-hosted tool that runs plan/apply automatically on pull requests, turning infrastructure changes into a reviewable PR workflow
  • checkov / tfsec — static analysis tools that scan Terraform config for security misconfigurations before anything is ever applied
  • Terraform Registry — the public catalog of providers and community modules, the default place to find a starting module for common infrastructure patterns
  • OpenTofu — a Linux Foundation-backed fork of Terraform’s last MPL-licensed release, created in response to the 2023 BUSL license change, aiming to stay a drop-in open-source alternative
  • Sentinel / OPA — policy-as-code frameworks (HashiCorp’s own Sentinel, or the open-source Open Policy Agent) that let organizations enforce rules on plan output before apply runs, e.g. “no public S3 buckets,” blocking noncompliant infrastructure automatically rather than relying on manual review

Best Practices

  • Structure state and modules around team/service boundaries, not just technical resource type, so blast radius matches organizational ownership
  • Use remote state with locking from day one, even on a solo project, the habit matters more than the immediate need
  • Review every plan output before applying, especially destroy/replace operations, treating it with the same scrutiny as a code review diff
  • Pin provider versions with a lock file (terraform.lock.hcl, committed to version control) so terraform init is reproducible across machines and CI runs
  • Keep modules small and composable rather than one giant “do everything” module, mirroring good software engineering practice inside HCL
  • Adopt a policy-as-code check (Sentinel, OPA, or even a simple CI lint step) before infrastructure changes reach production, catching common misconfigurations mechanically instead of relying purely on human review

FAQ

Is Terraform the same as CloudFormation? No — CloudFormation is AWS-specific and native to that provider’s console/API, while Terraform is provider-agnostic and works the same way across AWS, GCP, Azure, and hundreds of others through its plugin providers.

Does Terraform configure software inside a server, like Ansible does? Not natively — Terraform provisions infrastructure (the VM exists, the network exists), while configuring what runs inside that VM is usually handed off to a tool like Ansible, a startup script, or a pre-baked machine image.

What happens if the state file is lost? Terraform loses track of what it manages; the underlying infrastructure still exists, but reconstructing state (via terraform import, resource by resource) is painful, which is exactly why remote state with versioning/backups is treated as non-negotiable for anything beyond a toy project.

Why did OpenTofu get created? HashiCorp changed Terraform’s license from open-source MPL to the more restrictive Business Source License in 2023, prompting a coalition of companies to fork the last MPL-licensed version under the Linux Foundation as OpenTofu, aiming to preserve a fully open-source path forward.

Can Terraform manage resources it didn’t originally create? Yes, via terraform import (or config-driven import blocks in newer versions), which brings an existing resource under Terraform’s management by writing it into state without recreating it — though a matching resource block still needs to be written or generated to match.

Common Interview Questions

  • “What’s the difference between terraform plan and terraform apply?” — expect an answer centered on plan being a read-only diff and apply being the step that actually executes changes
  • “How does Terraform handle infrastructure drift?” — expect discussion of state refresh detecting differences between the state file and real infrastructure, and plan surfacing that drift before any change is applied
  • “Why is committing a Terraform state file to version control a bad idea?” — expect an answer covering both merge-conflict risk on a JSON file and the security risk of state files often containing sensitive resource attributes in plaintext
  • “What’s the difference between a Terraform module and a provider?” — expect a clear distinction: a provider is a plugin that talks to a specific platform’s API (AWS, GCP), while a module is a reusable, user-authored bundle of resource configuration built using one or more providers

History

  • Released by HashiCorp in 2014, founded by Mitchell Hashimoto and Armon Dadgar, who had previously built Vagrant
  • Grew alongside the broader shift toward Infrastructure as Code and multi-cloud strategy through the mid-to-late 2010s, becoming a default choice well before most competitors matured
  • HashiCorp went public via IPO in December 2021
  • Changed Terraform’s license from the open-source Mozilla Public License to the Business Source License in August 2023, a controversial move that directly triggered the OpenTofu fork
  • IBM announced its acquisition of HashiCorp in 2024, adding further uncertainty to Terraform’s long-term licensing and roadmap that accelerated OpenTofu’s adoption in some organizations
  • OpenTofu reached its own 1.0 stable release in 2024, establishing itself as a genuinely independent project rather than a temporary stopgap while the community waited to see how the HashiCorp/IBM deal would play out

Dig deeper