Ansible
Ansible
Definition: A configuration management and automation tool that connects to servers over SSH and runs tasks defined in simple YAML “playbooks,” with no agent installation required on the machines it manages. Created by Michael DeHaan and released in 2012, it was acquired by Red Hat in 2015 and remains the most widely adopted agentless automation tool for configuring, patching, and deploying software across fleets of existing servers. Where Terraform focuses on provisioning the infrastructure itself, Ansible typically takes over once that infrastructure exists — installing packages, managing config files, and rolling out application deployments.
Core Services & Concepts
- Playbooks — YAML files describing an ordered sequence of tasks to run against one or more groups of target machines, the primary unit of Ansible automation
- Inventory — a list (static file or dynamically generated from a cloud provider’s API) of the hosts Ansible manages, organized into groups like
webserversordb - Modules — the individual units of work a task invokes (
apt,copy,service, hundreds more), each written to check current state before acting so it only makes changes when needed - Idempotency — the property that re-running the same playbook against a host that’s already correctly configured produces no changes, making playbooks safe to run repeatedly rather than only once
- Roles — a standardized directory structure for bundling related tasks, handlers, templates, and variables into a reusable, shareable unit
- Handlers — tasks that only fire when notified by another task that reported a change, deduplicated and run once at the end of a play (e.g. “restart nginx” only if its config actually changed)
- Jinja2 templating — the templating engine used throughout playbooks and config file templates, letting variables, conditionals, and loops drive dynamic output
- Facts — system information (OS, IP addresses, memory, CPU) automatically gathered from each host at the start of a play and exposed as variables for tasks to use
- Ansible Vault — built-in encryption for sensitive variables (passwords, API keys) so secrets can live in version control alongside the rest of a playbook
How It Works: The Agentless Push Model
- The control node connects to each managed node over SSH (or WinRM for Windows targets) using existing credentials, no persistent daemon or agent needs to be installed or kept running on the target
- For each task, Ansible copies a small, self-contained Python script (the module) to the target over SSH, executes it, captures the JSON result, then removes it, leaving nothing resident behind
- Hosts are processed in parallel up to a configurable fork limit (5 by default), while within a single host, tasks execute strictly in the order written in the playbook
- Idempotency is enforced at the module level — the
aptmodule checks whether a package is already installed before attempting to install it, so a “no change” run reports nothing modified - Facts gathering (the implicit
setupmodule) runs automatically at the start of a play unless disabled, giving every subsequent task access to variables likeansible_distributionoransible_default_ipv4 - Handlers queue during a play but only actually execute after all regular tasks finish, and only once even if multiple tasks notify the same handler, avoiding redundant restarts
How Pricing Works
- Ansible Core (the CLI and open-source module library) is free and open-source under the GPLv3 license, no cost to install or run
- Red Hat Ansible Automation Platform — the commercial successor to Ansible Tower — adds a web UI, role-based access control, job scheduling, and centralized logging, priced by subscription tied to the number of managed nodes
- AWX is the free, upstream open-source project that Ansible Automation Platform is built from, giving teams the same UI/API layer without a support contract
- Because Ansible configures existing infrastructure rather than provisioning new resources, it doesn’t generate direct cloud billing the way running
terraform applydoes, though playbooks that call cloud modules can create billable resources - Most real-world cost is engineering time spent writing and maintaining playbooks/roles, plus an optional Automation Platform subscription for larger, governed deployments
Pros
- No agents to install, update, or keep running on target machines, reducing both attack surface and operational overhead
- Simple, readable YAML syntax gives it a notably lower learning curve than tools requiring a dedicated DSL or full programming language
- Huge module library — several thousand modules covering cloud providers, network devices, databases, and general Linux/Windows administration
- Works equally well for ad hoc one-off commands (
ansible all -m ping) and for structured, repeatable playbooks - Strong fit as the “glue” layer between infrastructure provisioning and application deployment, commonly paired directly with Terraform
- Roles and Ansible Galaxy make it easy to reuse and share configuration logic across projects and teams
Cons
- SSH-based execution scales less gracefully than agent/pull-based tools at very large fleets (many thousands of hosts), where fork limits and per-host connection overhead start to matter
- YAML plus Jinja2 templating can become difficult to read once conditionals, loops, and nested variables pile up, edging toward the complexity of a real programming language without the tooling
- Slower per-task execution than compiled, agent-resident tools, since every task round-trips a module over SSH
- Ansible Vault is functional but basic compared to dedicated secrets managers, larger orgs often layer HashiCorp Vault or a cloud KMS on top
- Historically weaker, slower Windows support via WinRM compared to the maturity of its SSH-based Linux automation
Comparison: Ansible vs Terraform vs Puppet/Chef
| Ansible | Terraform | Puppet / Chef | |
|---|---|---|---|
| Primary purpose | Configuration management + ad hoc automation | Declarative infrastructure provisioning | Configuration management, agent-based |
| Architecture | Agentless, push over SSH | Agentless, talks to provider APIs | Agent-based, typically pull-based |
| Language | YAML playbooks | HCL (declarative) | Puppet DSL / Ruby (Chef) |
| State tracking | Stateless — checks actual system each run | Explicit state file | Agent reports state to a central server |
| Best fit | Configuring servers after they exist, app deployment | Provisioning cloud resources across providers | Large, long-lived fleets needing continuous enforcement |
Best For
- Configuring, patching, and maintaining fleets of servers after they’ve already been provisioned
- Orchestrating multi-tier application deployments, including rolling updates across a group of hosts
- Quick, ad hoc automation across many machines without setting up a full agent-based system first
Real Examples
- Commonly deployed at organizations managing large hybrid or on-prem fleets, including reported use at NASA JPL for automation workflows
- A default choice for network automation, with dedicated modules for Cisco, Juniper, and other network vendor devices
- Red Hat uses it as the automation backbone underneath its own Ansible Automation Platform and much of its enterprise Linux tooling story
Use Cases
- Server configuration management — installing packages, managing users, editing config files consistently across a fleet
- Application deployment automation, including coordinated rolling restarts across web tiers
- Patch management and security updates applied uniformly across many hosts on a schedule
- Network device automation — pushing configuration changes to routers and switches through the same playbook model
- Security compliance and hardening, running CIS-benchmark-style playbooks to enforce a baseline configuration
Integration Notes & Common Pitfalls
- Keep inventory dynamic when targets live in the cloud, a static inventory file drifts out of date the moment autoscaling adds or removes instances
- Prefer purpose-built modules over
shell/commandtasks wherever one exists, hand-rolled shell steps are rarely truly idempotent and can silently break re-runs - Store secrets in Ansible Vault (or an external secrets manager referenced via lookup plugin), never as plaintext variables committed to version control
- Pin collection and role versions in
requirements.yml, an unpinned collection updating mid-project can change module behavior unexpectedly - Tune the
forkssetting and connection pipelining for large inventories, the default of 5 parallel hosts is conservative and often too slow for fleets in the hundreds - Use
--checkand--diffto dry-run a playbook against production before actually applying it, catching unintended changes before they happen
Code Example
# site.yml — install and configure nginx across the web group
- name: Configure web servers
hosts: webservers
become: true
vars:
nginx_port: 80
tasks:
- name: Install nginx
apt:
name: nginx
state: present
update_cache: true
- name: Deploy nginx config from template
template:
src: templates/nginx.conf.j2
dest: /etc/nginx/nginx.conf
notify: Restart nginx
- name: Ensure nginx is running and enabled
service:
name: nginx
state: started
enabled: true
handlers:
- name: Restart nginx
service:
name: nginx
state: restarted
Code Example: Ad Hoc Commands and Vault Workflow
# Quick ad hoc command, no playbook needed
ansible webservers -i inventory.ini -m ping
# Check disk usage across an entire group
ansible db -m shell -a "df -h"
# Encrypt a secrets file with Ansible Vault
ansible-vault encrypt group_vars/prod/secrets.yml
# Run a playbook, prompting for the vault password, dry-run first
ansible-playbook site.yml --ask-vault-pass --check --diff
# Then apply for real, limited to one host group
ansible-playbook site.yml --ask-vault-pass --limit webservers
Ecosystem
- Ansible Galaxy — the public hub for sharing and downloading community roles and collections, the default starting point before writing a role from scratch
- AWX / Ansible Automation Platform — the web UI, API, and RBAC layer built on top of Ansible Core, AWX being the free upstream project and Automation Platform its supported commercial form
- Molecule — a testing framework purpose-built for validating roles in isolated containers before they’re trusted against real infrastructure
- Collections — the namespaced packaging format (introduced in Ansible 2.10) that split modules out of Ansible Core into independently versioned, installable bundles
- Execution Environments — containerized, reproducible run environments (built with
ansible-builder) that package Ansible plus exactly the collections and dependencies a given automation job needs
Best Practices
- Organize automation into roles rather than one sprawling playbook, mirroring good software structure inside YAML
- Favor idempotent modules over
shell/commandtasks whenever a dedicated module exists for the job - Keep secrets in Vault or an external secrets manager, never committed to version control in plaintext
- Pin role and collection versions in
requirements.ymlsoansible-galaxy installis reproducible across machines and CI - Tag tasks (
tags: [nginx, config]) so selective runs are possible without executing an entire playbook for a small change - Always dry-run against production with
--check --diffbefore applying, treating it as the equivalent of Terraform’s plan step
FAQ
Does Ansible need an agent installed on managed machines? No — it connects over standard SSH (or WinRM for Windows) and pushes modules temporarily, which is the core distinction from agent-based tools like Puppet or Chef.
How is Ansible different from Terraform? Terraform provisions and manages the infrastructure itself (VMs, networks, load balancers), while Ansible typically configures what runs inside infrastructure that already exists — the two are commonly used together, not as substitutes for each other.
How does Ansible handle secrets? Through Ansible Vault, which encrypts variable files or specific values so they can be safely stored in version control and decrypted only at run time with a vault password or key.
Can Ansible manage Windows servers? Yes, via WinRM instead of SSH, though Windows support has historically been slower and less feature-complete than its Linux/SSH automation.
What’s the difference between a role and a playbook? A playbook is the entry point that defines which hosts to target and what to run, while a role is a reusable, self-contained bundle of tasks, templates, and variables that a playbook can include.
Common Interview Questions
- “What does idempotency mean in the context of Ansible, and why does it matter?” — expect an answer covering safe repeated execution and how individual modules check state before acting
- “Explain Ansible’s agentless architecture and its trade-offs” — expect discussion of SSH-based push execution versus agent-based pull models, and how that affects scale
- “What’s the difference between an ad hoc command and a playbook?” — expect a distinction between one-off
ansibleCLI commands and structured, repeatable YAML playbooks - “How do handlers work, and why not just put a restart task directly after a config change?” — expect an explanation of notify/handler deduplication running once at the end of a play
History
- Created by Michael DeHaan and released in 2012, with the name borrowed from Ursula K. Le Guin’s science fiction, where an “ansible” is a device enabling instantaneous communication across space
- Ansible, Inc. was acquired by Red Hat in October 2015, folding it into Red Hat’s broader automation and enterprise Linux strategy
- Ansible Tower, the original commercial web UI/API layer, launched to give enterprise teams centralized scheduling, RBAC, and logging on top of the open-source core
- Ansible 2.10 (2020) restructured the project, splitting most modules out of the core package into independently versioned Collections
- AWX was open-sourced as the free upstream project underlying Tower, letting the community run the same UI/API layer without a support subscription
- Red Hat rebranded its commercial offering as Red Hat Ansible Automation Platform, consolidating Tower, execution environments, and content management under one product name
Related Terms
Referenced by