Ansible is a tool for describing how servers should be configured and then making them so. You write the desired state in YAML (which packages are installed, which files contain what, which services run), and Ansible connects to each machine over SSH, compares reality with that description and changes only what differs. There is no agent to install on the machines you manage, which is why teams can adopt it in an afternoon.
This introduction goes from first principles to a safe production run: how a run actually executes, inventories, modules, playbooks, idempotence, variables, roles, secrets, rolling updates, testing and the mistakes that bite newcomers. It assumes you can use SSH and a Linux shell.
The problem it solves
Configuring one server by hand is fine. Configuring forty identically, and keeping them identical while people make urgent fixes, is not. Shell scripts help but have a flaw: they describe steps, not state. Run a script that appends a line to a file twice and you get the line twice. Ansible's modules describe the end state instead, such as 'this line is present in this file', and do nothing when it already holds. That property, idempotence, is what lets you run the same playbook every day as a check and a fix.
Ansible sits beside other tools rather than replacing them. Provisioning tools create cloud resources such as networks and virtual machines; Ansible configures what runs on them. Container images, described in the Docker introduction, bake configuration into an artefact, and orchestrators such as Kubernetes run them. Ansible remains the natural fit for long-lived hosts, network devices, the nodes underneath a cluster, and one-off operational tasks across a fleet.
How a run works
The machine where you run Ansible is the control node; the machines it configures are managed nodes. For each task, Ansible renders the module code with its arguments, copies it to a temporary directory on the managed node over SSH, runs it with the node's Python interpreter, reads back a JSON result and deletes the temporary file. Most modules therefore need Python on the managed node; the raw module, which sends a plain command over SSH, is the exception and is how you bootstrap Python onto a bare machine.
By default Ansible works on 5 hosts in parallel (the forks setting) and uses the linear strategy: each task finishes on every targeted host before the next task starts. Hosts that fail are removed from the rest of the play, while the others continue. Knowing this explains most of the behaviour you will see in output.
Inventory: which machines
The inventory lists hosts and groups them. It can be a static INI or YAML file, or a dynamic inventory plugin that queries a cloud provider so newly launched instances appear automatically.
# inventory/production.ini
[web]
web1.example.com
web2.example.com
[db]
db1.example.com ansible_user=admin
[production:children]
web
dbGroups can contain other groups through :children, and every host is also in the implicit all group. Variables for a group go in group_vars/web.yml and for one host in host_vars/web1.example.com.yml, next to the inventory. Run ansible-inventory -i inventory/production.ini --graph to see how Ansible understood your groups before you target anything.
Modules and ad-hoc commands
A module is a unit of work: install a package, write a file, manage a user, call an HTTP endpoint. Modules are named with fully qualified collection names such as ansible.builtin.copy or community.general.ufw, which removes any ambiguity about which implementation runs. Each returns whether it changed anything, whether it failed, and data you can use later.
You can run a single module across hosts without writing a playbook. These ad-hoc commands are ideal for checks and one-off actions:
ansible all -i inventory/production.ini -m ansible.builtin.ping
ansible web -i inventory/production.ini -b -m ansible.builtin.apt -a "name=htop state=present"
ansible db -i inventory/production.ini -m ansible.builtin.command -a "uptime"The command and shell modules run arbitrary commands, and Ansible cannot know whether they changed anything, so they always report changed unless you tell it otherwise. Prefer a purpose-built module whenever one exists.
Playbooks: desired state as code
A playbook is a YAML list of plays. Each play maps a group of hosts to an ordered list of tasks. This one installs nginx, writes its configuration from a template and keeps the service running:
# site.yml
- name: Configure web servers
hosts: web
become: true
vars:
app_port: 8080
tasks:
- name: Install nginx
ansible.builtin.apt:
name: nginx
state: present
update_cache: true
- name: Write site configuration
ansible.builtin.template:
src: templates/site.conf.j2
dest: /etc/nginx/conf.d/site.conf
mode: "0644"
notify: Reload nginx
- name: Ensure nginx is running and enabled
ansible.builtin.service:
name: nginx
state: started
enabled: true
handlers:
- name: Reload nginx
ansible.builtin.service:
name: nginx
state: reloadedbecome: true runs tasks with privilege escalation, by default through sudo. The template module renders a Jinja2 file, so {{ app_port }} inside it becomes 8080. The notify line triggers a handler, a task that runs once after the current block of tasks finishes (or earlier with meta: flush_handlers) and only if something notified it. Change the configuration and nginx reloads once; change nothing and it does not reload at all, which is exactly the behaviour you want on a production server.
Run it with ansible-playbook -i inventory/production.ini site.yml. The first run reports changes; the second should report everything as ok and nothing changed.
Idempotence, check mode and diff
That second run is the real test of a playbook. If tasks still report changed on an unchanged system, something is not idempotent: usually a command or shell task. Fix it with creates: or removes: arguments, which skip the command when a file exists or does not, or with changed_when: and failed_when: expressions that inspect the output. The idempotency article explains why repeatable operations matter in distributed systems generally.
Before touching production, run ansible-playbook site.yml --check --diff. Check mode asks each module to report what it would change without changing it, and diff shows the file differences. Modules that cannot predict their effect, including command, are skipped in check mode, so a clean check is good evidence, not proof.
Variables and precedence
Variables come from many places: role defaults, inventory, group_vars and host_vars, play vars, facts gathered from each host (such as ansible_facts['distribution']), registered task results, and extra variables passed with -e on the command line. When the same name is defined in several places, a documented precedence order decides; role defaults are the weakest and extra variables always win.
The full order has more than twenty levels, and memorising it is not the answer. Keep each variable defined in one obvious place, put tunable values in role defaults, put environment differences in group variables, and reserve -e for one-off overrides. When a value surprises you, ansible-inventory -i inventory/production.ini --host web1.example.com shows the inventory-level variables a host receives, and a temporary ansible.builtin.debug task inside the play shows the final value after play variables are applied.
Roles and collections
Once a playbook grows, split it into roles: directories with a fixed layout that Ansible loads automatically. A role named app has tasks/main.yml, handlers/main.yml, templates/, files/, defaults/main.yml for overridable values and meta/main.yml for dependencies. A play then just lists roles: [app].
Collections are the distribution format for modules, plugins and roles, published on Ansible Galaxy or a private hub. Pin the ones you use in requirements.yml and install them with ansible-galaxy collection install -r requirements.yml, so every machine running your playbooks uses the same module versions.
Secrets with Vault
Passwords and API keys must not sit in plain text in Git. Ansible Vault encrypts whole files or single values with a password, and decrypts them in memory during a run. ansible-vault encrypt_string 'S3cret' --name db_password produces a block you paste into a variables file; run playbooks with --ask-vault-pass or a vault password file supplied by your CI system. Add no_log: true to tasks that handle secrets, because by default Ansible prints module arguments and results when they fail or when output is verbose.
Rolling changes safely
Pushing a change to every host at once means a bad change breaks every host at once. The serial keyword runs the whole play on a batch of hosts before moving to the next batch, and max_fail_percentage aborts the run when too many hosts in a batch fail. Combined with draining each host from the load balancer first, this gives a rolling deployment:
- name: Rolling upgrade of the app tier
hosts: web
become: true
serial: 2 # two hosts at a time
max_fail_percentage: 0 # stop the whole run if any host in a batch fails
pre_tasks:
- name: Drain host from the load balancer
ansible.builtin.command: /usr/local/bin/lb-drain {{ inventory_hostname }}
delegate_to: lb1.example.com
changed_when: true
roles:
- app
post_tasks:
- name: Wait for health check
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/health"
status_code: 200
register: health
retries: 10
delay: 3
until: health.status == 200
- name: Return host to the load balancer
ansible.builtin.command: /usr/local/bin/lb-enable {{ inventory_hostname }}
delegate_to: lb1.example.com
changed_when: truedelegate_to runs a task on a different machine on behalf of the current host, here the load balancer. Set serial so that the hosts removed at once never exceed your spare capacity.
Reading the recap
Every run ends with a PLAY RECAP line per host listing counts of ok, changed, unreachable, failed, skipped, rescued and ignored. In a scheduled run on a stable fleet, any non-zero changed count means something drifted, someone changed a host by hand, or a task is not idempotent; each deserves a look. Unreachable usually means SSH or DNS, not your playbook.
Failure modes
- Shell everywhere. Playbooks that are scripts in YAML lose idempotence and check mode. Use modules.
- Drift outside Ansible. Manual fixes are overwritten on the next run or, worse, never captured. Run playbooks on a schedule and treat changes as alerts.
- Precedence surprises. A host variable silently overrides a group value. Define each variable once.
- Missing Python. Minimal images fail on the first module. Bootstrap with
rawor use images with Python. - Slow runs. Fact gathering and many SSH round trips dominate. Set
gather_facts: falsewhere facts are unused, raiseforks, and enable SSH pipelining. - Leaked secrets. Failed tasks print arguments. Use Vault and
no_log.
Testing and CI
Treat playbooks like code. Run ansible-playbook --syntax-check and ansible-lint on every pull request; lint catches non-idempotent patterns, missing names and deprecated syntax. Molecule tests a role by creating a throwaway container or VM, applying the role, applying it again to prove nothing changes, and running assertions. Then apply to staging with --check --diff before production. The CI/CD article shows where these stages fit in a pipeline.
When Ansible is the wrong tool
Ansible pushes changes when someone runs it, so it does not continuously enforce state the way an agent-based tool or a Kubernetes controller does; if drift must be corrected within minutes, schedule runs or use a reconciling system. It is also a poor fit for creating and destroying cloud resources with dependencies between them, where a provisioning tool that tracks state plans changes more safely, and for immutable fleets where hosts are replaced rather than modified. In those setups Ansible often still has a job: building the image that gets deployed.
What to do next
- Install Ansible on a control machine and write an inventory for two test hosts; confirm with
ansible all -m ansible.builtin.ping. - Write the nginx playbook above, run it twice and confirm the second run reports no changes.
- Replace any
commandorshelltask with a module, or addcreatesorchanged_when. - Move the playbook into a role with defaults, and pin collections in
requirements.yml. - Encrypt every secret with Vault and mark secret-handling tasks with
no_log. - Add syntax check, ansible-lint and a Molecule idempotence test to CI, then roll out with
serial.