Roll Out Linux Patches with Ansible
Run this example from a Linux control node with Python 3 and its virtual-environment support installed, SSH access to the hosts, and Python 3 on each target. Ansible can make Linux patching repeatable, but a safe rollout still needs a canary, a maintenance window, and a recovery path. This example updates Debian or Ubuntu hosts in one-host batches, checks a service after package work, and reboots only when the host reports that one is required.
Start with a non-production canary and a tested backup or snapshot. The playbook fails on an unreachable host rather than silently skipping it.
Start With a Canary Group
InventorySave this inventory as inventory.ini. Keep the first run limited to a canary. Use approved SSH credentials, validate privilege escalation, and confirm the target OS is Debian-family before using APT.
python3 -m venv "$HOME/ansible-control"source "$HOME/ansible-control/bin/activate"python -m pip install "ansible-core>=2.11"ansible --version
cat > inventory.ini <<'EOF'[canary]app01 ansible_host=10.20.40.15
[linux_patch_wave1]app01app02 ansible_host=10.20.40.16EOF
ansible -i inventory.ini canary -m ansible.builtin.ping❯ View Expected Console Output
app01 | SUCCESS => {"changed": false, "ping": "pong"}
Figure 1: Review the canary result before widening the rollout.
Patch One Host at a Time and Check Service Health
Controlled RolloutSave as patch-debian.yml. Use ansible-core 2.11 or newer for the package-removal guard below. The serial setting limits concurrent hosts. Replace the sample service check with one that represents each server’s role.
---- name: Patch Debian-family Linux hosts hosts: linux_patch_wave1 become: true serial: 1 max_fail_percentage: 0 tasks: - name: Refresh APT package cache ansible.builtin.apt: update_cache: true cache_valid_time: 3600
- name: Upgrade installed packages ansible.builtin.apt: upgrade: dist fail_on_autoremove: true
- name: Check whether a reboot is required ansible.builtin.stat: path: /var/run/reboot-required register: reboot_required
- name: Reboot when requested by the operating system ansible.builtin.reboot: msg: Rebooting after approved patch maintenance reboot_timeout: 900 when: reboot_required.stat.exists
- name: Gather service state ansible.builtin.service_facts:
- name: Verify the application service is active ansible.builtin.assert: that: - ansible_facts.services['nginx.service'].state == 'running' fail_msg: "nginx.service is not running after patching"❯ View Expected Console Output
PLAY RECAPapp01 : ok=5 changed=1 unreachable=0 failed=0Review Check Mode Before Applying Patches
Change ReviewCheck mode previews the canary plan without applying it, but it cannot predict every package-manager action. Review the target list and output, confirm the maintenance window and recovery snapshot, then start the separate apply step only when approved.
ansible-playbook -i inventory.ini patch-debian.yml --limit canary --check --diff❯ View Expected Console Output
Check mode completed. Review the predicted changes before continuing.Apply the Approved Canary Patch
CanaryAfter reviewing check mode and confirming the change window, apply the playbook to the canary only. The APT removal guard makes the task fail if the upgrade would remove packages; review that failure with the package owner before changing the plan. Confirm the service check and application health before expanding the rollout.
ansible-playbook -i inventory.ini patch-debian.yml --limit canary❯ View Expected Console Output
Canary completed. Confirm the application health check before continuing.Proceed in Approved Batches
RolloutAfter the canary owner confirms application health, run the next approved group during its maintenance window. Record the inventory revision, playbook revision, changed hosts, reboot outcomes, and failed checks. pipefail ensures a failed Ansible run is not hidden by tee.
set -o pipefailansible-playbook -i inventory.ini patch-debian.yml --limit 'linux_patch_wave1:!canary' --forks 1 --diff | tee patch-wave1.log❯ View Expected Console Output
Each host reports ok, changed, unreachable, and failed task counts.Stop the rollout if a health check fails or a host becomes unreachable.