Introduction
The serial keyword controls how many hosts Ansible updates at once. Using percentages (serial: "25%") or stepped batches (serial: [1, 5, "50%"]) enables rolling deployments, canary releases, and risk-controlled updates. Without serial, Ansible runs on all hosts simultaneously — fine for configuration, risky for deployments.
Basic Percentage
---
- name: Rolling update — 25% at a time
hosts: webservers
serial: "25%"
tasks:
- name: Deploy application
ansible.builtin.copy:
src: app.tar.gz
dest: /opt/app/app.tar.gz
- name: Restart service
ansible.builtin.systemd:
name: myapp
state: restarted
With 20 hosts and serial: "25%", Ansible runs 5 hosts per batch (4 batches total).
Percentage Rounding
| Hosts | serial | Batch Size | Batches |
|---|---|---|---|
| 10 | "25%" | 3 (rounds up) | 4 |
| 10 | "30%" | 3 | 4 |
| 10 | "50%" | 5 | 2 |
| 100 | "10%" | 10 | 10 |
| 7 | "25%" | 2 (rounds up) | 4 |
| 1 | "50%" | 1 (minimum 1) | 1 |
Percentages always round up to at least 1.
Fixed Batch Size
# Run 5 hosts at a time
- name: Batch update
hosts: all
serial: 5
tasks:
- name: Update packages
ansible.builtin.apt:
upgrade: dist
Stepped Batches (Canary Pattern)
# Start small, grow confidence, finish fast
- name: Canary deployment
hosts: webservers
serial:
- 1 # First: 1 host (canary)
- 5 # Then: 5 hosts
- "25%" # Then: 25% of remaining
- "100%" # Finally: all remaining
tasks:
- name: Deploy new version
ansible.builtin.copy:
src: app-v2.tar.gz
dest: /opt/app/
- name: Restart
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Health check
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/health"
status_code: 200
register: health
until: health.status == 200
retries: 10
delay: 5
With 100 hosts:
- Batch 1: 1 host (canary)
- Batch 2: 5 hosts
- Batch 3: 24 hosts (25% of 94 remaining)
- Batch 4: 70 hosts (all remaining)
Fail Fast with max_fail_percentage
- name: Safe rolling update
hosts: webservers
serial: "20%"
max_fail_percentage: 10
tasks:
- name: Deploy
ansible.builtin.copy:
src: app.tar.gz
dest: /opt/app/
- name: Health check
ansible.builtin.uri:
url: "http://localhost:8080/health"
status_code: 200
retries: 5
delay: 3
If >10% of a batch fails, the entire play aborts — preventing a bad deploy from spreading.
Zero-Downtime with Load Balancer
- name: Zero-downtime deploy
hosts: webservers
serial: "25%"
max_fail_percentage: 0 # Any failure stops everything
pre_tasks:
- name: Remove from LB
ansible.builtin.uri:
url: "https://lb.example.com/api/remove/{{ inventory_hostname }}"
method: POST
delegate_to: localhost
- name: Drain connections
ansible.builtin.pause:
seconds: 30
roles:
- deploy_app
post_tasks:
- name: Health check
ansible.builtin.uri:
url: "http://localhost:8080/health"
register: check
until: check.status == 200
retries: 12
delay: 5
- name: Add back to LB
ansible.builtin.uri:
url: "https://lb.example.com/api/add/{{ inventory_hostname }}"
method: POST
delegate_to: localhost
order Directive
# Control host processing order within each batch
- name: Ordered rolling update
hosts: webservers
serial: 2
order: sorted # alphabetical order
tasks:
- name: Update
ansible.builtin.debug:
msg: "Updating {{ inventory_hostname }}"
| Order | Description |
|---|---|
inventory | Default — order defined in inventory |
sorted | Alphabetical by hostname |
reverse_sorted | Reverse alphabetical |
shuffle | Random order |
reverse_inventory | Reverse of inventory order |
Comparison of Strategies
| Strategy | serial | Use Case | Risk |
|---|---|---|---|
| All at once | (not set) | Config management | High |
| Fixed batch | 5 | Known cluster size | Medium |
| Percentage | "25%" | Scales with fleet | Medium |
| Canary | [1, 5, "50%"] | Production deploys | Low |
| One-at-a-time | 1 | Database migrations | Lowest |
Troubleshooting
| Issue | Solution |
|---|---|
Too slow with serial: 1 | Increase batch or use percentage |
| All hosts fail at once | Add serial — you're running without it |
| Batch too small with percentage | Percentage rounds up; use fixed number instead |
| Remaining hosts don't run after failure | Expected — max_fail_percentage or any failure stops subsequent batches |
| Need different serial per play | Use multiple plays with different serial values |
Best Practices
- Always use
serialfor deployments — never deploy to all hosts at once - Canary pattern for production —
[1, 5, "25%", "100%"] - Combine with
max_fail_percentage— stop early on failures - Health checks after restart —
uri+untilloop - LB integration in pre/post_tasks — drain before, add after
- Use
order: shuffle— avoid always hitting the same hosts first
Conclusion
serial with percentages and stepped batches is how you do safe production deployments. Start with a canary (1 host), grow the batch size as confidence builds, and combine with max_fail_percentage to stop the blast radius. The pattern serial: [1, 5, "25%", "100%"] with health checks is the gold standard for rolling updates.