Ansible Serial Keyword — Rolling Update Batch Size Strategy
Introduction
By default, Ansible runs tasks on all hosts simultaneously. The serial keyword changes this — it processes hosts in batches, enabling rolling updates, canary deployments, and zero-downtime upgrades. This is essential for production deployments where you can't take all servers offline at once.
Basic Usage
---
- name: Rolling update - 2 hosts at a time
hosts: webservers
serial: 2
become: true
tasks:
- name: Deploy application
ansible.builtin.git:
repo: https://github.com/org/app.git
dest: /opt/app
version: "{{ app_version }}"
- name: Restart service
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Wait for service to be healthy
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/health"
status_code: 200
retries: 10
delay: 5
register: health
until: health.status == 200
With 10 webservers: runs on hosts 1-2, waits for completion, then 3-4, then 5-6, etc.
Percentage-Based Batches
---
- name: Update 25% of fleet at a time
hosts: webservers
serial: "25%"
tasks:
- name: Upgrade packages
ansible.builtin.apt:
upgrade: safe
register: upgrade_result
- name: Reboot if needed
ansible.builtin.reboot:
reboot_timeout: 300
when: upgrade_result.changed
Canary Deployment Pattern
Start with 1 host, then ramp up:
---
- name: Canary deployment
hosts: webservers
serial:
- 1 # First: deploy to 1 host (canary)
- 5 # Then: 5 hosts at a time
- "100%" # Finally: all remaining
become: true
tasks:
- name: Deploy new version
ansible.builtin.git:
repo: https://github.com/org/app.git
dest: /opt/app
version: "{{ app_version }}"
- name: Restart application
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Verify deployment
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/health"
status_code: 200
retries: 12
delay: 5
until: result.status == 200
register: result
- name: Smoke test
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:8080/api/v1/status"
return_content: true
register: smoke
failed_when: "'version' not in smoke.content"
Fail-Safe with max_fail_percentage
Stop the entire deployment if too many hosts fail:
---
- name: Safe rolling deploy
hosts: webservers
serial: 3
max_fail_percentage: 30
become: true
tasks:
- name: Deploy and verify
block:
- name: Pull new code
ansible.builtin.git:
repo: https://github.com/org/app.git
dest: /opt/app
version: "{{ app_version }}"
- name: Restart
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Health check
ansible.builtin.uri:
url: "http://localhost:8080/health"
retries: 5
delay: 3
register: health
until: health.status == 200
rescue:
- name: Rollback this host
ansible.builtin.git:
repo: https://github.com/org/app.git
dest: /opt/app
version: "{{ previous_version }}"
- name: Restart with old version
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Mark host as failed
ansible.builtin.fail:
msg: "Deploy failed on {{ inventory_hostname }}"
Load Balancer Integration
Remove hosts from load balancer during updates:
---
- name: Rolling update with LB drain
hosts: webservers
serial: 2
become: true
pre_tasks:
- name: Disable host in HAProxy
community.general.haproxy:
state: disabled
host: "{{ inventory_hostname }}"
backend: app-backend
delegate_to: "{{ groups['loadbalancers'][0] }}"
- name: Wait for connections to drain
ansible.builtin.wait_for:
timeout: 30
tasks:
- name: Deploy application
ansible.builtin.git:
repo: https://github.com/org/app.git
dest: /opt/app
version: "{{ app_version }}"
- name: Restart application
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Wait for service to be ready
ansible.builtin.uri:
url: "http://localhost:8080/health"
retries: 10
delay: 3
register: health
until: health.status == 200
post_tasks:
- name: Re-enable host in HAProxy
community.general.haproxy:
state: enabled
host: "{{ inventory_hostname }}"
backend: app-backend
delegate_to: "{{ groups['loadbalancers'][0] }}"
- name: Wait for LB health check
ansible.builtin.pause:
seconds: 10
Serial vs Forks vs Throttle
| Keyword | Scope | Purpose |
|---|---|---|
serial | Play | Process N hosts before moving to next batch |
forks | Global | Max parallel SSH connections (default: 5) |
throttle | Task | Max parallel executions of a single task |
# These work together:
---
- name: Combined example
hosts: all # 100 hosts
serial: 10 # Process 10 at a time
# forks = 5 in ansible.cfg → 5 of the 10 run in parallel
tasks:
- name: API call (rate limited)
ansible.builtin.uri:
url: "https://api.example.com/register"
method: POST
throttle: 2 # Only 2 API calls at once (within the batch of 10)
any_errors_fatal with Serial
---
- name: Critical deployment - stop on any failure
hosts: webservers
serial: 2
any_errors_fatal: true # If any host in the batch fails, stop everything
tasks:
- name: Deploy critical update
ansible.builtin.command:
cmd: /opt/deploy.sh
Run Once per Batch
- name: Notify monitoring (once per batch, not per host)
ansible.builtin.uri:
url: "https://monitoring.example.com/api/maintenance"
method: POST
body_format: json
body:
message: "Deploying batch {{ ansible_play_batch | join(',') }}"
run_once: true
delegate_to: localhost
Troubleshooting
| Issue | Solution |
|---|---|
| Too slow | Increase serial value or forks |
| Failures stop everything | Use max_fail_percentage instead of any_errors_fatal |
| Need rollback | Add block/rescue pattern within serial play |
| Hosts out of LB too long | Reduce drain wait time, optimize deploy tasks |
| Wrong batch ordering | Use order: sorted or order: shuffle on the play |
Best Practices
- Always health-check after deploy —
urimodule with retries - Canary first —
serial: [1, 5, "100%"]catches issues early - Integrate with load balancer — drain before update, re-enable after
- Set
max_fail_percentage— don't deploy to all hosts if some fail - Use
order: shuffle— randomize batch order to spread risk - Combine with
throttle— rate-limit API calls within batches
Conclusion
The serial keyword is your primary tool for zero-downtime deployments. Combined with canary patterns, load balancer integration, and max_fail_percentage, you can safely deploy to thousands of hosts with automatic rollback on failure. Start with conservative batches and increase as confidence grows.