Ansible max_fail_percentage — Control Failure Tolerance
Introduction
max_fail_percentage sets the maximum percentage of hosts that can fail before Ansible aborts the entire play. Combined with serial, it creates a safety valve for rolling deployments: if too many hosts fail in a batch, the play stops before the bad deploy reaches your entire fleet.
Basic Usage
---
- name: Safe rolling deploy
hosts: webservers
serial: "25%"
max_fail_percentage: 10
tasks:
- name: Deploy application
ansible.builtin.copy:
src: app.tar.gz
dest: /opt/app/
- name: Restart service
ansible.builtin.systemd:
name: myapp
state: restarted
- name: Health check
ansible.builtin.uri:
url: "http://localhost:8080/health"
status_code: 200
retries: 5
delay: 3
With 100 hosts, serial: "25%", and max_fail_percentage: 10:
- Each batch = 25 hosts
- If >2 hosts (10% of 25) fail in a batch → entire play aborts
- Remaining batches won't run
How It Works
Batch 1 (25 hosts): 24 pass, 1 fails (4%) → Continue ✅
Batch 2 (25 hosts): 22 pass, 3 fails (12%) → ABORT ❌
Batch 3 (25 hosts): Never runs
Batch 4 (25 hosts): Never runs
The percentage is calculated per batch, not across the entire play.
Common Settings
| Setting | Effect | Use Case |
|---|---|---|
max_fail_percentage: 0 | Any failure aborts | Critical infrastructure |
max_fail_percentage: 10 | 10% tolerance | Standard deployments |
max_fail_percentage: 25 | 25% tolerance | Flaky environments |
max_fail_percentage: 49 | Nearly half can fail | Non-critical updates |
max_fail_percentage: 100 | Never aborts (default) | Fire-and-forget |
Zero Tolerance
- name: Database migration (zero tolerance)
hosts: database
serial: 1
max_fail_percentage: 0
tasks:
- name: Backup database
ansible.builtin.command:
cmd: pg_dump myapp > /backups/pre-migrate.sql
- name: Run migration
ansible.builtin.command:
cmd: /opt/myapp/bin/migrate
- name: Verify schema
ansible.builtin.command:
cmd: /opt/myapp/bin/migrate --check
changed_when: false
Canary with Failure Gates
- name: Canary deploy with failure gates
hosts: webservers
serial: [1, 5, "25%", "100%"]
max_fail_percentage: 0 # Zero tolerance on canary
pre_tasks:
- name: Remove from load balancer
ansible.builtin.uri:
url: "https://lb.example.com/api/remove/{{ inventory_hostname }}"
method: POST
delegate_to: localhost
roles:
- deploy_app
post_tasks:
- name: Health check
ansible.builtin.uri:
url: "http://localhost:8080/health"
register: health
until: health.status == 200
retries: 12
delay: 5
- name: Add back to load balancer
ansible.builtin.uri:
url: "https://lb.example.com/api/add/{{ inventory_hostname }}"
method: POST
delegate_to: localhost
If the canary (first host) fails, the entire deploy stops.
any_errors_fatal vs max_fail_percentage
# any_errors_fatal: ALL hosts stop if ANY host fails
- name: All-or-nothing
hosts: webservers
any_errors_fatal: true
tasks:
- name: Critical task
ansible.builtin.command:
cmd: /opt/app/check
# max_fail_percentage: threshold-based
- name: Threshold-based
hosts: webservers
serial: 10
max_fail_percentage: 20
tasks:
- name: Tolerant task
ansible.builtin.command:
cmd: /opt/app/update
| Feature | any_errors_fatal | max_fail_percentage: 0 |
|---|---|---|
| Scope | All hosts immediately | Per batch with serial |
| Tolerance | Zero (any failure) | Zero (but per batch) |
| Without serial | Stops all hosts | Same as default |
| Best for | Pre-flight checks | Rolling deployments |
Calculation Examples
20 hosts, serial: 5, max_fail_percentage: 30
→ Batch size: 5
→ Max failures per batch: 1 (30% of 5 = 1.5, rounds down to 1)
→ If 2+ hosts fail in a batch → abort
100 hosts, serial: "10%", max_fail_percentage: 20
→ Batch size: 10
→ Max failures per batch: 2 (20% of 10)
→ If 3+ hosts fail in a batch → abort
50 hosts, serial: 1, max_fail_percentage: 0
→ Batch size: 1
→ Any failure → abort immediately
Troubleshooting
| Issue | Solution |
|---|---|
| Play doesn't stop on failures | Check max_fail_percentage is set (default is 100) |
| Too aggressive — stops on 1 failure | Increase percentage or batch size |
| Failures in one batch, success in next | That's expected — each batch is independent |
| Need to continue after abort | Fix the issue, re-run; Ansible skips successful hosts with --limit @retry_file |
max_fail_percentage without serial | Works but applies to entire host list as one "batch" |
Best Practices
- Always pair with
serial— meaningless without batching - Use 0% for canary deploys — any failure should stop propagation
- Use 10-20% for standard rolling updates — allows for flaky hosts
- Add health checks —
uri+untilto catch delayed failures - Combine with
--limit @playbook.retry— re-run only failed hosts - Log failures — use
block/rescueto capture what went wrong before aborting
Conclusion
max_fail_percentage is your deployment circuit breaker. Set it to 0 for canary deploys, 10-20% for standard rolling updates, and always combine it with serial for batched execution. When too many hosts fail in a batch, the play stops — protecting the rest of your fleet from a bad deploy.