Ansible any_errors_fatal and max_fail_percentage — Playbook Failure Control
Introduction
By default, when a task fails on one host, Ansible removes that host from the play but continues on other hosts. This is fine for independent servers, but dangerous for clustered services — if a database migration fails on one node, you don't want it to succeed on others. any_errors_fatal and max_fail_percentage give you control over when failures should stop everything.
any_errors_fatal
One host fails → all hosts stop immediately.
---
- name: Critical database migration
hosts: database_cluster
any_errors_fatal: true # ONE failure stops ALL hosts
tasks:
- name: Run schema migration
ansible.builtin.command:
cmd: /opt/db/migrate.sh
# If this fails on db01, db02 and db03 STOP immediately
- name: Verify data integrity
ansible.builtin.command:
cmd: /opt/db/verify.sh
Per-Task any_errors_fatal
- name: Mixed criticality
hosts: all
tasks:
- name: Non-critical cleanup (can fail on some hosts)
ansible.builtin.file:
path: /tmp/old-cache
state: absent
ignore_errors: true
- name: Critical deployment (must succeed everywhere)
ansible.builtin.command:
cmd: /opt/deploy.sh
any_errors_fatal: true
# If ANY host fails here, ALL hosts stop
- name: Post-deploy verification
ansible.builtin.uri:
url: "http://localhost:8080/health"
With serial (Rolling Updates)
- name: Rolling deployment
hosts: webservers # 20 hosts
serial: 5 # Process 5 at a time
any_errors_fatal: true
tasks:
- name: Deploy new version
ansible.builtin.command:
cmd: /opt/deploy.sh v2.0
- name: Verify health
ansible.builtin.uri:
url: "http://localhost:8080/health"
status_code: 200
# Batch 1: [web01-05] — if web03 fails, web04/05 stop AND batch 2 never starts
# Without any_errors_fatal: web03 removed, web04/05 continue, batches 2-4 proceed
max_fail_percentage
Allow some failures before stopping:
- name: Tolerate some failures
hosts: webservers # 20 hosts
serial: 10
max_fail_percentage: 20 # Stop if >20% of batch fails
tasks:
- name: Deploy application
ansible.builtin.command:
cmd: /opt/deploy.sh
# Batch of 10:
# 1 failure (10%) → continue
# 2 failures (20%) → continue (at threshold)
# 3 failures (30%) → STOP ALL (exceeded 20%)
Setting to 0
max_fail_percentage: 0
# Equivalent to any_errors_fatal: true
# ANY failure stops everything
Setting to 100
max_fail_percentage: 100
# Never stops due to failures (all hosts can fail)
# Effectively disables failure-based stopping
Comparison
| Feature | Default | any_errors_fatal: true | max_fail_percentage: 20 |
|---|---|---|---|
| 1 of 10 fails | Continue 9 | Stop all | Continue |
| 3 of 10 fail | Continue 7 | Stop all | Stop all |
| 10 of 10 fail | All fail | Stop all | Stop all |
| Use case | Independent servers | Critical operations | Rolling updates |
| Scope | Per-play or task | Per-play or task | Per-play only |
Practical Examples
Database Cluster
- name: Database cluster upgrade
hosts: db_cluster
serial: 1 # One node at a time
any_errors_fatal: true
pre_tasks:
- name: Check cluster health before starting
ansible.builtin.command:
cmd: pg_isready
changed_when: false
tasks:
- name: Stop PostgreSQL
ansible.builtin.systemd:
name: postgresql
state: stopped
- name: Upgrade packages
ansible.builtin.package:
name: postgresql-16
state: latest
- name: Start PostgreSQL
ansible.builtin.systemd:
name: postgresql
state: started
- name: Verify replication
ansible.builtin.command:
cmd: psql -c "SELECT state FROM pg_stat_replication;"
register: repl
until: "'streaming' in repl.stdout"
retries: 30
delay: 5
changed_when: false
Load-Balanced Web Tier
- name: Zero-downtime deployment
hosts: webservers
serial: "30%" # 30% of hosts per batch
max_fail_percentage: 10
pre_tasks:
- name: Remove from load balancer
ansible.builtin.uri:
url: "https://lb.example.com/api/remove"
method: POST
body: '{"host": "{{ inventory_hostname }}"}'
tasks:
- name: Deploy
ansible.builtin.command:
cmd: /opt/deploy.sh
- name: Verify health
ansible.builtin.uri:
url: "http://localhost:8080/health"
status_code: 200
register: health
until: health.status == 200
retries: 12
delay: 5
post_tasks:
- name: Add back to load balancer
ansible.builtin.uri:
url: "https://lb.example.com/api/add"
method: POST
body: '{"host": "{{ inventory_hostname }}"}'
Mixed Criticality
- name: Server configuration
hosts: all
tasks:
# Phase 1: Non-critical (partial failure OK)
- name: Install monitoring agent
ansible.builtin.package:
name: monitoring-agent
state: present
ignore_errors: true
# Phase 2: Critical (all or nothing)
- name: Apply security patches
ansible.builtin.package:
name: "*"
state: latest
security: true
any_errors_fatal: true
# Phase 3: Verification
- name: Verify system health
ansible.builtin.command:
cmd: /opt/healthcheck.sh
any_errors_fatal: true
Combining with block/rescue
- name: Deployment with recovery
hosts: webservers
serial: 5
any_errors_fatal: true
tasks:
- block:
- name: Deploy new version
ansible.builtin.command:
cmd: /opt/deploy.sh v2.0
- name: Verify deployment
ansible.builtin.uri:
url: "http://localhost:8080/health"
rescue:
- name: Rollback
ansible.builtin.command:
cmd: /opt/deploy.sh rollback
- name: Verify rollback
ansible.builtin.uri:
url: "http://localhost:8080/health"
- name: Fail after rollback to stop remaining batches
ansible.builtin.fail:
msg: "Deployment failed on {{ inventory_hostname }}, rolled back, stopping remaining batches"
Troubleshooting
| Issue | Solution |
|---|---|
| One bad host stops everything | Remove any_errors_fatal or use max_fail_percentage |
| Failures don't stop playbook | Add any_errors_fatal: true or set max_fail_percentage: 0 |
| Want to continue after some failures | Use max_fail_percentage with appropriate threshold |
| Need per-task failure control | Set any_errors_fatal on specific tasks, not the play |
max_fail_percentage math wrong | It's percentage of current batch (with serial), not total hosts |
Best Practices
- Use
any_errors_fatalfor database/cluster operations — consistency is critical - Use
max_fail_percentagefor web tiers — some failures are tolerable - Combine with
serial— process hosts in batches for safer rollouts - Add
block/rescuefor rollback — automatic recovery before stopping - Test with
--check— verify which hosts would be affected - Log failure details — use rescue block to capture diagnostics
Conclusion
any_errors_fatal and max_fail_percentage protect your infrastructure from cascading failures. Use any_errors_fatal when consistency matters (databases, clusters) — one failure means stop everything. Use max_fail_percentage when partial availability is acceptable (web servers behind a load balancer). Combined with serial for rolling updates, these controls give you safe, production-grade deployment automation.