Ansible any_errors_fatal and max_fail_percentage — Playbook Failure Control

Introduction

By default, when a task fails on one host, Ansible removes that host from the play but continues on other hosts. This is fine for independent servers, but dangerous for clustered services — if a database migration fails on one node, you don't want it to succeed on others. any_errors_fatal and max_fail_percentage give you control over when failures should stop everything.

any_errors_fatal

One host fails → all hosts stop immediately.

---
- name: Critical database migration
  hosts: database_cluster
  any_errors_fatal: true  # ONE failure stops ALL hosts

  tasks:
    - name: Run schema migration
      ansible.builtin.command:
        cmd: /opt/db/migrate.sh
      # If this fails on db01, db02 and db03 STOP immediately

    - name: Verify data integrity
      ansible.builtin.command:
        cmd: /opt/db/verify.sh

Per-Task any_errors_fatal

- name: Mixed criticality
  hosts: all
  tasks:
    - name: Non-critical cleanup (can fail on some hosts)
      ansible.builtin.file:
        path: /tmp/old-cache
        state: absent
      ignore_errors: true

    - name: Critical deployment (must succeed everywhere)
      ansible.builtin.command:
        cmd: /opt/deploy.sh
      any_errors_fatal: true
      # If ANY host fails here, ALL hosts stop

    - name: Post-deploy verification
      ansible.builtin.uri:
        url: "http://localhost:8080/health"

With serial (Rolling Updates)

- name: Rolling deployment
  hosts: webservers    # 20 hosts
  serial: 5            # Process 5 at a time
  any_errors_fatal: true

  tasks:
    - name: Deploy new version
      ansible.builtin.command:
        cmd: /opt/deploy.sh v2.0

    - name: Verify health
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
        status_code: 200

# Batch 1: [web01-05] — if web03 fails, web04/05 stop AND batch 2 never starts
# Without any_errors_fatal: web03 removed, web04/05 continue, batches 2-4 proceed

max_fail_percentage

Allow some failures before stopping:

- name: Tolerate some failures
  hosts: webservers    # 20 hosts
  serial: 10
  max_fail_percentage: 20  # Stop if >20% of batch fails

  tasks:
    - name: Deploy application
      ansible.builtin.command:
        cmd: /opt/deploy.sh

# Batch of 10:
#   1 failure (10%) → continue
#   2 failures (20%) → continue (at threshold)
#   3 failures (30%) → STOP ALL (exceeded 20%)

Setting to 0

  max_fail_percentage: 0
  # Equivalent to any_errors_fatal: true
  # ANY failure stops everything

Setting to 100

  max_fail_percentage: 100
  # Never stops due to failures (all hosts can fail)
  # Effectively disables failure-based stopping

Comparison

FeatureDefaultany_errors_fatal: truemax_fail_percentage: 20
1 of 10 failsContinue 9Stop allContinue
3 of 10 failContinue 7Stop allStop all
10 of 10 failAll failStop allStop all
Use caseIndependent serversCritical operationsRolling updates
ScopePer-play or taskPer-play or taskPer-play only

Practical Examples

Database Cluster

- name: Database cluster upgrade
  hosts: db_cluster
  serial: 1           # One node at a time
  any_errors_fatal: true

  pre_tasks:
    - name: Check cluster health before starting
      ansible.builtin.command:
        cmd: pg_isready
      changed_when: false

  tasks:
    - name: Stop PostgreSQL
      ansible.builtin.systemd:
        name: postgresql
        state: stopped

    - name: Upgrade packages
      ansible.builtin.package:
        name: postgresql-16
        state: latest

    - name: Start PostgreSQL
      ansible.builtin.systemd:
        name: postgresql
        state: started

    - name: Verify replication
      ansible.builtin.command:
        cmd: psql -c "SELECT state FROM pg_stat_replication;"
      register: repl
      until: "'streaming' in repl.stdout"
      retries: 30
      delay: 5
      changed_when: false

Load-Balanced Web Tier

- name: Zero-downtime deployment
  hosts: webservers
  serial: "30%"         # 30% of hosts per batch
  max_fail_percentage: 10

  pre_tasks:
    - name: Remove from load balancer
      ansible.builtin.uri:
        url: "https://lb.example.com/api/remove"
        method: POST
        body: '{"host": "{{ inventory_hostname }}"}'

  tasks:
    - name: Deploy
      ansible.builtin.command:
        cmd: /opt/deploy.sh

    - name: Verify health
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
        status_code: 200
      register: health
      until: health.status == 200
      retries: 12
      delay: 5

  post_tasks:
    - name: Add back to load balancer
      ansible.builtin.uri:
        url: "https://lb.example.com/api/add"
        method: POST
        body: '{"host": "{{ inventory_hostname }}"}'

Mixed Criticality

- name: Server configuration
  hosts: all
  tasks:
    # Phase 1: Non-critical (partial failure OK)
    - name: Install monitoring agent
      ansible.builtin.package:
        name: monitoring-agent
        state: present
      ignore_errors: true

    # Phase 2: Critical (all or nothing)
    - name: Apply security patches
      ansible.builtin.package:
        name: "*"
        state: latest
        security: true
      any_errors_fatal: true

    # Phase 3: Verification
    - name: Verify system health
      ansible.builtin.command:
        cmd: /opt/healthcheck.sh
      any_errors_fatal: true

Combining with block/rescue

- name: Deployment with recovery
  hosts: webservers
  serial: 5
  any_errors_fatal: true

  tasks:
    - block:
        - name: Deploy new version
          ansible.builtin.command:
            cmd: /opt/deploy.sh v2.0

        - name: Verify deployment
          ansible.builtin.uri:
            url: "http://localhost:8080/health"
      rescue:
        - name: Rollback
          ansible.builtin.command:
            cmd: /opt/deploy.sh rollback

        - name: Verify rollback
          ansible.builtin.uri:
            url: "http://localhost:8080/health"

        - name: Fail after rollback to stop remaining batches
          ansible.builtin.fail:
            msg: "Deployment failed on {{ inventory_hostname }}, rolled back, stopping remaining batches"

Troubleshooting

IssueSolution
One bad host stops everythingRemove any_errors_fatal or use max_fail_percentage
Failures don't stop playbookAdd any_errors_fatal: true or set max_fail_percentage: 0
Want to continue after some failuresUse max_fail_percentage with appropriate threshold
Need per-task failure controlSet any_errors_fatal on specific tasks, not the play
max_fail_percentage math wrongIt's percentage of current batch (with serial), not total hosts

Best Practices

  1. Use any_errors_fatal for database/cluster operations — consistency is critical
  2. Use max_fail_percentage for web tiers — some failures are tolerable
  3. Combine with serial — process hosts in batches for safer rollouts
  4. Add block/rescue for rollback — automatic recovery before stopping
  5. Test with --check — verify which hosts would be affected
  6. Log failure details — use rescue block to capture diagnostics

Conclusion

any_errors_fatal and max_fail_percentage protect your infrastructure from cascading failures. Use any_errors_fatal when consistency matters (databases, clusters) — one failure means stop everything. Use max_fail_percentage when partial availability is acceptable (web servers behind a load balancer). Combined with serial for rolling updates, these controls give you safe, production-grade deployment automation.