Ansible max_fail_percentage — Control Failure Tolerance

Introduction

max_fail_percentage sets the maximum percentage of hosts that can fail before Ansible aborts the entire play. Combined with serial, it creates a safety valve for rolling deployments: if too many hosts fail in a batch, the play stops before the bad deploy reaches your entire fleet.

Basic Usage

---
- name: Safe rolling deploy
  hosts: webservers
  serial: "25%"
  max_fail_percentage: 10

  tasks:
    - name: Deploy application
      ansible.builtin.copy:
        src: app.tar.gz
        dest: /opt/app/

    - name: Restart service
      ansible.builtin.systemd:
        name: myapp
        state: restarted

    - name: Health check
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
        status_code: 200
      retries: 5
      delay: 3

With 100 hosts, serial: "25%", and max_fail_percentage: 10:

  • Each batch = 25 hosts
  • If >2 hosts (10% of 25) fail in a batch → entire play aborts
  • Remaining batches won't run

How It Works

Batch 1 (25 hosts): 24 pass, 1 fails (4%) → Continue ✅
Batch 2 (25 hosts): 22 pass, 3 fails (12%) → ABORT ❌
Batch 3 (25 hosts): Never runs
Batch 4 (25 hosts): Never runs

The percentage is calculated per batch, not across the entire play.

Common Settings

SettingEffectUse Case
max_fail_percentage: 0Any failure abortsCritical infrastructure
max_fail_percentage: 1010% toleranceStandard deployments
max_fail_percentage: 2525% toleranceFlaky environments
max_fail_percentage: 49Nearly half can failNon-critical updates
max_fail_percentage: 100Never aborts (default)Fire-and-forget

Zero Tolerance

- name: Database migration (zero tolerance)
  hosts: database
  serial: 1
  max_fail_percentage: 0

  tasks:
    - name: Backup database
      ansible.builtin.command:
        cmd: pg_dump myapp > /backups/pre-migrate.sql

    - name: Run migration
      ansible.builtin.command:
        cmd: /opt/myapp/bin/migrate

    - name: Verify schema
      ansible.builtin.command:
        cmd: /opt/myapp/bin/migrate --check
      changed_when: false

Canary with Failure Gates

- name: Canary deploy with failure gates
  hosts: webservers
  serial: [1, 5, "25%", "100%"]
  max_fail_percentage: 0  # Zero tolerance on canary

  pre_tasks:
    - name: Remove from load balancer
      ansible.builtin.uri:
        url: "https://lb.example.com/api/remove/{{ inventory_hostname }}"
        method: POST
      delegate_to: localhost

  roles:
    - deploy_app

  post_tasks:
    - name: Health check
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
      register: health
      until: health.status == 200
      retries: 12
      delay: 5

    - name: Add back to load balancer
      ansible.builtin.uri:
        url: "https://lb.example.com/api/add/{{ inventory_hostname }}"
        method: POST
      delegate_to: localhost

If the canary (first host) fails, the entire deploy stops.

any_errors_fatal vs max_fail_percentage

# any_errors_fatal: ALL hosts stop if ANY host fails
- name: All-or-nothing
  hosts: webservers
  any_errors_fatal: true
  tasks:
    - name: Critical task
      ansible.builtin.command:
        cmd: /opt/app/check

# max_fail_percentage: threshold-based
- name: Threshold-based
  hosts: webservers
  serial: 10
  max_fail_percentage: 20
  tasks:
    - name: Tolerant task
      ansible.builtin.command:
        cmd: /opt/app/update
Featureany_errors_fatalmax_fail_percentage: 0
ScopeAll hosts immediatelyPer batch with serial
ToleranceZero (any failure)Zero (but per batch)
Without serialStops all hostsSame as default
Best forPre-flight checksRolling deployments

Calculation Examples

20 hosts, serial: 5, max_fail_percentage: 30
→ Batch size: 5
→ Max failures per batch: 1 (30% of 5 = 1.5, rounds down to 1)
→ If 2+ hosts fail in a batch → abort

100 hosts, serial: "10%", max_fail_percentage: 20
→ Batch size: 10
→ Max failures per batch: 2 (20% of 10)
→ If 3+ hosts fail in a batch → abort

50 hosts, serial: 1, max_fail_percentage: 0
→ Batch size: 1
→ Any failure → abort immediately

Troubleshooting

IssueSolution
Play doesn't stop on failuresCheck max_fail_percentage is set (default is 100)
Too aggressive — stops on 1 failureIncrease percentage or batch size
Failures in one batch, success in nextThat's expected — each batch is independent
Need to continue after abortFix the issue, re-run; Ansible skips successful hosts with --limit @retry_file
max_fail_percentage without serialWorks but applies to entire host list as one "batch"

Best Practices

  1. Always pair with serial — meaningless without batching
  2. Use 0% for canary deploys — any failure should stop propagation
  3. Use 10-20% for standard rolling updates — allows for flaky hosts
  4. Add health checks — uri + until to catch delayed failures
  5. Combine with --limit @playbook.retry — re-run only failed hosts
  6. Log failures — use block/rescue to capture what went wrong before aborting

Conclusion

max_fail_percentage is your deployment circuit breaker. Set it to 0 for canary deploys, 10-20% for standard rolling updates, and always combine it with serial for batched execution. When too many hosts fail in a batch, the play stops — protecting the rest of your fleet from a bad deploy.