Ansible failed_when — Custom Failure Conditions

Introduction

Ansible decides task success or failure based on module return codes. But real-world commands don't always follow conventions — grep returns 1 for "not found" (not an error), APIs return 200 with error payloads, and scripts use custom exit codes. failed_when lets you define exactly what constitutes a failure.

Basic Usage

---
- name: Custom failure conditions
  hosts: all
  tasks:
    # Default: rc != 0 is failure
    # With failed_when: YOU decide what's a failure

    - name: Search for config entry
      ansible.builtin.command:
        cmd: grep "max_connections" /etc/postgresql/16/main/postgresql.conf
      register: grep_result
      failed_when: grep_result.rc > 1
      # rc=0 → found (ok)
      # rc=1 → not found (ok, not a failure)
      # rc=2 → file error (real failure)

Never Fail

    - name: Check service status (informational only)
      ansible.builtin.command:
        cmd: systemctl is-active myapp
      register: service_check
      failed_when: false  # Never fails, regardless of return code

    - name: Report status
      ansible.builtin.debug:
        msg: "Service is {{ 'running' if service_check.rc == 0 else 'stopped' }}"

Fail on Output Content

    # Fail based on command output, not return code
    - name: Run database health check
      ansible.builtin.command:
        cmd: /opt/db/healthcheck.sh
      register: health
      failed_when: "'CRITICAL' in health.stdout"

    # Fail if output contains error indicators
    - name: Deploy application
      ansible.builtin.shell:
        cmd: /opt/deploy.sh 2>&1
      register: deploy
      failed_when:
        - "'ERROR' in deploy.stdout"
        - "'ROLLBACK' not in deploy.stdout"
      # Fails if ERROR in output AND no ROLLBACK happened

    # Fail if expected output is missing
    - name: Verify installation
      ansible.builtin.command:
        cmd: /opt/app/bin/myapp --version
      register: version
      failed_when: "'v2.' not in version.stdout"

Multiple Conditions

    # ALL conditions must be true to fail (AND logic)
    - name: Check API response
      ansible.builtin.uri:
        url: "https://api.example.com/health"
        return_content: true
      register: api_health
      failed_when:
        - api_health.status != 200
        - api_health.json.status != "healthy"
      # Fails only if BOTH status is not 200 AND json status is not healthy

    # ANY condition triggers failure (OR logic)
    - name: Validate configuration
      ansible.builtin.command:
        cmd: /opt/app/validate-config
      register: validate
      failed_when: >
        validate.rc != 0 or
        'WARN' in validate.stderr or
        'deprecated' in validate.stdout

Complex Expressions

    # Numeric thresholds
    - name: Check disk usage
      ansible.builtin.command:
        cmd: df --output=pcent / | tail -1 | tr -d ' %'
      register: disk_pct
      failed_when: disk_pct.stdout | int > 90

    # List/count checks
    - name: Count running processes
      ansible.builtin.shell:
        cmd: pgrep -c myapp
      register: proc_count
      failed_when: proc_count.stdout | int < 2

    # Regex matching
    - name: Check SSL certificate expiry
      ansible.builtin.command:
        cmd: openssl x509 -enddate -noout -in /etc/ssl/certs/app.pem
      register: cert_info
      failed_when: cert_info.stdout is search('notAfter=.*202[0-5]')

With Loops

    - name: Verify all services are healthy
      ansible.builtin.uri:
        url: "http://{{ item.host }}:{{ item.port }}/health"
        return_content: true
      register: health_checks
      failed_when: >
        health_result.status != 200 or
        health_result.json.status != 'ok'
      loop:
        - { host: web01, port: 8080 }
        - { host: web02, port: 8080 }
        - { host: api01, port: 9090 }
      loop_control:
        loop_var: item
        label: "{{ item.host }}"

failed_when + changed_when Together

    # Full task status control
    - name: Apply database migration
      ansible.builtin.command:
        cmd: /opt/app/migrate.sh
      register: migration
      changed_when: "'Migrated' in migration.stdout"
      failed_when: "'ERROR' in migration.stderr"
      # "No migrations" → ok (not changed, not failed)
      # "Migrated 3 tables" → changed
      # "ERROR: duplicate key" → failed

Practical Examples

API Error Handling

    - name: Create user via API
      ansible.builtin.uri:
        url: "https://api.example.com/users"
        method: POST
        body_format: json
        body:
          username: "{{ new_user }}"
          email: "{{ new_email }}"
        status_code: [200, 201, 409]  # Accept "already exists"
      register: api_response
      failed_when:
        - api_response.status == 409
        - api_response.json.error != "user_already_exists"
      # 409 with "user_already_exists" is OK (idempotent)
      # 409 with any other error is a real failure

Package Version Check

    - name: Verify minimum package version
      ansible.builtin.command:
        cmd: rpm -q --queryformat '%{VERSION}' nginx
      register: pkg_version
      failed_when: pkg_version.stdout is version('1.24', '<')
      changed_when: false

Cluster Quorum Check

    - name: Check cluster has quorum
      ansible.builtin.command:
        cmd: crm status
      register: cluster
      failed_when: >
        'partition with quorum' not in cluster.stdout or
        'OFFLINE' in cluster.stdout
      changed_when: false

Troubleshooting

IssueSolution
failed_when expression errorWrap complex expressions in quotes or use > block scalar
Task succeeds when it shouldn'tCheck condition logic — and vs or in multi-conditions
Variable undefined in expressionEnsure register is on the same task
Can't compare versionsUse is version('1.0', '>=') Jinja2 test
Integer comparison failsCast with `

Best Practices

  1. Use failed_when over ignore_errors — precise failure control is better than ignoring all errors
  2. Always register the result — you need the output to make failure decisions
  3. Document your logic — add comments explaining why custom failure conditions exist
  4. Test edge cases — what happens when the command returns empty output?
  5. Combine with changed_when — give Ansible accurate information about both change and failure
  6. Use failed_when: false sparingly — only for truly informational tasks

Conclusion

failed_when transforms Ansible from "did the command return 0?" to "did the operation actually succeed?" It's essential for shell commands, API calls, and any task where the default success/failure detection doesn't match reality. Combined with changed_when, you get playbooks that accurately report what happened and what changed.