Introduction

The serial keyword controls how many hosts Ansible updates at once. Using percentages (serial: "25%") or stepped batches (serial: [1, 5, "50%"]) enables rolling deployments, canary releases, and risk-controlled updates. Without serial, Ansible runs on all hosts simultaneously — fine for configuration, risky for deployments.

Basic Percentage

---
- name: Rolling update — 25% at a time
  hosts: webservers
  serial: "25%"

  tasks:
    - name: Deploy application
      ansible.builtin.copy:
        src: app.tar.gz
        dest: /opt/app/app.tar.gz

    - name: Restart service
      ansible.builtin.systemd:
        name: myapp
        state: restarted

With 20 hosts and serial: "25%", Ansible runs 5 hosts per batch (4 batches total).

Percentage Rounding

HostsserialBatch SizeBatches
10"25%"3 (rounds up)4
10"30%"34
10"50%"52
100"10%"1010
7"25%"2 (rounds up)4
1"50%"1 (minimum 1)1

Percentages always round up to at least 1.

Fixed Batch Size

# Run 5 hosts at a time
- name: Batch update
  hosts: all
  serial: 5

  tasks:
    - name: Update packages
      ansible.builtin.apt:
        upgrade: dist

Stepped Batches (Canary Pattern)

# Start small, grow confidence, finish fast
- name: Canary deployment
  hosts: webservers
  serial:
    - 1        # First: 1 host (canary)
    - 5        # Then: 5 hosts
    - "25%"    # Then: 25% of remaining
    - "100%"   # Finally: all remaining

  tasks:
    - name: Deploy new version
      ansible.builtin.copy:
        src: app-v2.tar.gz
        dest: /opt/app/

    - name: Restart
      ansible.builtin.systemd:
        name: myapp
        state: restarted

    - name: Health check
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}:8080/health"
        status_code: 200
      register: health
      until: health.status == 200
      retries: 10
      delay: 5

With 100 hosts:

  1. Batch 1: 1 host (canary)
  2. Batch 2: 5 hosts
  3. Batch 3: 24 hosts (25% of 94 remaining)
  4. Batch 4: 70 hosts (all remaining)

Fail Fast with max_fail_percentage

- name: Safe rolling update
  hosts: webservers
  serial: "20%"
  max_fail_percentage: 10

  tasks:
    - name: Deploy
      ansible.builtin.copy:
        src: app.tar.gz
        dest: /opt/app/

    - name: Health check
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
        status_code: 200
      retries: 5
      delay: 3

If >10% of a batch fails, the entire play aborts — preventing a bad deploy from spreading.

Zero-Downtime with Load Balancer

- name: Zero-downtime deploy
  hosts: webservers
  serial: "25%"
  max_fail_percentage: 0  # Any failure stops everything

  pre_tasks:
    - name: Remove from LB
      ansible.builtin.uri:
        url: "https://lb.example.com/api/remove/{{ inventory_hostname }}"
        method: POST
      delegate_to: localhost

    - name: Drain connections
      ansible.builtin.pause:
        seconds: 30

  roles:
    - deploy_app

  post_tasks:
    - name: Health check
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
      register: check
      until: check.status == 200
      retries: 12
      delay: 5

    - name: Add back to LB
      ansible.builtin.uri:
        url: "https://lb.example.com/api/add/{{ inventory_hostname }}"
        method: POST
      delegate_to: localhost

order Directive

# Control host processing order within each batch
- name: Ordered rolling update
  hosts: webservers
  serial: 2
  order: sorted      # alphabetical order

  tasks:
    - name: Update
      ansible.builtin.debug:
        msg: "Updating {{ inventory_hostname }}"
OrderDescription
inventoryDefault — order defined in inventory
sortedAlphabetical by hostname
reverse_sortedReverse alphabetical
shuffleRandom order
reverse_inventoryReverse of inventory order

Comparison of Strategies

StrategyserialUse CaseRisk
All at once(not set)Config managementHigh
Fixed batch5Known cluster sizeMedium
Percentage"25%"Scales with fleetMedium
Canary[1, 5, "50%"]Production deploysLow
One-at-a-time1Database migrationsLowest

Troubleshooting

IssueSolution
Too slow with serial: 1Increase batch or use percentage
All hosts fail at onceAdd serial — you're running without it
Batch too small with percentagePercentage rounds up; use fixed number instead
Remaining hosts don't run after failureExpected — max_fail_percentage or any failure stops subsequent batches
Need different serial per playUse multiple plays with different serial values

Best Practices

  1. Always use serial for deployments — never deploy to all hosts at once
  2. Canary pattern for production — [1, 5, "25%", "100%"]
  3. Combine with max_fail_percentage — stop early on failures
  4. Health checks after restart — uri + until loop
  5. LB integration in pre/post_tasks — drain before, add after
  6. Use order: shuffle — avoid always hitting the same hosts first

Conclusion

serial with percentages and stepped batches is how you do safe production deployments. Start with a canary (1 host), grow the batch size as confidence builds, and combine with max_fail_percentage to stop the blast radius. The pattern serial: [1, 5, "25%", "100%"] with health checks is the gold standard for rolling updates.