Ansible Serial Keyword — Rolling Update Batch Size Strategy

Introduction

By default, Ansible runs tasks on all hosts simultaneously. The serial keyword changes this — it processes hosts in batches, enabling rolling updates, canary deployments, and zero-downtime upgrades. This is essential for production deployments where you can't take all servers offline at once.

Basic Usage

---
- name: Rolling update - 2 hosts at a time
  hosts: webservers
  serial: 2
  become: true
  tasks:
    - name: Deploy application
      ansible.builtin.git:
        repo: https://github.com/org/app.git
        dest: /opt/app
        version: "{{ app_version }}"

    - name: Restart service
      ansible.builtin.systemd:
        name: myapp
        state: restarted

    - name: Wait for service to be healthy
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}:8080/health"
        status_code: 200
      retries: 10
      delay: 5
      register: health
      until: health.status == 200

With 10 webservers: runs on hosts 1-2, waits for completion, then 3-4, then 5-6, etc.

Percentage-Based Batches

---
- name: Update 25% of fleet at a time
  hosts: webservers
  serial: "25%"
  tasks:
    - name: Upgrade packages
      ansible.builtin.apt:
        upgrade: safe
      register: upgrade_result

    - name: Reboot if needed
      ansible.builtin.reboot:
        reboot_timeout: 300
      when: upgrade_result.changed

Canary Deployment Pattern

Start with 1 host, then ramp up:

---
- name: Canary deployment
  hosts: webservers
  serial:
    - 1        # First: deploy to 1 host (canary)
    - 5        # Then: 5 hosts at a time
    - "100%"   # Finally: all remaining
  become: true

  tasks:
    - name: Deploy new version
      ansible.builtin.git:
        repo: https://github.com/org/app.git
        dest: /opt/app
        version: "{{ app_version }}"

    - name: Restart application
      ansible.builtin.systemd:
        name: myapp
        state: restarted

    - name: Verify deployment
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}:8080/health"
        status_code: 200
      retries: 12
      delay: 5
      until: result.status == 200
      register: result

    - name: Smoke test
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}:8080/api/v1/status"
        return_content: true
      register: smoke
      failed_when: "'version' not in smoke.content"

Fail-Safe with max_fail_percentage

Stop the entire deployment if too many hosts fail:

---
- name: Safe rolling deploy
  hosts: webservers
  serial: 3
  max_fail_percentage: 30
  become: true

  tasks:
    - name: Deploy and verify
      block:
        - name: Pull new code
          ansible.builtin.git:
            repo: https://github.com/org/app.git
            dest: /opt/app
            version: "{{ app_version }}"

        - name: Restart
          ansible.builtin.systemd:
            name: myapp
            state: restarted

        - name: Health check
          ansible.builtin.uri:
            url: "http://localhost:8080/health"
          retries: 5
          delay: 3
          register: health
          until: health.status == 200

      rescue:
        - name: Rollback this host
          ansible.builtin.git:
            repo: https://github.com/org/app.git
            dest: /opt/app
            version: "{{ previous_version }}"

        - name: Restart with old version
          ansible.builtin.systemd:
            name: myapp
            state: restarted

        - name: Mark host as failed
          ansible.builtin.fail:
            msg: "Deploy failed on {{ inventory_hostname }}"

Load Balancer Integration

Remove hosts from load balancer during updates:

---
- name: Rolling update with LB drain
  hosts: webservers
  serial: 2
  become: true

  pre_tasks:
    - name: Disable host in HAProxy
      community.general.haproxy:
        state: disabled
        host: "{{ inventory_hostname }}"
        backend: app-backend
      delegate_to: "{{ groups['loadbalancers'][0] }}"

    - name: Wait for connections to drain
      ansible.builtin.wait_for:
        timeout: 30

  tasks:
    - name: Deploy application
      ansible.builtin.git:
        repo: https://github.com/org/app.git
        dest: /opt/app
        version: "{{ app_version }}"

    - name: Restart application
      ansible.builtin.systemd:
        name: myapp
        state: restarted

    - name: Wait for service to be ready
      ansible.builtin.uri:
        url: "http://localhost:8080/health"
      retries: 10
      delay: 3
      register: health
      until: health.status == 200

  post_tasks:
    - name: Re-enable host in HAProxy
      community.general.haproxy:
        state: enabled
        host: "{{ inventory_hostname }}"
        backend: app-backend
      delegate_to: "{{ groups['loadbalancers'][0] }}"

    - name: Wait for LB health check
      ansible.builtin.pause:
        seconds: 10

Serial vs Forks vs Throttle

KeywordScopePurpose
serialPlayProcess N hosts before moving to next batch
forksGlobalMax parallel SSH connections (default: 5)
throttleTaskMax parallel executions of a single task
# These work together:
---
- name: Combined example
  hosts: all          # 100 hosts
  serial: 10          # Process 10 at a time
  # forks = 5 in ansible.cfg → 5 of the 10 run in parallel

  tasks:
    - name: API call (rate limited)
      ansible.builtin.uri:
        url: "https://api.example.com/register"
        method: POST
      throttle: 2       # Only 2 API calls at once (within the batch of 10)

any_errors_fatal with Serial

---
- name: Critical deployment - stop on any failure
  hosts: webservers
  serial: 2
  any_errors_fatal: true  # If any host in the batch fails, stop everything

  tasks:
    - name: Deploy critical update
      ansible.builtin.command:
        cmd: /opt/deploy.sh

Run Once per Batch

    - name: Notify monitoring (once per batch, not per host)
      ansible.builtin.uri:
        url: "https://monitoring.example.com/api/maintenance"
        method: POST
        body_format: json
        body:
          message: "Deploying batch {{ ansible_play_batch | join(',') }}"
      run_once: true
      delegate_to: localhost

Troubleshooting

IssueSolution
Too slowIncrease serial value or forks
Failures stop everythingUse max_fail_percentage instead of any_errors_fatal
Need rollbackAdd block/rescue pattern within serial play
Hosts out of LB too longReduce drain wait time, optimize deploy tasks
Wrong batch orderingUse order: sorted or order: shuffle on the play

Best Practices

  1. Always health-check after deploy — uri module with retries
  2. Canary first — serial: [1, 5, "100%"] catches issues early
  3. Integrate with load balancer — drain before update, re-enable after
  4. Set max_fail_percentage — don't deploy to all hosts if some fail
  5. Use order: shuffle — randomize batch order to spread risk
  6. Combine with throttle — rate-limit API calls within batches

Conclusion

The serial keyword is your primary tool for zero-downtime deployments. Combined with canary patterns, load balancer integration, and max_fail_percentage, you can safely deploy to thousands of hosts with automatic rollback on failure. Start with conservative batches and increase as confidence grows.