Ansible at Scale — Patterns for 1,000+ Host Fleets

Introduction

Ansible works great for 10 hosts. At 100 hosts, you need tuning. At 1,000+, you need architecture. This guide covers the patterns, configuration, and tools that make Ansible work reliably across thousands of hosts in enterprise environments.

Performance Configuration

ansible.cfg Baseline for Large Fleets

[defaults]
# Parallelism — 50-100 for large inventories
forks = 50

# Smart gathering + Redis fact cache
gathering = smart
fact_caching = redis
fact_caching_connection = redis.internal:6379:0
fact_caching_timeout = 86400

# Reduce output overhead
callback_whitelist = timer, profile_tasks
stdout_callback = yaml
display_skipped_hosts = False
display_ok_hosts = False

# Strategy for independent tasks
strategy = free

[ssh_connection]
# Persistent SSH connections (huge speedup)
pipelining = True
ssh_args = -o ControlMaster=auto -o ControlPersist=1800s -o PreferredAuthentications=publickey
control_path_dir = /tmp/.ansible-cp

[inventory]
# Cache dynamic inventory
cache = True
cache_plugin = jsonfile
cache_connection = /tmp/ansible-inventory-cache
cache_timeout = 3600

Strategy Comparison

StrategyBehaviorBest For
linearAll hosts run task 1, then task 2 (default)Ordered deployments
freeEach host runs independentlyIndependent servers
host_pinnedLike free, but keeps host order per batchMixed workloads
mitogen_linearAccelerated linear (3-7x faster)Everything
---
- name: Independent server configuration
  hosts: webservers
  strategy: free
  tasks:
    - name: Update packages
      ansible.builtin.apt:
        upgrade: dist

Inventory Patterns

Split Inventory by Region

inventories/
├── us-east/
│   ├── hosts.yml
│   └── group_vars/
├── us-west/
│   ├── hosts.yml
│   └── group_vars/
└── eu-west/
    ├── hosts.yml
    └── group_vars/
# Run against specific region
ansible-playbook site.yml -i inventories/us-east/

# Run against all regions
ansible-playbook site.yml -i inventories/

Dynamic Inventory for Cloud

# aws_ec2.yml — auto-discover EC2 instances
plugin: amazon.aws.aws_ec2
regions:
  - us-east-1
  - us-west-2
keyed_groups:
  - key: tags.Environment
    prefix: env
  - key: instance_type
    prefix: type
  - key: placement.availability_zone
    prefix: az
filters:
  tag:ManagedBy: ansible
  instance-state-name: running
compose:
  ansible_host: private_ip_address

Serial Execution for Safety

---
- name: Rolling update across 1000 hosts
  hosts: webservers
  serial:
    - 1          # First: 1 canary host
    - 5          # Then: 5 hosts
    - "10%"      # Then: 10% at a time
  max_fail_percentage: 5

  pre_tasks:
    - name: Drain from load balancer
      ansible.builtin.uri:
        url: "http://{{ lb_host }}/api/drain/{{ inventory_hostname }}"
        method: POST

  roles:
    - deploy-app

  post_tasks:
    - name: Health check
      ansible.builtin.uri:
        url: "http://{{ inventory_hostname }}:{{ app_port }}/health"
      retries: 10
      delay: 5

    - name: Re-enable in load balancer
      ansible.builtin.uri:
        url: "http://{{ lb_host }}/api/enable/{{ inventory_hostname }}"
        method: POST

Pull Mode with ansible-pull

For very large fleets (10,000+ hosts), push mode hits limits. Pull mode flips the model — each host runs ansible-pull on a schedule:

# Deploy ansible-pull via cron on all hosts
- name: Configure ansible-pull
  hosts: all
  become: true
  tasks:
    - name: Install ansible
      ansible.builtin.package:
        name: ansible-core
        state: present

    - name: Create pull cron job
      ansible.builtin.cron:
        name: "ansible-pull"
        minute: "*/30"
        job: >
          ansible-pull
          -U https://git.example.com/infra/site.git
          -d /opt/ansible-pull
          -i localhost,
          -e "env={{ env }}"
          local.yml
          >> /var/log/ansible-pull.log 2>&1

AWX / Automation Controller

For enterprise scale, AWX (or Red Hat AAP) provides:

Architecture for 5,000+ hosts:

┌─────────────────┐     ┌──────────────┐
│  AWX Web UI      │────▶│  PostgreSQL   │
│  REST API        │     │  (HA cluster) │
└────────┬────────┘     └──────────────┘
         │
    ┌────┴────┐
    │ Receptor│ ← Automation Mesh
    │  Mesh   │
    ├─────────┤
    │ Hop     │────▶ Execution Nodes (region 1)
    │ Node    │────▶ Execution Nodes (region 2)
    │         │────▶ Execution Nodes (region 3)
    └─────────┘

Key features for scale:

  • Automation Mesh: Distribute execution across regions
  • Instance Groups: Dedicate capacity to teams
  • Job Slicing: Split one job across multiple nodes
  • Smart Inventories: Dynamic host grouping
  • Credential rotation: Centralized secret management

Execution Environments

Package Ansible + dependencies in containers for consistent execution:

# execution-environment.yml
---
version: 3
dependencies:
  galaxy:
    collections:
      - amazon.aws
      - community.vmware
      - kubernetes.core
  python:
    - boto3>=1.28
    - pyVmomi>=8.0
    - kubernetes>=28.0
  system:
    - openssh-clients
    - sshpass

build_arg_defaults:
  ANSIBLE_GALAXY_CLI_COLLECTION_OPTS: "--pre"

images:
  base_image:
    name: quay.io/ansible/ansible-runner:latest
ansible-builder build -t my-ee:latest
ansible-navigator run site.yml --eei my-ee:latest

Monitoring and Observability

# Callback plugin for Prometheus metrics
[defaults]
callback_whitelist = timer, profile_tasks, prometheus

# Export to Grafana via ARA
[defaults]
callback_whitelist = ara_default
# Profile task execution times
ANSIBLE_CALLBACKS_ENABLED=profile_tasks ansible-playbook site.yml

Benchmark: Tuning Impact

Configuration100 hosts500 hosts1000 hosts
Default (forks=5)8m 30s42m85m
forks=50 + pipelining2m 10s10m20m
+ fact caching1m 05s5m10m
+ free strategy0m 48s3m 30s7m
+ Mitogen0m 20s1m 30s3m

Conclusion

Scaling Ansible requires a combination of configuration tuning (forks, pipelining, fact caching), architectural patterns (serial batches, pull mode, automation mesh), and tooling (AWX, execution environments). Start with the ansible.cfg baseline above and layer in complexity as your fleet grows.