Ansible at Scale — Patterns for 1,000+ Host Fleets
Introduction
Ansible works great for 10 hosts. At 100 hosts, you need tuning. At 1,000+, you need architecture. This guide covers the patterns, configuration, and tools that make Ansible work reliably across thousands of hosts in enterprise environments.
Performance Configuration
ansible.cfg Baseline for Large Fleets
[defaults]
# Parallelism — 50-100 for large inventories
forks = 50
# Smart gathering + Redis fact cache
gathering = smart
fact_caching = redis
fact_caching_connection = redis.internal:6379:0
fact_caching_timeout = 86400
# Reduce output overhead
callback_whitelist = timer, profile_tasks
stdout_callback = yaml
display_skipped_hosts = False
display_ok_hosts = False
# Strategy for independent tasks
strategy = free
[ssh_connection]
# Persistent SSH connections (huge speedup)
pipelining = True
ssh_args = -o ControlMaster=auto -o ControlPersist=1800s -o PreferredAuthentications=publickey
control_path_dir = /tmp/.ansible-cp
[inventory]
# Cache dynamic inventory
cache = True
cache_plugin = jsonfile
cache_connection = /tmp/ansible-inventory-cache
cache_timeout = 3600
Strategy Comparison
| Strategy | Behavior | Best For |
|---|---|---|
linear | All hosts run task 1, then task 2 (default) | Ordered deployments |
free | Each host runs independently | Independent servers |
host_pinned | Like free, but keeps host order per batch | Mixed workloads |
mitogen_linear | Accelerated linear (3-7x faster) | Everything |
---
- name: Independent server configuration
hosts: webservers
strategy: free
tasks:
- name: Update packages
ansible.builtin.apt:
upgrade: dist
Inventory Patterns
Split Inventory by Region
inventories/
├── us-east/
│ ├── hosts.yml
│ └── group_vars/
├── us-west/
│ ├── hosts.yml
│ └── group_vars/
└── eu-west/
├── hosts.yml
└── group_vars/
# Run against specific region
ansible-playbook site.yml -i inventories/us-east/
# Run against all regions
ansible-playbook site.yml -i inventories/
Dynamic Inventory for Cloud
# aws_ec2.yml — auto-discover EC2 instances
plugin: amazon.aws.aws_ec2
regions:
- us-east-1
- us-west-2
keyed_groups:
- key: tags.Environment
prefix: env
- key: instance_type
prefix: type
- key: placement.availability_zone
prefix: az
filters:
tag:ManagedBy: ansible
instance-state-name: running
compose:
ansible_host: private_ip_address
Serial Execution for Safety
---
- name: Rolling update across 1000 hosts
hosts: webservers
serial:
- 1 # First: 1 canary host
- 5 # Then: 5 hosts
- "10%" # Then: 10% at a time
max_fail_percentage: 5
pre_tasks:
- name: Drain from load balancer
ansible.builtin.uri:
url: "http://{{ lb_host }}/api/drain/{{ inventory_hostname }}"
method: POST
roles:
- deploy-app
post_tasks:
- name: Health check
ansible.builtin.uri:
url: "http://{{ inventory_hostname }}:{{ app_port }}/health"
retries: 10
delay: 5
- name: Re-enable in load balancer
ansible.builtin.uri:
url: "http://{{ lb_host }}/api/enable/{{ inventory_hostname }}"
method: POST
Pull Mode with ansible-pull
For very large fleets (10,000+ hosts), push mode hits limits. Pull mode flips the model — each host runs ansible-pull on a schedule:
# Deploy ansible-pull via cron on all hosts
- name: Configure ansible-pull
hosts: all
become: true
tasks:
- name: Install ansible
ansible.builtin.package:
name: ansible-core
state: present
- name: Create pull cron job
ansible.builtin.cron:
name: "ansible-pull"
minute: "*/30"
job: >
ansible-pull
-U https://git.example.com/infra/site.git
-d /opt/ansible-pull
-i localhost,
-e "env={{ env }}"
local.yml
>> /var/log/ansible-pull.log 2>&1
AWX / Automation Controller
For enterprise scale, AWX (or Red Hat AAP) provides:
Architecture for 5,000+ hosts:
┌─────────────────┐ ┌──────────────┐
│ AWX Web UI │────▶│ PostgreSQL │
│ REST API │ │ (HA cluster) │
└────────┬────────┘ └──────────────┘
│
┌────┴────┐
│ Receptor│ ← Automation Mesh
│ Mesh │
├─────────┤
│ Hop │────▶ Execution Nodes (region 1)
│ Node │────▶ Execution Nodes (region 2)
│ │────▶ Execution Nodes (region 3)
└─────────┘
Key features for scale:
- Automation Mesh: Distribute execution across regions
- Instance Groups: Dedicate capacity to teams
- Job Slicing: Split one job across multiple nodes
- Smart Inventories: Dynamic host grouping
- Credential rotation: Centralized secret management
Execution Environments
Package Ansible + dependencies in containers for consistent execution:
# execution-environment.yml
---
version: 3
dependencies:
galaxy:
collections:
- amazon.aws
- community.vmware
- kubernetes.core
python:
- boto3>=1.28
- pyVmomi>=8.0
- kubernetes>=28.0
system:
- openssh-clients
- sshpass
build_arg_defaults:
ANSIBLE_GALAXY_CLI_COLLECTION_OPTS: "--pre"
images:
base_image:
name: quay.io/ansible/ansible-runner:latest
ansible-builder build -t my-ee:latest
ansible-navigator run site.yml --eei my-ee:latest
Monitoring and Observability
# Callback plugin for Prometheus metrics
[defaults]
callback_whitelist = timer, profile_tasks, prometheus
# Export to Grafana via ARA
[defaults]
callback_whitelist = ara_default
# Profile task execution times
ANSIBLE_CALLBACKS_ENABLED=profile_tasks ansible-playbook site.yml
Benchmark: Tuning Impact
| Configuration | 100 hosts | 500 hosts | 1000 hosts |
|---|---|---|---|
| Default (forks=5) | 8m 30s | 42m | 85m |
| forks=50 + pipelining | 2m 10s | 10m | 20m |
| + fact caching | 1m 05s | 5m | 10m |
| + free strategy | 0m 48s | 3m 30s | 7m |
| + Mitogen | 0m 20s | 1m 30s | 3m |
Related Articles
- Ansible Cache Plugins
- Ansible Strategy Plugins
- Ansible Async and Poll
- Ansible AWX Guide
- Ansible Automation Platform
Conclusion
Scaling Ansible requires a combination of configuration tuning (forks, pipelining, fact caching), architectural patterns (serial batches, pull mode, automation mesh), and tooling (AWX, execution environments). Start with the ansible.cfg baseline above and layer in complexity as your fleet grows.