Introduction
When Ansible playbooks slow down or crash on large inventories, the culprit is almost always memory, not CPU. The control node holds inventory, facts, variables, and task state all in RAM simultaneously. Understanding this helps you right-size your control node and optimize playbooks for thousands of hosts.
Memory vs CPU: Where the Work Happens
| Component | Control Node (local) | Managed Nodes (remote) |
|---|---|---|
| Inventory loading | ✅ RAM-heavy | — |
| Fact gathering | ✅ Stored in RAM | ✅ Collected via SSH |
| Variable resolution | ✅ In-memory | — |
| Task coordination | ✅ Per-fork overhead | — |
| Module execution | Minimal | ✅ Runs here via SSH |
| Template rendering | ✅ Jinja2 in RAM | — |
Key insight: Ansible is agentless — modules execute on remote nodes via SSH. The control node only coordinates, but coordination requires holding everything in memory.
Why Memory Is the Bottleneck
1. Inventory in Memory
Ansible loads the entire inventory at startup:
1,000 hosts × ~5KB per host (vars, groups) = ~5MB
10,000 hosts × ~5KB = ~50MB
50,000 hosts × ~10KB (complex vars) = ~500MB
2. Fact Gathering Multiplier
Each host's facts consume ~100-300KB:
1,000 hosts × 200KB facts = ~200MB
10,000 hosts × 200KB = ~2GB
3. Forks (Parallel Execution)
Each fork is a separate Python process:
| Forks | Base Memory per Fork | 100 Hosts | 1,000 Hosts |
|---|---|---|---|
| 5 (default) | ~50MB | ~250MB | ~500MB |
| 20 | ~50MB | ~1GB | ~2GB |
| 50 | ~50MB | ~2.5GB | ~5GB+ |
4. Variables and Templates
Large variable files, Jinja2 templates, and registered results accumulate:
# This registers output for EVERY host in memory
- name: Get service status
ansible.builtin.command: systemctl status myapp
register: service_status # × 1,000 hosts = significant RAM
5. Callback Plugins and Logging
Plugins like json, junit, or custom callbacks store results in memory until playbook completion.
Measuring Memory Usage
Monitor During Playbook Run
# Watch memory usage in real-time
watch -n 1 'ps aux --sort=-rss | head -10'
# Track Ansible process memory
watch -n 1 'ps aux | grep ansible | grep -v grep | awk "{sum+=\$6} END {print sum/1024 \" MB\"}"'
Profile a Specific Playbook
# Using /usr/bin/time for peak memory
/usr/bin/time -v ansible-playbook -i inventory site.yml 2>&1 | grep "Maximum resident"
# Maximum resident set size (kbytes): 524288
Check Before and After
free -h
ansible-playbook site.yml &
sleep 30 && free -h
8 Optimization Strategies
1. Disable Unnecessary Fact Gathering
# Skip facts entirely
- hosts: all
gather_facts: false
tasks:
- name: Deploy config (no facts needed)
ansible.builtin.copy:
src: app.conf
dest: /etc/myapp/app.conf
2. Gather Only Needed Facts
# Collect only network facts
- hosts: all
gather_facts: true
gather_subset:
- "!all"
- "!min"
- network
| gather_subset | Approximate Size |
|---|---|
all (default) | 200-300KB/host |
min | 10-20KB/host |
network | 30-50KB/host |
hardware | 50-100KB/host |
3. Enable Fact Caching
# ansible.cfg
[defaults]
gathering = smart
fact_caching = jsonfile
fact_caching_connection = /tmp/ansible_facts_cache
fact_caching_timeout = 86400 # 24 hours
Or use Redis for multi-node setups:
[defaults]
fact_caching = redis
fact_caching_connection = localhost:6379:0
fact_caching_timeout = 86400
4. Right-Size Forks
# ansible.cfg
[defaults]
forks = 20 # Balance: more parallelism vs more memory
# Rule of thumb:
# Available RAM (GB) × 10 = approximate max forks
# 8GB → ~80 forks max
# 16GB → ~160 forks max
5. Use --limit for Large Inventories
# Process in batches
ansible-playbook site.yml --limit 'webservers[0:99]'
ansible-playbook site.yml --limit 'webservers[100:199]'
6. Avoid Storing Large Registered Variables
# BAD — stores all output in memory
- command: cat /var/log/syslog
register: huge_log
# GOOD — only check return code
- command: grep -q "ERROR" /var/log/syslog
register: error_check
failed_when: false
changed_when: false
7. Use free Strategy for Memory-Constrained Environments
- hosts: all
strategy: free # Don't wait for slowest host per task
tasks: [...]
8. Split Large Playbooks
# Instead of one playbook for 10,000 hosts
# Split by groups and run separately
- import_playbook: webservers.yml
- import_playbook: databases.yml
- import_playbook: monitoring.yml
Control Node Sizing Guide
| Managed Hosts | Minimum RAM | Recommended RAM | Forks |
|---|---|---|---|
| < 100 | 2 GB | 4 GB | 10-20 |
| 100-500 | 4 GB | 8 GB | 20-50 |
| 500-2,000 | 8 GB | 16 GB | 30-50 |
| 2,000-10,000 | 16 GB | 32 GB | 50-100 |
| 10,000+ | 32 GB+ | 64 GB+ | Use AAP/Tower |
Related Articles
- Ansible Best Practices Guide
- Ansible Configuration Settings
- Ansible Roles Explained
- Ansible Automation Platform Guide
Conclusion
Memory is Ansible's bottleneck because the control node holds inventory, facts, variables, and fork state all in RAM. Optimize by disabling unnecessary fact gathering (gather_subset), enabling fact caching (Redis/jsonfile), right-sizing forks to available RAM, avoiding large registered variables, and using --limit for batch processing. Size your control node based on managed host count — 1GB RAM per 100-200 hosts is a good starting estimate.