Introduction

When Ansible playbooks slow down or crash on large inventories, the culprit is almost always memory, not CPU. The control node holds inventory, facts, variables, and task state all in RAM simultaneously. Understanding this helps you right-size your control node and optimize playbooks for thousands of hosts.

Memory vs CPU: Where the Work Happens

ComponentControl Node (local)Managed Nodes (remote)
Inventory loading✅ RAM-heavy—
Fact gathering✅ Stored in RAM✅ Collected via SSH
Variable resolution✅ In-memory—
Task coordination✅ Per-fork overhead—
Module executionMinimal✅ Runs here via SSH
Template rendering✅ Jinja2 in RAM—

Key insight: Ansible is agentless — modules execute on remote nodes via SSH. The control node only coordinates, but coordination requires holding everything in memory.

Why Memory Is the Bottleneck

1. Inventory in Memory

Ansible loads the entire inventory at startup:

1,000 hosts × ~5KB per host (vars, groups) = ~5MB
10,000 hosts × ~5KB = ~50MB
50,000 hosts × ~10KB (complex vars) = ~500MB

2. Fact Gathering Multiplier

Each host's facts consume ~100-300KB:

1,000 hosts × 200KB facts = ~200MB
10,000 hosts × 200KB = ~2GB

3. Forks (Parallel Execution)

Each fork is a separate Python process:

ForksBase Memory per Fork100 Hosts1,000 Hosts
5 (default)~50MB~250MB~500MB
20~50MB~1GB~2GB
50~50MB~2.5GB~5GB+

4. Variables and Templates

Large variable files, Jinja2 templates, and registered results accumulate:

# This registers output for EVERY host in memory
- name: Get service status
  ansible.builtin.command: systemctl status myapp
  register: service_status  # × 1,000 hosts = significant RAM

5. Callback Plugins and Logging

Plugins like json, junit, or custom callbacks store results in memory until playbook completion.

Measuring Memory Usage

Monitor During Playbook Run

# Watch memory usage in real-time
watch -n 1 'ps aux --sort=-rss | head -10'

# Track Ansible process memory
watch -n 1 'ps aux | grep ansible | grep -v grep | awk "{sum+=\$6} END {print sum/1024 \" MB\"}"'

Profile a Specific Playbook

# Using /usr/bin/time for peak memory
/usr/bin/time -v ansible-playbook -i inventory site.yml 2>&1 | grep "Maximum resident"
# Maximum resident set size (kbytes): 524288

Check Before and After

free -h
ansible-playbook site.yml &
sleep 30 && free -h

8 Optimization Strategies

1. Disable Unnecessary Fact Gathering

# Skip facts entirely
- hosts: all
  gather_facts: false
  tasks:
    - name: Deploy config (no facts needed)
      ansible.builtin.copy:
        src: app.conf
        dest: /etc/myapp/app.conf

2. Gather Only Needed Facts

# Collect only network facts
- hosts: all
  gather_facts: true
  gather_subset:
    - "!all"
    - "!min"
    - network
gather_subsetApproximate Size
all (default)200-300KB/host
min10-20KB/host
network30-50KB/host
hardware50-100KB/host

3. Enable Fact Caching

# ansible.cfg
[defaults]
gathering = smart
fact_caching = jsonfile
fact_caching_connection = /tmp/ansible_facts_cache
fact_caching_timeout = 86400  # 24 hours

Or use Redis for multi-node setups:

[defaults]
fact_caching = redis
fact_caching_connection = localhost:6379:0
fact_caching_timeout = 86400

4. Right-Size Forks

# ansible.cfg
[defaults]
forks = 20  # Balance: more parallelism vs more memory

# Rule of thumb:
# Available RAM (GB) × 10 = approximate max forks
# 8GB → ~80 forks max
# 16GB → ~160 forks max

5. Use --limit for Large Inventories

# Process in batches
ansible-playbook site.yml --limit 'webservers[0:99]'
ansible-playbook site.yml --limit 'webservers[100:199]'

6. Avoid Storing Large Registered Variables

# BAD — stores all output in memory
- command: cat /var/log/syslog
  register: huge_log

# GOOD — only check return code
- command: grep -q "ERROR" /var/log/syslog
  register: error_check
  failed_when: false
  changed_when: false

7. Use free Strategy for Memory-Constrained Environments

- hosts: all
  strategy: free  # Don't wait for slowest host per task
  tasks: [...]

8. Split Large Playbooks

# Instead of one playbook for 10,000 hosts
# Split by groups and run separately
- import_playbook: webservers.yml
- import_playbook: databases.yml
- import_playbook: monitoring.yml

Control Node Sizing Guide

Managed HostsMinimum RAMRecommended RAMForks
< 1002 GB4 GB10-20
100-5004 GB8 GB20-50
500-2,0008 GB16 GB30-50
2,000-10,00016 GB32 GB50-100
10,000+32 GB+64 GB+Use AAP/Tower

Conclusion

Memory is Ansible's bottleneck because the control node holds inventory, facts, variables, and fork state all in RAM. Optimize by disabling unnecessary fact gathering (gather_subset), enabling fact caching (Redis/jsonfile), right-sizing forks to available RAM, avoiding large registered variables, and using --limit for batch processing. Size your control node based on managed host count — 1GB RAM per 100-200 hosts is a good starting estimate.