Magento is resource-hungry by design - an application server under every request, not a static site. Undersized infrastructure shows up as slow pages at normal traffic and outages at peak; oversized infrastructure is just wasted money. Sizing is measurable, not mystical. Here is the practical guide.
The Starting Matrix
For a single-web-node production store (separate database):
| Store profile | vCPU | RAM |
|---|---|---|
| Small (few thousand visits/day) | 2 | 8GB |
| Mid (tens of thousands/day) | 4-8 | 16GB |
| Large / peak-season retail | 8-16 | 32GB+ |
RAM breaks down roughly: PHP-FPM workers are the big consumer (80-150MB per active worker on real stores), then MySQL’s buffer pool, Redis, OpenSearch, and OS. A 4GB “everything box” running Magento + MySQL + Redis + OpenSearch is not a production plan - it is a demo.
PHP-FPM: The Knob That Matters Most
pm.max_children decides how many concurrent PHP requests you serve. The calculation:
max_children = (RAM available for PHP) / (measured worker size)
Measure worker size under real load (ps -o rss -C php-fpm averaged over busy minutes) - 80-150MB is typical. On a 16GB box reserving 6GB for PHP: roughly 40-70 workers.
The failure modes at both extremes:
- Too many children: memory exhaustion under load, swapping, death spiral - worse than queuing
- Too few: requests queue behind workers while CPU idles; the fix people then reach for (raising children without measuring) creates the first failure
Use pm = dynamic with sensible start_servers/min_spare/max_spare, or pm = ondemand for quieter stores. Always set pm.max_requests (500-1000) so workers recycle before Magento’s memory leaks accumulate.
CPU and Concurrency Reality
Magento’s PHP is largely single-threaded per request: more cores, more concurrent requests. CPU saturation shows as rising response time with healthy memory - the signal to scale out (add web nodes) or up. Watch Load average / vCPU count: sustained values above ~0.7 per core mean you are out of headroom.
Plan for Peak, Not Average
Black Friday traffic is 5-20x a normal Tuesday. Options, in order of sanity:
- Headroom sizing: size for peak, pay for it year-round - simple, expensive
- Autoscaling web tier: scale-out on CPU/queue depth - requires stateless web nodes (shared media, central sessions in Redis, deploys to all nodes)
- Pre-event scaling: scheduled capacity increases for known peaks - the pragmatic favourite
Whatever the plan, load-test it: a synthetic checkout flow at target concurrency, watching FPM saturation and database connections. The first time your sizing meets reality should not be with real customers.
Sizing is a measurement habit: know your worker size, your peak ratio, your saturation signals. The stores that sail through peak season are the ones that rehearsed it.