Installing a web server, choosing between the common ones, and confirming each layer works before adding the next. Setting up web applications on a VPS deals with that sequence and it is the same here.
What decides how much traffic the machine actually carries is a number in the PHP configuration that most people never open.
The pool is the real limit
Each PHP request is handled by a worker process. The number of workers is a configured maximum. It is not derived from the hardware and it does not grow because the machine is large.
A server with thirty-two cores and a pool of five processes five requests at a time. The sixth waits.
The symptom is unmistakable once you know it: fast when quiet, stalling under load, with the processor and memory both apparently idle. Nothing is exhausted, because the constraint is a count instead of a resource.
This is regularly diagnosed as slow code, and hours go into profiling an application that is working correctly. Understanding load average walks through reading the machine while it happens.
Measure what one worker uses
The sizing calculation is simple and depends on one number you must measure rather than assume:
ps -ylC php-fpm --sort:rss | awk '{print $8/1024 " MB"}' | tail -20
Published figures are not yours. A lean application may use 30 MB per worker; a large site with many plugins may use 150 MB. The difference changes the answer by a factor of five.
Measure under real traffic, not on an idle machine, and take the larger figures in place of the average.
The calculation
Decide how much of the machine's memory the pool may have: total memory, minus what the database needs, minus a real reserve for the operating system. Then divide by the per-worker figure.
On a 16 GB machine giving the database 4 GB and reserving 2 GB, the pool has 10 GB. At 100 MB per worker that is about 100 workers.
The reserve is not optional. A machine with no headroom starts terminating processes when anything spikes, and the process chosen is frequently the database.
Too many is worse than too few
Raising the limit is the reflex when a site queues, and past a point it converts a slow site into a broken one.
Too few workers means requests wait and then succeed. Too many means memory is exhausted, the system begins killing processes, and requests fail outright, while the machine also spends its time swapping rather than serving.
Slow is recoverable. Killed is an outage. When uncertain, size conservatively and raise it deliberately while watching memory.
Static or dynamic
A dynamic pool starts a few workers and adds them under load, which saves memory on a machine doing other things.
A static pool starts them all and keeps them. It uses the memory permanently and removes the delay of starting a worker at the moment traffic arrives.
On a dedicated server whose job is serving this site, static is usually right. The memory is not needed for anything else, and the behaviour is predictable rather than varying with load.
One pool per site
If the machine serves several sites, give each its own pool running as its own user.
Two things follow. One site's traffic cannot consume every worker and starve the others. And a compromise of one application does not reach another's files, because the processes run as different users.
Sizing then means dividing the memory between pools rather than setting one large number, which is more work and is the arrangement that keeps one busy site from taking down the rest.
Caching changes the arithmetic entirely
Every page served from cache is a request that never occupies a worker at all.
A site with effective caching may serve ten times the traffic on the same pool, because only the uncached requests reach PHP. That is a larger improvement than any pool tuning achieves, and it should come first.
Understanding caching layers goes over it, and optimizing server performance deals with where the remaining effort belongs.
Confirm the setting is the one in use
Changing a pool file and restarting the wrong service is a common way to conclude that tuning does nothing.
systemctl reload php-fpm ps -C php-fpm --no-headers | wc -l
Under load, that count should approach your limit. If it sits at a lower number regardless of traffic, the configuration being read is not the one you edited.
Read the queue, not only the workers
A pool sized correctly still fails if requests arrive faster than they complete, and the queue is where that becomes visible before anyone complains.
curl -s http://127.0.0.1/status?full 2>/dev/null | grep -E 'listen queue|idle processes|active processes' grep -c 'server reached pm.max_children' /var/log/php-fpm/error.log 2>/dev/null
A listen queue that is regularly above zero means requests are waiting for a worker. That is felt as a site which is fine most of the time and occasionally very slow, with no error anywhere.
The log message about reaching the maximum is the direct statement of the same problem. One appearance is a spike; a pattern of them means the pool is too small for the traffic or the requests are taking too long.
Find out why requests are slow before adding workers
Raising the worker count is the obvious response and it multiplies memory use to work around something else.
curl -s http://127.0.0.1/status?full 2>/dev/null | grep -A4 'request duration' | head ps -o pid,rss,etime,cmd -C php-fpm --sort=-rss 2>/dev/null | head -5
A worker occupied for several seconds is usually waiting on a database query or an external service rather than computing. Ten more workers waiting on the same slow thing consume ten times the memory and finish no faster.
Where the duration is the problem, fixing that reduces the required pool size rather than requiring a larger one, which is the cheaper outcome in both directions.
Confirm the pool is restarting cleanly
Workers are recycled after a number of requests, and a value set wrongly produces either memory growth or constant restarts.
grep -E 'pm.max_requests|pm.max_children|pm =' /etc/php-fpm.d/*.conf 2>/dev/null systemctl status php-fpm --no-pager | head -5
Without recycling, a small leak in the application accumulates until the machine runs out. With too aggressive a value, workers spend their time starting rather than serving.
Read the memory of the oldest workers against the newest. A large difference means something is accumulating, and the recycle value is masking a leak rather than solving it.