Blocking scanner traffic at Nginx for a leaner Statamic cache
Blocking scanner traffic at NGINX is the second-biggest win I have had auditing Statamic static caches, after stripping UTM tracking parameters. On a site with around 340 real pages, the cache had grown to 1,407 URLs. The bulk of the gap was bots probing for WordPress paths and PHP exploits, each one cached as if it were a real page.
What was actually in there
I ran the same Statamic\Facades\StaticCache::driver()->getUrls() script as last time, but tweaked the grouping to bucket URLs by their first path segment instead of by query string:
php artisan tinker --execute='
$urls = Statamic\Facades\StaticCache::driver()->getUrls();
$bySegment = $urls->map(function ($u) {
$path = parse_url($u, PHP_URL_PATH) ?? "/";
$parts = array_values(array_filter(explode("/", $path)));
return $parts[0] ?? "(root)";
})->countBy()->sortDesc();
echo "Top 30 path segments:\n";
$bySegment->take(30)->each(fn($c, $s) => printf("%6d /%s\n", $c, $s));
'The top of the list told the whole story:
Top 30 path segments:
103 /blog
96 /wp-includes
56 /wp-admin
48 /wp-content
44 /case-studies
41 /guides
28 /glossary
20 /__shared-errors
13 /services
13 /testimonials
12 /perch
12 /api
6 /admin
6 /config
5 /_ignition
5 /cgi-bin
5 /index.php
4 /vendorWordPress probes, Perch CMS scanners, Laravel Ignition exploit attempts, random PHP shell filenames. None of them belong to anything the site actually serves. All of them had been generated, returned to a scanner, and written to disk as cache entries.
I keep an eye on this kind of thing as part of website maintenance on the Statamic sites I look after.
Block at the edge, not at the cache
Statamic does ship a static_caching.exclude.urls config that keeps specific paths out of the cache. That works, but it is the wrong layer.
By the time Statamic checks an exclude rule, the request has already passed through NGINX, booted Laravel, started a PHP-FPM worker, and run through middleware. For a bot hitting fifty random URLs in a few seconds, that is a lot of work for a request you are going to throw away.
The cleaner fix is to never let those requests reach PHP.
First, stop bogus .php requests reaching PHP at all
Before any blocklist, there is one line that handles most of the volume. A Statamic site never serves a public .php file except the front controller, so the PHP location has no business falling back to it.
location ~ \.php$ {
try_files $uri =404; # NOT: try_files $uri /index.php =404;
# ...existing fastcgi_pass / fastcgi_param lines unchanged...
}With the fallback in place, a request for /shell.php is handed to the front controller, rendered by Statamic as a 404, and cached. Without it, NGINX answers the 404 itself and PHP is never touched. Real pages are unaffected, because they route through location /'s own try_files $uri $uri/ /index.php?$query_string.
The line that cancels that fix
This is the part that cost me an afternoon. Ploi's stock Laravel vhost carries error_page 404 /index.php; at server level. It intercepts the 404 that try_files ... =404 produces, so the request takes the long way round to exactly where you did not want it:
GET /shell.php
-> location ~ \.php$ , $uri missing -> =404
-> error_page 404 internally redirects to /index.php
-> re-matches location ~ \.php$ , where $uri is now /index.php, which exists
-> FastCGI -> Statamic renders a 404 -> cachedComment it out. Nothing on a Statamic site depends on it: location / already falls through to the front controller unconditionally, and Statamic renders its own 404 from PHP. The only 404s error_page was catching were the deliberate ones you just added.
One command tells you which of the two you have:
curl -sI https://example.com/nonexistent.phpJudge it on the headers, not the body size. A response that reached PHP always carries x-powered-by: Statamic and two set-cookie headers, for the XSRF token and the session. A bare NGINX 404 has neither. If you see the cookies, error_page is still winning and the try_files line is doing nothing.
You will often see content-length: 146 on the good response, which is NGINX's stock 404 page. Do not rely on it. With gzip enabled NGINX chunks that page instead, so the header disappears and vary: Accept-Encoding shows up in its place. Two sites I fixed the same week disagreed on the byte count and agreed on the headers.
Re-check both lines after any PHP version change or certificate operation in Ploi. Those regenerate the vhost from the panel's template, which restores error_page and silently re-breaks the fix, with no diff anywhere to catch it.
The NGINX config
On a Ploi-managed site, custom rules live in the per-site Server include directory. In the Ploi NGINX management UI, that is the Server section in the right sidebar. Add a new file called scanner-blocks.conf:
# scanner-blocks.conf - non-.php probe blocklist for Statamic.
# (*.php probes are handled by the PHP location's try_files $uri =404.)
# Backup / dump / config / secret file extensions.
location ~* \.(sql|sqlite|bak|old|env|log|ini|sh|swp|swo)$ { access_log off; log_not_found off; return 444; }
# Web-shell and env-exposure probes.
location ~* ^/(ALFA_DATA|manage/env|management/env)([/.]|$) { access_log off; log_not_found off; return 444; }
# WordPress / xmlrpc.
location ~* ^/(wp-admin|wp-content|wp-includes|wp-json|wp-config|wp-login|xmlrpc|wordpress)([/.]|$) { access_log off; log_not_found off; return 444; }
# CMS and control panels, plus Joomla directory names.
location ~* ^/(perch|cpanel|phpmyadmin|cgi-bin|administrator|joomla|drupal|modules|plugins|components|templates)([/.]|$) { access_log off; log_not_found off; return 444; }
# Laravel Ignition RCE and app internals.
location ~* ^/(_ignition|bootstrap|artisan)([/.]|$) { access_log off; log_not_found off; return 444; }
# Stale Statamic shared-error cache entries.
location ~* ^/__shared-errors([/.]|$) { access_log off; log_not_found off; return 444; }
# Credential / config probe filenames anywhere in the tree.
location ~* /(database|config|configuration|sa-private-key|aws|credentials|secrets|backup|dump)\.(php|json|yml|yaml|sql|env|bak|log|txt|old)$ { access_log off; log_not_found off; return 444; }
# Framework, panel and debug probes.
location ~* ^/(_profiler|_debugbar|symfony|debug|backend|portal|app|livewire|telescope|actuator)([/.]|$) { access_log off; log_not_found off; return 444; }
# VPN and enterprise scanners.
location ~* ^/(dana-na|RDWeb|owa|exchange)([/.]|$) { access_log off; log_not_found off; return 444; }
# Misrouted /public/*.
location ~* ^/public([/.]|$) { access_log off; log_not_found off; return 444; }
# --- Check these three against your own site before enabling ---
# Only if the site has no front-end member area (the Statamic CP is /cp).
location ~* ^/(auth|user|users|profile|password|register|signup|signin)([/.]|$) { access_log off; log_not_found off; return 444; }
# Only if the REST and GraphQL APIs are disabled.
location ~* ^/api([/.]|$) { access_log off; log_not_found off; return 444; }
# Allow ONLY the real /sitemap.xml; 404 every other sitemap-ish guess.
location ~* ^/(?!sitemap\.xml$).*sitemap[^/]*\.xml(\.gz)?$ { access_log off; log_not_found off; return 404; }Save in Ploi and it runs nginx -t then reloads automatically. If the config has a syntax error, Ploi shows you the error and refuses to apply it.
The return 444 directive is the interesting bit. It closes the connection without sending any response at all. No status code, no body, no headers. From the scanner's point of view, the request just dies. That is exactly what I want for hostile traffic. If you would rather see rejections in your error log for audit, swap 444 for 404 and remove the access_log off lines.
Two things about those patterns are worth spelling out, because getting either wrong is quiet rather than loud.
First, the ([/.]|$) on the end is doing real work. Without it the rule prefix-matches, so a list containing a generic English word like templates or app will silently kill a future page at /templates-for-events or /apple-store. Use the character class rather than plain (/|$) though: allowing the dot is what keeps /wp-login.php blocked, and that is one of the commonest probes on the internet.
Second, order matters. NGINX tries regex locations in definition order and stops at the first match, so appending a narrower duplicate of a rule you already have is dead config rather than extra protection. It will pass nginx -t quite happily and never execute.
Before deploying a list like this, check it against your own content. I ran the full rule set against every real URL on the site I was hardening, including collection entries, pagination, /cp, /img, /assets and /!/, and confirmed zero false positives, then against a list of known probe shapes to confirm zero misses. That takes a few minutes and it is the difference between hardening a site and 404ing a real page that nobody notices for a year.
Never block these on a Statamic site
These serve real traffic, so keep them clear of any rule: /vendor and /storage (Control Panel JS and uploads), /cp, /assets, /build, /img for Glide transforms, and /!/, which is Statamic's form, action and nocache endpoint. Also leave .well-known, robots.txt and sitemap.xml alone.
The /!/ one catches people out. Block it and every form on the site stops submitting, with no error anywhere in the application log, because the request never arrives.
Clear the cache and watch the count fall
With NGINX now rejecting scanner traffic before it reaches Statamic, clear the polluted entries:
php please static:clearRe-run the original tinker a day later. On this site the URL count dropped from 1,407 to around 380. That number lines up with the actual content: collection entries, taxonomy pages, pagination, plus a handful of index pages.
If your static cache is bigger than your sitemap, scanners are probably the reason.
What Statamic 6.31 changed, and what it did not
There is a related fix upstream worth knowing about if you run the half measure rather than the full one, because it covers the other half of this problem.
Under half measure the cacher stores 404s as well as 200s and records their URLs in a tracked set. With background recache switched on, every content save then re-fetched that whole set, junk included, and each dead URL threw on the way back and landed in failed_jobs. On one site I looked at, roughly 97 per cent of a 20,900-URL tracked set was scanner debris, and editors were waiting on it before their changes went live.
I reported that as statamic/cms#15054 and it was fixed by #15062, released in v6.31.0. The change is small: AbstractCacher::refreshUrl() now checks whether the cached response was an error and invalidates it instead of dispatching a warm job, on the reasoning that warming an error response would only fail the same way again.
What it deliberately did not change is the part that matters here. 404s are still cached and still tracked - the middleware still allows [200, 404] - because upstream wants them reachable by invalidation, so a stale 404 clears when you finally publish a page at that URL. The junk still arrives. Statamic just stops replaying it.
Which is the whole argument for doing this at the web server. Upstream cleans up after the traffic; NGINX stops it existing. A probe that reaches PHP still costs a full render and an FPM worker, and that is true whichever caching strategy you run, or none at all.
Why this matters
Two reasons, beyond the URL count itself.
First, every cached scanner URL is a file on disk. On a server hosting several Statamic sites, that turns into a lot of meaningless files in public/static/ to back up, rsync, and eventually clean up. Multiply by the number of scanner waves over a year and the disk impact is real.
Second, even cached responses to scanner traffic cost something. A polluted cache means more entries to invalidate when content changes, more files to walk during warming, and more PHP work overall. Blocking at NGINX returns the request in microseconds with no PHP touched, so the scanner gets nothing useful back, the cache stays clean and the Control Panel stays fast.
Same as the UTM fix, this is the kind of drift that does not show up until you go looking. The cache works, the site works, nobody complains. A year later the cache is full of /wp-admin entries and the framework cache directory is several gigabytes, all from traffic the site should never have answered.
You might also like...
- Tracer 2.0: my Statamic UTM builder gets site-wide settings
- Sentinel 2.3.0: faster scans for my Statamic security addon
- Warming Statamic static cache behind basic auth
- Over 250 installs on: Sentinel and CMS security for Statamic
- Statamic background recache smooths out bulk cache invalidations
- What a software supply chain attack means for your Statamic website
- A 7-day cooldown against npm supply-chain attacks
- Browser console errors: when the noise hides the signal
- Why my Statamic static cache hit 2.3GB (and how I fixed it)
- Introducing Sentinel: the Statamic monitoring tool I built for myself, now free for everyone