.htaccess gzip and Brotli: browser caching and blocking AI bots in Apache
Published 7 October 20268 min readBy the SmoothSeen editorial team
In Apache .htaccess, compression is switched on with AddOutputFilterByType for Brotli and gzip, on text types only. Browser caching is set with mod_expires per file type, plus a one-year immutable Cache-Control reserved for fingerprinted file names. AI training bots are blocked with mod_rewrite and a 403, while the search crawlers that decide whether AI assistants cite you stay allowed.
Key points
- With mod_brotli and mod_deflate, a 6,143-byte HTML page travelled as 301 bytes with Brotli and 414 with gzip in our test on Apache httpd 2.4.69.
- WebP, JPEG and AVIF images are already compressed; adding them to the compression list costs CPU and saves no bytes.
- A one-year cache with immutable is only safe for fingerprinted file names; HTML should carry no-cache so it is revalidated on every visit.
- Blocking GPTBot, ClaudeBot or CCBot is an editorial choice that does not remove you from ChatGPT or Claude; blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot does.
- A user-agent block does not stop anyone who spoofs it, and Google-Extended and Applebot-Extended can only be controlled in robots.txt.
To check it on your own site: SEO audit
On this page
Apache decides three things in this guide: whether a file travels compressed, how long the browser may keep it and who gets a 403. It does not control what a CDN in front of it does, and it cannot tell a genuine bot from an impostor. Every snippet was tested on 7 October 2026 with Apache httpd 2.4.69 in its official Docker image (httpd:2.4), using curl from a second container. HTTPS redirects and security headers are covered in the guide to .htaccess for HTTPS, HSTS and security headers.
Which modules do you need, and how do you know they are there?
- Module
- mod_deflate
- Purpose
- gzip compression
- Note
- Sends
Vary: Accept-Encodingon its own1
- Module
- mod_brotli
- Purpose
- Brotli compression
- Note
- Available since Apache 2.4.262
- Module
- mod_filter
- Purpose
- The
AddOutputFilterByTypedirective - Note
- Loaded in the stock
httpd.conf
- Module
- mod_expires
- Purpose
Expiresand themax-ageinCache-Control, by type- Note
- Allowed in .htaccess3
- Module
- mod_headers
- Purpose
- Any other header, such as
immutable - Note
- Module
- mod_rewrite
- Purpose
- The user-agent block
- Note
The official image ships all six, but its stock httpd.conf leaves mod_deflate, mod_brotli, mod_expires and mod_rewrite commented out. On shared hosting you cannot see this directly: ask support or use the checks at the end. If a module is not loaded, the directive causes a 500 error, which is a good thing: you find out. That is why this .htaccess does not use <IfModule>, which would hide the failure.
A reminder from the Apache documentation: if you can edit the main configuration, put this in the VirtualHost, because .htaccess is read on every request4.
Step 1: Brotli and gzip compression
AddOutputFilterByType BROTLI_COMPRESS text/html text/plain text/css text/javascript application/javascript application/json application/xml image/svg+xml
AddOutputFilterByType DEFLATE text/html text/plain text/css text/javascript application/javascript application/json application/xml image/svg+xmlIn our test, Apache answered with Brotli when the request accepted br and gzip, with gzip when it only accepted gzip, and uncompressed when it asked for no compression. These are the bytes transferred:
- File
- HTML (
index.html) - Uncompressed
- 6,143
- gzip
- 414
- Brotli
- 301
- File
- CSS
- Uncompressed
- 4,710
- gzip
- 286
- Brotli
- 156
- File
- JavaScript
- Uncompressed
- 5,670
- gzip
- 302
- Brotli
- 192
- File
- WebP image
- Uncompressed
- 4,000
- gzip
- 4,000
- Brotli
- 4,000
The test files are highly repetitive and compress better than a real site; what matters is the order of magnitude and that Brotli beat gzip on all three text files.
Three details worth knowing:
- List both
text/javascriptandapplication/javascript. Apache 2.4.69'smime.typesserves.jsastext/javascript; many older examples only listapplication/javascript, and the JavaScript goes out uncompressed. - Leave images out. WebP, JPEG, PNG and AVIF are already compressed, and the mod_deflate documentation excludes them in its examples1. SVG is text and does compress.
- BREACH. Both modules' documentation warns that some web applications are vulnerable to information disclosure when a TLS connection carries compressed data, pointing to the BREACH family of attacks12. If your site handles tokens or personal data on dynamic pages, review this before enabling it there.
Step 2: browser caching without serving stale files
web.dev's rule is simple: URLs with a version or fingerprint in the name can be cached for a year (max-age=31536000), and unversioned URLs such as HTML should carry no-cache, which does not mean "do not store" but "revalidate before use"5. MDN adds that immutable says the response will not change while it is fresh, designed precisely for fingerprinted files6.
ExpiresActive On
ExpiresByType text/css "access plus 1 week"
ExpiresByType text/javascript "access plus 1 week"
ExpiresByType image/webp "access plus 1 month"
ExpiresByType image/avif "access plus 1 month"
ExpiresByType image/png "access plus 1 month"
ExpiresByType image/jpeg "access plus 1 month"
ExpiresByType font/woff2 "access plus 1 year"
# HTML may be stored, but is revalidated on every visit
Header set Cache-Control "no-cache" "expr=%{CONTENT_TYPE} =~ m#^text/html#"
# Fingerprinted file names (app.3f9a1c2b.js): one year, immutable
<FilesMatch "\.[0-9a-f]{8,}\.(css|js|mjs|woff2|webp|avif|png|jpe?g|svg)$">
Header set Cache-Control "max-age=31536000, immutable"
</FilesMatch>The headers Apache 2.4.69 returned:
- Request
/(HTML)- Cache-Control returned
no-cache
- Request
/assets/styles.css- Cache-Control returned
max-age=604800
- Request
/assets/app.3f9a1c2b.js- Cache-Control returned
max-age=31536000, immutable
- Request
/assets/photo.webp- Cache-Control returned
max-age=2592000
Two choices in this block:
- HTML is matched by content type, not extension. The
expr=condition on the content type covers any HTML response, whether or not its file ends in.html. If your CMS already sends its ownCache-Control, check withcurlwhich one wins. - The fingerprinted JavaScript also gets a one-week
Expiresfrom mod_expires. That is harmless: when a response carriesmax-age, the HTTP standard requires caches to ignoreExpires7.
The mistake that is hardest to spot: giving immutable to an unfingerprinted styles.css. Change that file tomorrow and anyone who already has it cached keeps the old version for up to a year. If your CMS does not fingerprint file names, stick to a week.
Step 3: AI bots, keeping search and training apart
AI bots do different jobs, and blocking the wrong one has very different consequences:
- Bot
- OAI-SearchBot
- Company
- OpenAI
- What its owner uses it for
- Showing sites in ChatGPT search
- If you block it
- You will not appear in ChatGPT search answers8
- Bot
- GPTBot
- Company
- OpenAI
- What its owner uses it for
- Training models
- If you block it
- Your content is not used for training; you can still appear in ChatGPT8
- Bot
- Claude-SearchBot and Claude-User
- Company
- Anthropic
- What its owner uses it for
- Search and user queries in Claude
- If you block it
- It may reduce your visibility in Claude9
- Bot
- ClaudeBot
- Company
- Anthropic
- What its owner uses it for
- Training models
- If you block it
- Your future content is excluded from training9
- Bot
- PerplexityBot
- Company
- Perplexity
- What its owner uses it for
- Linking sites in its results; not used for training
- If you block it
- Your pages may stop appearing as links in its results10
- Bot
- CCBot
- Company
- Common Crawl
- What its owner uses it for
- Open web archive
- If you block it
- You leave that archive11
The block we tested keeps out three training crawlers (GPTBot, ClaudeBot, meta-externalagent) plus Common Crawl's, and always leaves robots.txt open:
RewriteEngine On
RewriteCond %{REQUEST_URI} !^/robots\.txt$
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|CCBot|meta-externalagent) [NC]
RewriteRule ^ - [F]With user-agents like those these bots send, the four blocked ones got 403 on /blog/ and 200 on /robots.txt. OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot, Googlebot and an ordinary Chrome got 200 on both. The robots.txt exception matters: Anthropic explains that blocking its bots by IP stops them reading your robots.txt and may not guarantee an opt-out9.
Four limits before you use it:
- User-agents can be spoofed. Common Crawl says it knows of crawlers pretending to be CCBot and recommends checking the IP ranges it publishes11. OpenAI and Perplexity publish theirs too810.
- Google-Extended and Applebot-Extended are not user-agents. Google-Extended is only a robots.txt token, does not affect Google Search and has no user-agent of its own12. Applebot-Extended does not crawl pages13. An .htaccess rule never sees them: control them in robots.txt.
- ChatGPT-User and Perplexity-User may not follow robots.txt, according to their owners, because they act when a user asks810. Block them at the server and you disappear from those requests.
- A 403 is invisible in robots.txt. If a colleague only checks robots.txt, they will not know the block exists. Document it.
For bots that honour robots.txt, robots.txt is the cleanest route: each company documents it and anyone can read it. The .htaccess block is the backstop:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
User-agent: *
Allow: /How to check it
- Compression.
curl -sI -H "Accept-Encoding: br, gzip" https://www.example.com/should showContent-Encoding: br; with"Accept-Encoding: gzip",Content-Encoding: gzip. The Network panel in your browser's developer tools shows it too. - Caching.
curl -sIthe home page, a CSS file and a fingerprinted file, and compare eachCache-Controlwith the table in step 2. - Bots.
curl -sI -A "GPTBot/1.4" https://www.example.com/should give 403;curl -sI -A "OAI-SearchBot/1.4" https://www.example.com/, 200; and/robots.txtwith the blocked user-agent, 200. - Request a page after every change.
httpd -tdoes not read .htaccess files: an error only shows up as a 500 when a page is requested.
What SmoothSeen does with this
Within its SEO analysis, SmoothSeen checks whether your site serves text compressed with gzip or Brotli. In its AI search analysis it reads your robots.txt keeping search crawlers apart from training crawlers, without counting a training block as a failure, and requests the page with each AI crawler's user-agent, a browser's and Googlebot's to spot a server or firewall answering 403 to any of them.
What to do this week
Request your home page with curl -sI -H "Accept-Encoding: br, gzip" and with -A "OAI-SearchBot/1.4". If it is not compressed, or ChatGPT's search crawler gets a 403 nobody decided on, start there. To check compression, robots.txt and server-level blocks in one go, analyse your site with SmoothSeen.
Frequently asked questions
Brotli or gzip: which should I enable?
Both. The browser announces on each request which algorithms it accepts; with both filters active, Apache used Brotli in our test whenever it was accepted and gzip otherwise. That way modern browsers get Brotli and the rest get gzip. Brotli cut our HTML to 301 bytes against 414 with gzip. If your host has no mod_brotli, gzip alone is still a big improvement.
Why are my CSS changes not showing after enabling caching?
Because the browser keeps using its stored copy until it expires. That is what happens when you give a long cache, or immutable, to a file whose name does not change when its content does. The fix is to add a fingerprint or version number to the file name on every release, or to cut the cache for those files to a week.
Does blocking GPTBot remove me from ChatGPT?
No. OpenAI explains that GPTBot is used to train models while ChatGPT search relies on OAI-SearchBot, with independent settings for each. You can opt out of training and still appear in answers that use search. What does remove you is blocking OAI-SearchBot, either in robots.txt or with a 403 at the server.
Is it better to block bots in robots.txt or in .htaccess?
For bots that honour robots.txt, use robots.txt: it is the route their owners document, it is visible from outside and a typo in a bot name breaks nothing. .htaccess is useful as a backstop or for bots that ignore robots.txt, but it only stops bots that identify themselves honestly and is invisible to anyone reviewing robots.txt.
Sources
- 1Apache Module mod_deflate, Apache Software Foundation, accessed 7 October 2026.
- 2Apache Module mod_brotli, Apache Software Foundation, accessed 7 October 2026.
- 3Apache Module mod_expires, Apache Software Foundation, accessed 7 October 2026.
- 4Apache HTTP Server Tutorial: .htaccess files, Apache Software Foundation, accessed 7 October 2026.
- 5Prevent unnecessary network requests with the HTTP Cache, web.dev (Google), accessed 7 October 2026.
- 6Cache-Control, MDN Web Docs, updated 17 September 2026.
- 7RFC 9111: HTTP Caching, section 5.3 (Expires), IETF, June 2022.
- 8Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
- 9Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
- 10Perplexity Crawlers, Perplexity, accessed 7 October 2026.
- 11CCBot, Common Crawl, accessed 7 October 2026.
- 12Google's common crawlers: Google-Extended, Google Search Central, updated 14 July 2026.
- 13About Applebot, Apple, published 4 September 2026.
How to cite this article
SmoothSeen. (2026, October 7). .htaccess gzip and Brotli: browser caching and blocking AI bots in Apache. https://smoothseen.com/en/blog/apache-htaccess-compression-caching-bots/
Keep reading
What is SEO? How search engine optimisation works and how to improve it in 2026
What SEO is, how Google decides which pages to show and a prioritised checklist to improve your rankings with free tools.
.htaccess force HTTPS: redirect to https, enable HSTS and add security headers in Apache
How to force HTTPS in .htaccess or an Apache VirtualHost, roll out HSTS safely and add security headers. Every snippet tested on Apache 2.4.69.
Broken links: what they are, how to find them and how to fix them
What a broken link is, what Google actually says about 404 errors, and how to find and fix broken links with Search Console, a crawler and curl.