Skip to content

.htaccess gzip and Brotli: browser caching and blocking AI bots in Apache

Published 7 October 20268 min readBy the SmoothSeen editorial team

In Apache .htaccess, compression is switched on with AddOutputFilterByType for Brotli and gzip, on text types only. Browser caching is set with mod_expires per file type, plus a one-year immutable Cache-Control reserved for fingerprinted file names. AI training bots are blocked with mod_rewrite and a 403, while the search crawlers that decide whether AI assistants cite you stay allowed.

Key points

  • With mod_brotli and mod_deflate, a 6,143-byte HTML page travelled as 301 bytes with Brotli and 414 with gzip in our test on Apache httpd 2.4.69.
  • WebP, JPEG and AVIF images are already compressed; adding them to the compression list costs CPU and saves no bytes.
  • A one-year cache with immutable is only safe for fingerprinted file names; HTML should carry no-cache so it is revalidated on every visit.
  • Blocking GPTBot, ClaudeBot or CCBot is an editorial choice that does not remove you from ChatGPT or Claude; blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot does.
  • A user-agent block does not stop anyone who spoofs it, and Google-Extended and Applebot-Extended can only be controlled in robots.txt.

To check it on your own site: SEO audit

On this page

Apache decides three things in this guide: whether a file travels compressed, how long the browser may keep it and who gets a 403. It does not control what a CDN in front of it does, and it cannot tell a genuine bot from an impostor. Every snippet was tested on 7 October 2026 with Apache httpd 2.4.69 in its official Docker image (httpd:2.4), using curl from a second container. HTTPS redirects and security headers are covered in the guide to .htaccess for HTTPS, HSTS and security headers.

Which modules do you need, and how do you know they are there?

Module
mod_deflate
Purpose
gzip compression
Note
Sends Vary: Accept-Encoding on its own1
Module
mod_brotli
Purpose
Brotli compression
Note
Available since Apache 2.4.262
Module
mod_filter
Purpose
The AddOutputFilterByType directive
Note
Loaded in the stock httpd.conf
Module
mod_expires
Purpose
Expires and the max-age in Cache-Control, by type
Note
Allowed in .htaccess3
Module
mod_headers
Purpose
Any other header, such as immutable
Note
Module
mod_rewrite
Purpose
The user-agent block
Note

The official image ships all six, but its stock httpd.conf leaves mod_deflate, mod_brotli, mod_expires and mod_rewrite commented out. On shared hosting you cannot see this directly: ask support or use the checks at the end. If a module is not loaded, the directive causes a 500 error, which is a good thing: you find out. That is why this .htaccess does not use <IfModule>, which would hide the failure.

A reminder from the Apache documentation: if you can edit the main configuration, put this in the VirtualHost, because .htaccess is read on every request4.

Step 1: Brotli and gzip compression

AddOutputFilterByType BROTLI_COMPRESS text/html text/plain text/css text/javascript application/javascript application/json application/xml image/svg+xml
AddOutputFilterByType DEFLATE text/html text/plain text/css text/javascript application/javascript application/json application/xml image/svg+xml

In our test, Apache answered with Brotli when the request accepted br and gzip, with gzip when it only accepted gzip, and uncompressed when it asked for no compression. These are the bytes transferred:

File
HTML (index.html)
Uncompressed
6,143
gzip
414
Brotli
301
File
CSS
Uncompressed
4,710
gzip
286
Brotli
156
File
JavaScript
Uncompressed
5,670
gzip
302
Brotli
192
File
WebP image
Uncompressed
4,000
gzip
4,000
Brotli
4,000

The test files are highly repetitive and compress better than a real site; what matters is the order of magnitude and that Brotli beat gzip on all three text files.

Three details worth knowing:

  • List both text/javascript and application/javascript. Apache 2.4.69's mime.types serves .js as text/javascript; many older examples only list application/javascript, and the JavaScript goes out uncompressed.
  • Leave images out. WebP, JPEG, PNG and AVIF are already compressed, and the mod_deflate documentation excludes them in its examples1. SVG is text and does compress.
  • BREACH. Both modules' documentation warns that some web applications are vulnerable to information disclosure when a TLS connection carries compressed data, pointing to the BREACH family of attacks12. If your site handles tokens or personal data on dynamic pages, review this before enabling it there.

Step 2: browser caching without serving stale files

web.dev's rule is simple: URLs with a version or fingerprint in the name can be cached for a year (max-age=31536000), and unversioned URLs such as HTML should carry no-cache, which does not mean "do not store" but "revalidate before use"5. MDN adds that immutable says the response will not change while it is fresh, designed precisely for fingerprinted files6.

ExpiresActive On
ExpiresByType text/css "access plus 1 week"
ExpiresByType text/javascript "access plus 1 week"
ExpiresByType image/webp "access plus 1 month"
ExpiresByType image/avif "access plus 1 month"
ExpiresByType image/png "access plus 1 month"
ExpiresByType image/jpeg "access plus 1 month"
ExpiresByType font/woff2 "access plus 1 year"

# HTML may be stored, but is revalidated on every visit
Header set Cache-Control "no-cache" "expr=%{CONTENT_TYPE} =~ m#^text/html#"

# Fingerprinted file names (app.3f9a1c2b.js): one year, immutable
<FilesMatch "\.[0-9a-f]{8,}\.(css|js|mjs|woff2|webp|avif|png|jpe?g|svg)$">
    Header set Cache-Control "max-age=31536000, immutable"
</FilesMatch>

The headers Apache 2.4.69 returned:

Request
/ (HTML)
Cache-Control returned
no-cache
Request
/assets/styles.css
Cache-Control returned
max-age=604800
Request
/assets/app.3f9a1c2b.js
Cache-Control returned
max-age=31536000, immutable
Request
/assets/photo.webp
Cache-Control returned
max-age=2592000

Two choices in this block:

  1. HTML is matched by content type, not extension. The expr= condition on the content type covers any HTML response, whether or not its file ends in .html. If your CMS already sends its own Cache-Control, check with curl which one wins.
  2. The fingerprinted JavaScript also gets a one-week Expires from mod_expires. That is harmless: when a response carries max-age, the HTTP standard requires caches to ignore Expires7.

The mistake that is hardest to spot: giving immutable to an unfingerprinted styles.css. Change that file tomorrow and anyone who already has it cached keeps the old version for up to a year. If your CMS does not fingerprint file names, stick to a week.

Step 3: AI bots, keeping search and training apart

AI bots do different jobs, and blocking the wrong one has very different consequences:

Bot
OAI-SearchBot
Company
OpenAI
What its owner uses it for
Showing sites in ChatGPT search
If you block it
You will not appear in ChatGPT search answers8
Bot
GPTBot
Company
OpenAI
What its owner uses it for
Training models
If you block it
Your content is not used for training; you can still appear in ChatGPT8
Bot
Claude-SearchBot and Claude-User
Company
Anthropic
What its owner uses it for
Search and user queries in Claude
If you block it
It may reduce your visibility in Claude9
Bot
ClaudeBot
Company
Anthropic
What its owner uses it for
Training models
If you block it
Your future content is excluded from training9
Bot
PerplexityBot
Company
Perplexity
What its owner uses it for
Linking sites in its results; not used for training
If you block it
Your pages may stop appearing as links in its results10
Bot
CCBot
Company
Common Crawl
What its owner uses it for
Open web archive
If you block it
You leave that archive11

The block we tested keeps out three training crawlers (GPTBot, ClaudeBot, meta-externalagent) plus Common Crawl's, and always leaves robots.txt open:

RewriteEngine On
RewriteCond %{REQUEST_URI} !^/robots\.txt$
RewriteCond %{HTTP_USER_AGENT} (GPTBot|ClaudeBot|CCBot|meta-externalagent) [NC]
RewriteRule ^ - [F]

With user-agents like those these bots send, the four blocked ones got 403 on /blog/ and 200 on /robots.txt. OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot, Googlebot and an ordinary Chrome got 200 on both. The robots.txt exception matters: Anthropic explains that blocking its bots by IP stops them reading your robots.txt and may not guarantee an opt-out9.

Four limits before you use it:

  • User-agents can be spoofed. Common Crawl says it knows of crawlers pretending to be CCBot and recommends checking the IP ranges it publishes11. OpenAI and Perplexity publish theirs too810.
  • Google-Extended and Applebot-Extended are not user-agents. Google-Extended is only a robots.txt token, does not affect Google Search and has no user-agent of its own12. Applebot-Extended does not crawl pages13. An .htaccess rule never sees them: control them in robots.txt.
  • ChatGPT-User and Perplexity-User may not follow robots.txt, according to their owners, because they act when a user asks810. Block them at the server and you disappear from those requests.
  • A 403 is invisible in robots.txt. If a colleague only checks robots.txt, they will not know the block exists. Document it.

For bots that honour robots.txt, robots.txt is the cleanest route: each company documents it and anyone can read it. The .htaccess block is the backstop:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

User-agent: *
Allow: /

How to check it

  1. Compression. curl -sI -H "Accept-Encoding: br, gzip" https://www.example.com/ should show Content-Encoding: br; with "Accept-Encoding: gzip", Content-Encoding: gzip. The Network panel in your browser's developer tools shows it too.
  2. Caching. curl -sI the home page, a CSS file and a fingerprinted file, and compare each Cache-Control with the table in step 2.
  3. Bots. curl -sI -A "GPTBot/1.4" https://www.example.com/ should give 403; curl -sI -A "OAI-SearchBot/1.4" https://www.example.com/, 200; and /robots.txt with the blocked user-agent, 200.
  4. Request a page after every change. httpd -t does not read .htaccess files: an error only shows up as a 500 when a page is requested.

What SmoothSeen does with this

Within its SEO analysis, SmoothSeen checks whether your site serves text compressed with gzip or Brotli. In its AI search analysis it reads your robots.txt keeping search crawlers apart from training crawlers, without counting a training block as a failure, and requests the page with each AI crawler's user-agent, a browser's and Googlebot's to spot a server or firewall answering 403 to any of them.

What to do this week

Request your home page with curl -sI -H "Accept-Encoding: br, gzip" and with -A "OAI-SearchBot/1.4". If it is not compressed, or ChatGPT's search crawler gets a 403 nobody decided on, start there. To check compression, robots.txt and server-level blocks in one go, analyse your site with SmoothSeen.

Frequently asked questions

Brotli or gzip: which should I enable?

Both. The browser announces on each request which algorithms it accepts; with both filters active, Apache used Brotli in our test whenever it was accepted and gzip otherwise. That way modern browsers get Brotli and the rest get gzip. Brotli cut our HTML to 301 bytes against 414 with gzip. If your host has no mod_brotli, gzip alone is still a big improvement.

Why are my CSS changes not showing after enabling caching?

Because the browser keeps using its stored copy until it expires. That is what happens when you give a long cache, or immutable, to a file whose name does not change when its content does. The fix is to add a fingerprint or version number to the file name on every release, or to cut the cache for those files to a week.

Does blocking GPTBot remove me from ChatGPT?

No. OpenAI explains that GPTBot is used to train models while ChatGPT search relies on OAI-SearchBot, with independent settings for each. You can opt out of training and still appear in answers that use search. What does remove you is blocking OAI-SearchBot, either in robots.txt or with a 403 at the server.

Is it better to block bots in robots.txt or in .htaccess?

For bots that honour robots.txt, use robots.txt: it is the route their owners document, it is visible from outside and a typo in a bot name breaks nothing. .htaccess is useful as a backstop or for bots that ignore robots.txt, but it only stops bots that identify themselves honestly and is invisible to anyone reviewing robots.txt.

Sources

  1. 1Apache Module mod_deflate, Apache Software Foundation, accessed 7 October 2026.
  2. 2Apache Module mod_brotli, Apache Software Foundation, accessed 7 October 2026.
  3. 3Apache Module mod_expires, Apache Software Foundation, accessed 7 October 2026.
  4. 4Apache HTTP Server Tutorial: .htaccess files, Apache Software Foundation, accessed 7 October 2026.
  5. 5Prevent unnecessary network requests with the HTTP Cache, web.dev (Google), accessed 7 October 2026.
  6. 6Cache-Control, MDN Web Docs, updated 17 September 2026.
  7. 7RFC 9111: HTTP Caching, section 5.3 (Expires), IETF, June 2022.
  8. 8Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
  9. 9Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
  10. 10Perplexity Crawlers, Perplexity, accessed 7 October 2026.
  11. 11CCBot, Common Crawl, accessed 7 October 2026.
  12. 12Google's common crawlers: Google-Extended, Google Search Central, updated 14 July 2026.
  13. 13About Applebot, Apple, published 4 September 2026.

How to cite this article

SmoothSeen. (2026, October 7). .htaccess gzip and Brotli: browser caching and blocking AI bots in Apache. https://smoothseen.com/en/blog/apache-htaccess-compression-caching-bots/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First version.

Keep reading

.htaccess gzip, Brotli, browser caching and AI bots