Skip to content

Nginx gzip and Brotli: compression, browser caching and blocking AI bots

Published 7 October 20268 min readBy the SmoothSeen editorial team

In nginx, gzip is enabled with gzip on, gzip_types for the text types and gzip_vary on. Brotli is not in the nginx.org packages: it needs the ngx_brotli module, compiled separately. Browser caching is set per location with expires or add_header Cache-Control, and AI training bots are blocked with map and return 403, leaving AI search crawlers alone.

Key points

  • With gzip on, gzip_types and gzip_vary, a 6,143-byte HTML page travelled as 413 bytes in our test on nginx 1.30.5 and 1.31.6.
  • The nginx.org packages do not include Brotli. The ngx_brotli module is a Google project on GitHub that you have to compile; NGINX Plus offers it as a package.
  • A regular expression with braces in a location must be quoted; without quotes, nginx -t fails with "unknown directive".
  • A caching add_header in a location wipes out the server block's security headers; add_header_inherit merge in the server block prevents that since nginx 1.29.3.
  • Blocking GPTBot or ClaudeBot with map and return 403 does not remove you from ChatGPT or Claude; blocking OAI-SearchBot or Claude-SearchBot does.

To check it on your own site: SEO audit

On this page

Nginx decides three things here: whether a response travels compressed, how long the browser may keep it and who gets a 403. It does not control what a CDN in front of it does, and it cannot tell a genuine bot from an impostor. Every snippet was tested on 7 October 2026 with nginx 1.30.5 (stable) and nginx 1.31.6 (mainline) in their official Docker images, using curl from a second container; the exception is Brotli, and we explain why. HTTPS and security headers are in the guide to nginx HTTPS, HSTS and security headers.

Step 1: enable gzip

These lines go in the http block, or in a file under conf.d/, which the stock configuration includes inside http:

gzip on;
gzip_vary on;
gzip_min_length 1000;
gzip_types text/plain text/css text/javascript application/javascript application/json application/xml image/svg+xml;
Directive
gzip
What it does
Turns compression on
Default1
off
Directive
gzip_vary
What it does
Adds Vary: Accept-Encoding, so a proxy does not hand the compressed version to a client that cannot read it
Default1
off
Directive
gzip_min_length
What it does
Skips responses smaller than this, in bytes
Default1
20
Directive
gzip_types
What it does
Types to compress, on top of text/html, which is always compressed
Default1
text/html
Directive
gzip_comp_level
What it does
Level from 1 (fast) to 9 (smallest)
Default1
1
Directive
gzip_proxied
What it does
Whether to compress requests that arrive through a proxy
Default1
off

Bytes transferred in our test, identical on both versions:

File
HTML
Uncompressed
6,143
gzip
413
File
CSS
Uncompressed
4,710
gzip
277
File
JavaScript
Uncompressed
5,670
gzip
297
File
WebP image
Uncompressed
4,000
gzip
4,000

Three practical details:

  • Do not list text/html in gzip_types: it is always compressed.
  • List both JavaScript types. nginx's mime.types serves .js as application/javascript, but Apache's uses text/javascript; if nginx proxies another server, you may receive either.
  • Behind a CDN, look at gzip_proxied. nginx treats a request as proxied when it carries a Via header, and with the default off it does not compress it1.

The nginx documentation also warns that responses compressed over SSL/TLS may be vulnerable to BREACH attacks1: if your site handles tokens on dynamic pages, review this.

What about Brotli?

nginx does not ship Brotli. The nginx.org packages page lists its dynamic modules (geoip, image-filter, njs, perl, xslt, otel and acme) and Brotli is not among them2. The official Docker image has no Brotli module in /usr/lib/nginx/modules/ either: we checked.

There are two routes:

  1. ngx_brotli, a Google project on GitHub. You compile it as a dynamic module against your exact nginx version (./configure --with-compat --add-dynamic-module=...) and load it with load_module3.
  2. NGINX Plus, the commercial edition, which offers it as the nginx-plus-module-brotli package4.

This block, using the module and directive names the project documents, is untested, because the official images do not include the module:

# In nginx.conf, outside any block (main context)
load_module modules/ngx_http_brotli_filter_module.so;
load_module modules/ngx_http_brotli_static_module.so;

# In the http block
brotli on;
brotli_types text/plain text/css text/javascript application/javascript application/json application/xml image/svg+xml;

If compiling a third-party module, and redoing it on every upgrade, is not worth it for you, a well-configured gzip already covers the essentials.

Pre-compressed files with gzip_static

gzip_static serves the .gz file sitting next to the original instead of compressing on every request. The module is not built by default, but the official packages include it (--with-http_gzip_static_module shows up in nginx -V)5.

# In the http block, next to the gzip lines
gzip_static on;

In our test, with app.3f9a1c2b.js.gz created by gzip -9 -k, nginx returned Content-Encoding: gzip and the 318 bytes of the pre-compressed file, with its location's Cache-Control intact; without Accept-Encoding, it served the 5,670-byte original. Where there is no .gz, it keeps compressing on the fly. The documentation asks for the original and the .gz to share the same modification time5; gzip -k preserves it. The gain is CPU time rather than bytes: the same file compressed on the fly came to 297.

Step 2: browser caching per file type

web.dev's rule: URLs with a version or fingerprint in the name can be cached for a year (max-age=31536000), and those without one, such as HTML, should carry no-cache, meaning "revalidate before use"6. MDN defines immutable as a promise that the response will not change while it is fresh7.

# HTML and everything else: may be stored, but is revalidated
location / {
    add_header Cache-Control "no-cache";
}

# Fingerprinted file names: one year, immutable
# (the expression has braces, so it must be quoted)
location ~* "\.[0-9a-f]{8,}\.(css|js|mjs|woff2|webp|avif|png|jpe?g|svg)$" {
    add_header Cache-Control "max-age=31536000, immutable";
}

# CSS and JS without a fingerprint: one week
location ~* \.(css|js|mjs)$ {
    expires 7d;
}

# Images and fonts without a fingerprint: one month
location ~* \.(webp|avif|png|jpe?g|gif|svg|woff2)$ {
    expires 30d;
}
Request
/
Cache-Control returned
no-cache
Request
/assets/styles.css
Cache-Control returned
max-age=604800 plus Expires
Request
/assets/app.3f9a1c2b.js
Cache-Control returned
max-age=31536000, immutable
Request
/assets/photo.webp
Cache-Control returned
max-age=2592000 plus Expires

Three things we learnt in testing:

  1. Braces break the configuration unless quoted. Our first version, with the expression unquoted, made nginx -t fail with unknown directive "8,}\.(css|js...: nginx read the brace as the start of a block.
  2. Order matters. nginx uses the first regular expression that matches, in the order they appear, which is why the fingerprint rule sits above the generic CSS and JS one.
  3. expires sets both Expires and Cache-Control: max-age8, and does not count as an add_header for inheritance: the location using expires kept receiving the server block's security headers.

And the usual trap: an add_header Cache-Control in a location wipes out the security headers you set in the server block, because add_header is only inherited when the current level has none8. Since nginx 1.29.3, one line in the server block is enough:

add_header_inherit merge;

With it, on both versions tested, /, the fingerprinted file and the CSS each carried their Cache-Control plus X-Content-Type-Options and Referrer-Policy. Without it, the two locations using add_header lost them.

Step 3: AI bots, keeping search and training apart

Search crawlers decide whether ChatGPT, Claude or Perplexity can cite you: OAI-SearchBot (block it and you will not appear in ChatGPT search answers9), Claude-SearchBot and Claude-User (blocking them may reduce your visibility in Claude10) and PerplexityBot, which is not used to train models11. Training crawlers such as GPTBot or ClaudeBot only decide whether your content is used for training; blocking them is an editorial choice that does not remove you from answers910.

This block, in the http context (the only place map is allowed12), keeps out three training crawlers plus CCBot, Common Crawl's, and always leaves robots.txt open:

map $http_user_agent $training_bot {
    default 0;
    "~*(GPTBot|ClaudeBot|CCBot|meta-externalagent)" 1;
}

map "$training_bot:$uri" $block_bot {
    default          0;
    "1:/robots.txt"  0;
    "~^1:"           1;
}

And in the server block:

if ($block_bot) {
    return 403;
}

The second map works because nginx checks exact strings first and regular expressions afterwards12: 1:/robots.txt beats ~^1:. Result on both versions: the four blocked bots got 403 on /blog/ and 200 on /robots.txt; OAI-SearchBot, ChatGPT-User, Claude-SearchBot, PerplexityBot, Googlebot and an ordinary Chrome got 200 on both. Anthropic explains that blocking its bots in a way that stops them reading robots.txt may not guarantee an opt-out10: hence the exception.

The limits of any user-agent block:

  • It can be spoofed. Common Crawl says it knows of crawlers pretending to be CCBot and publishes its IP ranges so you can check13; OpenAI and Perplexity publish theirs too.
  • Google-Extended is not a user-agent. It is a robots.txt token with no user-agent of its own, and it does not affect Google Search14. map never sees it: it belongs in robots.txt.
  • ChatGPT-User and Perplexity-User may not follow robots.txt, according to their owners911. Block them at the server and you drop out of the requests users make.

For bots that honour robots.txt, robots.txt is the clearest route and the one anyone can read; the map is the backstop:

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: meta-externalagent
User-agent: Google-Extended
Disallow: /

User-agent: *
Allow: /

How to check it

  1. Syntax. nginx -t before every reload.
  2. Compression. curl -sI -H "Accept-Encoding: gzip" https://www.example.com/ should show Content-Encoding: gzip and Vary: Accept-Encoding. Repeat with a .js and a .css file.
  3. Caching and headers. curl -sI the home page, a CSS file and a fingerprinted file: compare each Cache-Control with the table and check the security headers are still there.
  4. Bots. curl -sI -A "GPTBot/1.4" https://www.example.com/ should give 403, and with -A "OAI-SearchBot/1.4", 200.

What SmoothSeen does with this

Within its SEO analysis, SmoothSeen checks whether your site serves text compressed with gzip or Brotli. In its AI search analysis it reads your robots.txt keeping search crawlers apart from training crawlers, without counting a training block as a failure, and requests the page with each AI crawler's user-agent, a browser's and Googlebot's to spot a server answering 403 to any of them.

What to do this week

Request a .js file from your site with curl -sI -H "Accept-Encoding: gzip": if it is not compressed, check gzip_types. Then request the home page with -A "OAI-SearchBot/1.4" and confirm it gets a 200. To review compression, robots.txt and server-level blocks in one go, analyse your site with SmoothSeen.

Frequently asked questions

Why is nginx not compressing my JavaScript files?

Usually because the MIME type they go out with is not in gzip_types. nginx's mime.types uses application/javascript, but if nginx proxies another application it may receive text/javascript. Look at the Content-Type header with curl -sI and add exactly that value to the list. Also check the file is larger than gzip_min_length.

Do I have to recompile nginx to use Brotli?

With the nginx.org packages, yes: you compile the ngx_brotli module against your exact nginx version, load it with load_module and repeat the process on every upgrade. Some distributions package it for their own nginx builds, and NGINX Plus offers it as a package. If none of that fits, gzip is still a big improvement.

Does blocking GPTBot remove me from ChatGPT?

No. OpenAI separates GPTBot, which it uses to train models, from OAI-SearchBot, which decides whether your site appears in ChatGPT search, and each one is configured separately. You can opt out of training and still appear in answers that use search. What does remove you is blocking OAI-SearchBot.

Why are my CSS changes not showing after enabling caching?

Because the browser uses its stored copy until it expires. It happens when you give a year, or immutable, to a file whose name does not change when its content does. The fix is to have your build or publishing tool add a fingerprint or version number to the file name, or to limit those files to a week with expires 7d.

Sources

  1. 1Module ngx_http_gzip_module, nginx.org, accessed 7 October 2026.
  2. 2nginx: Linux packages, nginx.org, accessed 7 October 2026.
  3. 3google/ngx_brotli, Google, GitHub, accessed 7 October 2026.
  4. 4Brotli (dynamic module), F5 NGINX, accessed 7 October 2026.
  5. 5Module ngx_http_gzip_static_module, nginx.org, accessed 7 October 2026.
  6. 6Prevent unnecessary network requests with the HTTP Cache, web.dev (Google), accessed 7 October 2026.
  7. 7Cache-Control, MDN Web Docs, updated 17 September 2026.
  8. 8Module ngx_http_headers_module, nginx.org, accessed 7 October 2026.
  9. 9Overview of OpenAI Crawlers, OpenAI, accessed 7 October 2026.
  10. 10Does Anthropic crawl data from the web, and how can site owners block the crawler?, Anthropic, updated 7 April 2026.
  11. 11Perplexity Crawlers, Perplexity, accessed 7 October 2026.
  12. 12Module ngx_http_map_module, nginx.org, accessed 7 October 2026.
  13. 13CCBot, Common Crawl, accessed 7 October 2026.
  14. 14Google's common crawlers: Google-Extended, Google Search Central, updated 14 July 2026.

How to cite this article

SmoothSeen. (2026, October 7). Nginx gzip and Brotli: compression, browser caching and blocking AI bots. https://smoothseen.com/en/blog/nginx-gzip-brotli-caching-bots/

Who writes this

SmoothSeen is a website audit tool that measures visibility in search engines and AI assistants and delivers reports under the agency's own brand.

This blog belongs to SmoothSeen: when an article discusses the product, it does so knowing the product is ours. Third-party figures link to their original source.

Change history

  • First version.

Keep reading

Nginx gzip, Brotli, browser caching and blocking AI bots