Nginx as a Production Reverse Proxy: TLS, HTTP/2, Compression and Rate Limiting
Almost every Nginx configuration in production started as a snippet someone pasted from a blog post in 2015. It works, so nobody revisits it — which is how servers end up in 2026 still offering TLS 1.0, still not sending HSTS, and still buffering nothing so that one slow mobile client can occupy an application worker for ninety seconds.
This is a walkthrough of what each part of a modern reverse proxy configuration is actually doing, and which defaults are wrong for a proxied application. It assumes Nginx sits in front of an app server — Gunicorn, Puma, Node, anything — on the same host or a private network.
The shape of the configuration
Nginx configuration is hierarchical: directives set in the http block are inherited by every server block, and directives in a server block are inherited by every location inside it. Inheritance is by replacement, not merging, which is the source of a classic bug — the moment you add a single add_header inside a location, every add_header inherited from the parent block stops applying in that location.
add_header does not accumulate across levels. If you set security headers at the server level and then add one cache header inside a location block, the security headers silently vanish for that location. Use include for a shared headers file, or repeat them.
Put shared settings in /etc/nginx/conf.d/*.conf, which the default nginx.conf includes into the http block, and per-site settings in sites-available with a symlink from sites-enabled.
TLS that scores well and stays fast
The modern configuration is short. Offer TLS 1.2 and 1.3 only, let the client pick among a small set of AEAD ciphers, and enable session resumption so returning visitors skip a full handshake.
# /etc/nginx/conf.d/tls.conf
ssl_protocols TLSv1.2 TLSv1.3;
ssl_prefer_server_ciphers off;
ssl_ciphers ECDHE-ECDSA-AES128-GCM-SHA256:ECDHE-RSA-AES128-GCM-SHA256:ECDHE-ECDSA-AES256-GCM-SHA384:ECDHE-RSA-AES256-GCM-SHA384:ECDHE-ECDSA-CHACHA20-POLY1305:ECDHE-RSA-CHACHA20-POLY1305;
ssl_session_cache shared:SSL:10m;
ssl_session_timeout 1d;
ssl_session_tickets off;
ssl_stapling on;
ssl_stapling_verify on;
resolver 1.1.1.1 8.8.8.8 valid=300s;
resolver_timeout 5s;Two of these are counterintuitive. ssl_prefer_server_ciphers off is correct for TLS 1.3 and modern clients: the client generally knows better than you whether its hardware has AES acceleration, and will pick ChaCha20 on phones that lack it. And ssl_session_tickets off trades a small amount of resumption performance for forward secrecy, because ticket keys are long-lived unless you rotate them yourself.
| Directive | Effect | Cost of getting it wrong |
|---|---|---|
| ssl_protocols | Which TLS versions are offered | TLS 1.0/1.1 fail PCI and modern audits |
| ssl_session_cache | Server-side resumption for repeat visitors | Full handshake on every connection |
| ssl_stapling | Server fetches OCSP response for the client | Extra round trip to the CA per visitor |
| resolver | Required for OCSP stapling to work at all | Stapling silently does nothing |
That last row catches people constantly: ssl_stapling on without a resolver directive produces no error and no stapled response. Check it with openssl s_client -connect example.com:443 -status and look for a non-empty OCSP response section.
HTTP/2, and the syntax that changed
In Nginx 1.25.1 the way you enable HTTP/2 changed. The old form appended http2 to the listen directive; the new form is a separate directive. The old syntax still works but is deprecated and logs a warning, and mixing the two across server blocks on the same port produces genuinely confusing behaviour.
# Deprecated since 1.25.1
listen 443 ssl http2;
# Current
listen 443 ssl;
http2 on;HTTP/2 multiplexes many requests over one connection, which removes head-of-line blocking at the HTTP layer and makes the old practice of sharding assets across domains actively harmful. If your frontend still does domain sharding or aggressive file concatenation for HTTP/1.1, HTTP/2 is a good moment to stop.
Compression: gzip, and Brotli if you have it
Compression is the cheapest performance win available, and Nginx ships with it off for everything except text/html. The default gzip_types is one line and it is almost certainly not what you want.
gzip on;
gzip_vary on;
gzip_min_length 1024;
gzip_comp_level 5;
gzip_proxied any;
gzip_types
text/plain
text/css
text/javascript
application/javascript
application/json
application/xml
application/rss+xml
image/svg+xml
font/woff2;Compression level 5 rather than the maximum 9 is deliberate. Going from 5 to 9 typically buys two or three percent smaller output for roughly double the CPU per response, which on a busy proxy is a bad trade. Note also what is missing from that list: already-compressed formats. Gzipping a JPEG, a PNG, or a woff (not woff2) file wastes CPU and can make the response larger.
gzip_vary on is not optional when a CDN sits in front. Without it, a cache can serve a gzipped response to a client that did not ask for one, or cache the uncompressed variant and serve it to everyone.
Brotli compresses text roughly 15 to 20 percent better than gzip at equivalent CPU, but it is not built into stock Nginx. On Ubuntu, libnginx-mod-brotli provides it as a dynamic module. Enable both — clients that do not advertise br fall back to gzip automatically.
sudo apt install -y libnginx-mod-brotli
# In the http block:
brotli on;
brotli_comp_level 5;
brotli_types text/plain text/css application/javascript application/json image/svg+xml;Proxying, and the buffering settings that matter
The core reason to put Nginx in front of an application server is buffering. Application workers are expensive — a Gunicorn worker is a whole Python process — and a client on a slow mobile connection uploading a photo would otherwise hold one for the entire upload. Nginx absorbs that: it reads the request at the client's pace, and only hands a complete request to the app.
location / {
proxy_pass http://app_upstream;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_http_version 1.1;
proxy_set_header Connection "";
proxy_buffering on;
proxy_buffers 8 16k;
proxy_buffer_size 16k;
proxy_busy_buffers_size 32k;
proxy_connect_timeout 5s;
proxy_send_timeout 60s;
proxy_read_timeout 60s;
}proxy_http_version 1.1 with an empty Connection header enables keepalive to the upstream. Without it Nginx opens a fresh TCP connection to your app for every single request, which on a busy service is a startling amount of wasted work. Pair it with a keepalive directive in the upstream block:
upstream app_upstream {
server 127.0.0.1:8000;
keepalive 32;
keepalive_timeout 60s;
}One important exception: buffering must be off for streaming responses. Server-sent events, streamed LLM output, and long-polling endpoints all break in a way that looks like the server hanging, because Nginx is dutifully waiting to fill a buffer before sending anything.
location /events {
proxy_pass http://app_upstream;
proxy_buffering off;
proxy_cache off;
proxy_read_timeout 3600s;
add_header X-Accel-Buffering no;
}Rate limiting
Nginx implements rate limiting with the leaky bucket algorithm. You define a shared memory zone keyed on something — usually client IP — and a rate. Requests above the rate are delayed or rejected. The subtlety is in the burst and nodelay parameters, which almost everyone gets wrong on the first attempt.
# In the http block:
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
limit_req_status 429;
# In a location:
location /api/ {
limit_req zone=api burst=20 nodelay;
proxy_pass http://app_upstream;
}Read that as: sustained 10 requests per second, with a bucket that tolerates a burst of 20 above the rate. Without nodelay, burst requests are queued and served slowly, which turns a spike into high latency for legitimate users. With nodelay, burst requests are served immediately and only requests beyond the bucket are rejected — which is what you almost always want for an API.
| Configuration | Behaviour above the rate |
|---|---|
| limit_req zone=api; | Reject immediately — no tolerance for any burst |
| burst=20 | Queue up to 20, serve them slowly at the configured rate |
| burst=20 nodelay | Serve up to 20 immediately, reject beyond that |
| burst=20 delay=10 | Serve 10 immediately, delay the next 10, reject beyond |
$binary_remote_addr rather than $remote_addr is a memory optimisation that matters at scale: it stores 4 bytes for IPv4 instead of a string, so a 10 MB zone tracks roughly 160,000 addresses. And if Nginx is behind a load balancer or CDN, $remote_addr is the proxy, not the user — key on a validated forwarded address instead, or you will rate limit the entire internet as one client.
Test rate limits before shipping them. limit_req_status 429 is worth setting explicitly, because the Nginx default for a rejected request is 503, which clients and monitoring will interpret as your server being broken rather than as deliberate throttling.
Security headers
These belong at the proxy, not in application code, because they then apply uniformly to every response including static files and error pages.
# /etc/nginx/snippets/security-headers.conf
add_header Strict-Transport-Security "max-age=31536000; includeSubDomains" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-Frame-Options "SAMEORIGIN" always;
add_header Referrer-Policy "strict-origin-when-cross-origin" always;
add_header Permissions-Policy "geolocation=(), microphone=(), camera=()" always;The always parameter is what makes these apply to error responses too. Without it, a 500 or a 404 goes out with no security headers at all — which is exactly when an attacker is most interested in what your server does.
Be deliberate about HSTS. Once a browser has seen that header it will refuse to speak plain HTTP to your domain for a year, and there is no way to retract it early. Start with a short max-age, confirm every subdomain genuinely has a valid certificate, and only then raise it and consider preloading.
Verifying the configuration
sudo nginx -t
sudo nginx -T | less # full effective config, all includes resolved
sudo systemctl reload nginx
curl -sI https://example.com | grep -i -E 'strict-transport|content-type-options|content-encoding'
curl -s -o /dev/null -w '%{http_version} %{time_total}s\n' https://example.comnginx -T is the one to remember. It prints the entire configuration with every include expanded, which is how you find out that the setting you thought you changed is being overridden by a file in conf.d you forgot existed.
Should I use Nginx or Caddy for a new deployment?
Caddy obtains and renews certificates automatically and has a far shorter configuration file, which makes it excellent for straightforward sites. Nginx wins when you need fine-grained control — complex rate limiting, request buffering tuned per location, mature module ecosystem — and when you or your team already know it. Both are correct answers.
Does proxy_buffering off make my API faster?
No, and it usually makes the whole server slower under load. Turning buffering off means an application worker stays occupied until the slowest client has received every byte. Turn it off only for endpoints that genuinely stream, and leave it on everywhere else.
Why do my security headers disappear on some pages?
Because a location block on those pages contains its own add_header directive. add_header inherits by replacement: any add_header at a deeper level discards every header set at levels above it. Move them into a snippet and include it in each location that sets headers of its own.
How do I rate limit correctly behind CloudFront or an ALB?
Key the limit on the real client address rather than $remote_addr, which will be the proxy. Use the realip module with set_real_ip_from listing your load balancer ranges and real_ip_header X-Forwarded-For, so $binary_remote_addr resolves to the actual client before the limit is applied.
Is HTTP/3 worth enabling?
It helps most on lossy mobile networks, where QUIC recovers from packet loss without stalling every stream. Nginx supports it from 1.25 with listen 443 quic and an Alt-Svc header. It is a reasonable addition once HTTP/2, TLS and compression are already correct, and a distraction before that.