How to Implement Stale-While-Revalidate with nginx proxy_cache
This guide implements Stale-While-Revalidate Implementation at the origin, within Advanced Caching Strategies & CDN Architecture. Not every stack sits behind a CDN that caches HTML, and even those that do benefit from a cache in front of the application servers: it absorbs CDN misses, revalidations from many edge locations, and traffic that bypasses the CDN. nginx's proxy_cache provides exactly that, with three features that together implement stale-while-revalidate: proxy_cache_use_stale updating, proxy_cache_background_update, and proxy_cache_lock.
Configured correctly, nginx returns a cached (possibly stale) response immediately, refreshes it with a single background request, collapses concurrent misses into one upstream fetch, and keeps serving stale content when the application is down. Application servers see a small, steady stream of refreshes instead of every request, and TTFB for cached pages falls to a few milliseconds of nginx processing.
Rapid Diagnosis
- Check whether nginx caches at all. Look for
proxy_cachedirectives; addadd_header X-Cache-Status $upstream_cache_status;to see HIT, MISS, EXPIRED, STALE, UPDATING. - Watch upstream request rates during traffic spikes. If they track client request rates, there is no effective cache.
- Check for thundering herds. After expiry, do many simultaneous requests reach the app? Without
proxy_cache_lockthey will. - Check failure behaviour. Stop the app in staging; do cached pages still serve?
Root Cause Analysis
1. No caching for HTML. nginx proxies everything to the app by default.
2. Blocking revalidation. Without background update, the request that finds an expired entry waits for the upstream.
3. Concurrent misses. Without a lock, every request arriving during a miss goes upstream.
4. Cookies and headers preventing caching. nginx does not cache responses with Set-Cookie by default, and respects upstream Cache-Control: private/no-store.
Step-by-Step Resolution
1. Define a cache zone and enable caching for public routes
proxy_cache_path /var/cache/nginx/html levels=1:2 keys_zone=html:50m max_size=2g inactive=7d use_temp_path=off;
server {
location / {
proxy_pass http://app;
proxy_cache html;
proxy_cache_key "$scheme$host$uri$is_args$args";
proxy_cache_valid 200 301 10m;
proxy_cache_valid 404 1m;
add_header X-Cache-Status $upstream_cache_status always;
# trade-off: proxy_cache_valid only applies when the upstream sends no
# Cache-Control/Expires; if the app sends headers, nginx follows them unless
# proxy_ignore_headers is set. Decide which side owns the TTL.
}
}
2. Serve stale while updating, in the background
proxy_cache_use_stale error timeout updating http_500 http_502 http_503 http_504;
proxy_cache_background_update on;
proxy_cache_lock on;
proxy_cache_lock_timeout 5s;
# trade-off: 'updating' serves stale indefinitely while an update is in flight;
# if the upstream hangs, clients keep receiving the old version until the
# update completes or times out. Keep upstream timeouts sensible.
Expected outcome: expired entries are served instantly while one request refreshes them; concurrent misses wait for a single upstream fetch.
3. Honour stale-while-revalidate from the upstream
nginx understands stale-while-revalidate and stale-if-error extensions in upstream Cache-Control headers (since 1.11.10), so the application can set windows per response while nginx enforces them.
4. Bypass the cache for personal requests
map $http_cookie $skip_cache { default 0; ~*session= 1; }
proxy_cache_bypass $skip_cache;
proxy_no_cache $skip_cache;
# trade-off: bypassing on a session cookie keeps logged-in traffic uncached.
# Move to a shared-shell architecture to cache those users too.
Verification
Send repeated requests and watch X-Cache-Status: MISS, then HIT, then (after expiry) STALE or UPDATING followed by HIT. Load-test a page while its entry expires and confirm upstream logs show a single request. Stop the application and confirm cached pages still serve with STALE status. In RUM, TTFB p75 for cached routes should drop to near network latency.
Worked Example: A Self-Hosted CMS
A self-hosted CMS rendered pages in 300–600ms and had no CDN HTML caching. A campaign drove a 15x traffic spike that overwhelmed the application servers, with TTFB p95 exceeding 8 seconds. Adding nginx proxy_cache with a 5-minute validity, use_stale updating, background updates and locking (bypassing for editor sessions) reduced upstream traffic during the next spike to a few requests per second per page. TTFB p75 fell to 25ms for cached pages, and a later application crash went unnoticed by visitors for its 6-minute duration.
Purging nginx Caches
Open-source nginx has no built-in purge endpoint (the commercial version and third-party modules do). Common approaches: short validity combined with SWR so changes propagate within minutes; a cache key that includes a content version bumped on publish; or deleting cache files for a key via a small purge script. For many content sites, versioned keys plus SWR are the simplest reliable option.
Monitoring the Proxy Cache
Expose $upstream_cache_status in access logs alongside request time and upstream response time, then chart the share of HIT, STALE, UPDATING, MISS and BYPASS per route. A healthy configuration shows mostly HIT and STALE with occasional UPDATING; a high MISS share points to keys that are too specific or validity that is too short, and a high BYPASS share to cookies or headers excluding requests. Alert when upstream response time rises while STALE serving increases — that pattern means the cache is hiding an application problem that will surface as soon as stale windows run out.
Common Mistakes
- Caching responses that set cookies. nginx refuses by default; overriding that with
proxy_ignore_headers Set-Cookiecan leak sessions. - Missing
alwayson add_header. Without it, status headers are not added to some responses. - Too small a keys_zone. Each key consumes memory; a full zone evicts entries early.
- Cache keys without host. Multi-site servers can serve one site's pages for another.
Edge Cases
Vary. nginx caches variants based on Vary headers from the upstream; Vary: Cookie effectively disables sharing.
Large responses. Buffering and temp paths matter for big files; use_temp_path=off avoids extra copies.
Microcaching. Caching for one second with use_stale updating is a classic trick that absorbs bursts for highly dynamic pages while keeping content nearly live.
Behind a CDN. nginx then acts as a shield; align its TTLs with the CDN's so both do not revalidate at the same moments.
FAQ
Does proxy_cache_background_update need use_stale updating?
Yes — background updates only happen when nginx is allowed to serve the stale entry while updating, which proxy_cache_use_stale updating enables.
Is microcaching safe for dynamic pages?
For public pages, a one-second cache rarely shows meaningfully stale content and collapses bursts dramatically. Never apply it to personal responses.
Where should the cache live on disk?
On fast local storage (SSD or tmpfs for small caches). Network filesystems add latency and locking issues.
Can nginx revalidate with ETags?
With proxy_cache_revalidate on, nginx uses If-Modified-Since/If-None-Match when refreshing expired entries, so unchanged content returns a cheap 304 from the app.
How does this interact with a CDN's SWR?
They layer: the CDN serves stale from the edge while revalidating against nginx, and nginx serves stale while revalidating against the app. Each layer reduces load on the next.
What status codes should be cached?
200 and 301 for normal validity, 404 briefly, never 5xx as fresh content — but allow serving stale on 5xx via proxy_cache_use_stale.
Related
- Choosing SWR windows for HTML and APIs — choosing the values nginx enforces.
- Tiered caching and origin shield — the CDN-side equivalent of this shield.
- Fixing TTFB from serverless cold starts — another origin latency problem caching hides.