How to Use the Cache API in Cloudflare Workers
This guide implements Edge Compute & Dynamic Caching on one platform, within Advanced Caching Strategies & CDN Architecture. Cloudflare Workers sit in front of the CDN cache and can control it in two ways: implicitly, by passing cf options to fetch() (cache everything, override TTLs, set a custom cache key), and explicitly, through the Service-Worker-style Cache API — caches.default.match() and put() — which lets the Worker decide exactly what to store, under which key, for how long.
The explicit Cache API is what makes advanced patterns possible: caching responses from APIs that send no cache headers, caching POST-backed search results under a GET key, implementing stale-while-revalidate yourself, or storing transformed responses (personalised-at-the-edge variants, rewritten HTML). It also has behaviours that surprise people — it is local to a data centre, it respects certain headers on put, and it does not work on workers.dev domains — that are worth knowing before you rely on it.
Rapid Diagnosis
- Check whether responses are being cached at all. Look at
cf-cache-statuson responses (HIT, MISS, DYNAMIC, BYPASS).DYNAMICmeans the response was not eligible for the cache under default rules. - Check where the Worker runs. The Cache API is unavailable on
*.workers.dev; test on a custom domain route. - Check response headers on
put. Responses withSet-Cookie,Cache-Control: privateorno-storeare not stored bycache.put. - Check the key. Query strings, headers and URL variations produce different keys; inconsistent keys look like a low hit rate.
Root Cause Analysis
1. Default rules skip HTML and JSON. Without configuration, Cloudflare caches static file extensions but not HTML or API responses, which show as DYNAMIC.
2. Uncacheable origin headers. Origins often send private, no-cache or cookies; put respects some of these and silently stores nothing.
3. Per-data-centre caches. caches.default is local to the data centre handling the request; a new location starts cold. Hit rates look lower than with tiered caching.
4. Key fragmentation. Tracking parameters, varying Accept headers and inconsistent trailing slashes split the cache.
Step-by-Step Resolution
1. Use cf options for simple cases
export default {
async fetch(request) {
return fetch(request, {
cf: { cacheEverything: true, cacheTtlByStatus: { '200-299': 300, '404': 30, '500-599': 0 } },
});
},
};
// trade-off: cacheEverything caches HTML regardless of origin headers. Only use
// it on routes you know are not personalised, or you will cache private pages.
Expected outcome: HTML or API responses on the route become cacheable at the edge with TTLs you control.
2. Use the Cache API with an explicit key for custom logic
export default {
async fetch(request, env, ctx) {
const url = new URL(request.url);
url.searchParams.sort();
['utm_source', 'utm_medium', 'gclid', 'fbclid'].forEach((p) => url.searchParams.delete(p));
const key = new Request(url.toString(), { method: 'GET' });
const cache = caches.default;
let res = await cache.match(key);
if (res) return res;
res = await fetch(request);
if (res.ok) {
res = new Response(res.body, res);
res.headers.set('Cache-Control', 'public, max-age=300');
res.headers.delete('Set-Cookie');
ctx.waitUntil(cache.put(key, res.clone()));
}
return res;
},
};
// trade-off: deleting Set-Cookie is correct only for responses that are truly
// shared. If the origin sets a session cookie here, this route should not be
// cached at all.
Expected outcome: normalised keys, predictable TTLs and background storage that never delays the response.
3. Implement stale-while-revalidate in the Worker
Store a timestamp header with the cached response; when the entry is older than a soft TTL but younger than a hard TTL, return it immediately and refresh it in waitUntil.
const SOFT = 60_000, HARD = 86_400_000;
const cached = await cache.match(key);
if (cached) {
const age = Date.now() - Number(cached.headers.get('x-stored-at') ?? 0);
if (age < HARD) {
if (age > SOFT) ctx.waitUntil(refresh(key, request, cache));
return cached;
}
}
// trade-off: hand-rolled SWR is per data centre and can trigger several
// concurrent refreshes under load. Add a short lock (e.g. a KV flag) for
// expensive origins.
4. Purge correctly
Entries stored with cache.put can be purged by URL through the Cloudflare API; tag-based purge applies to responses cached by the CDN through fetch with Cache-Tag headers. Choose the caching path with your purge strategy in mind.
Verification
Check cf-cache-status (for fetch-based caching) or add your own x-worker-cache: hit|miss header for Cache API paths. Request the same URL with and without tracking parameters and confirm the second is a hit. In RUM, carry the status via Server-Timing and watch TTFB p75 per status. Load-test a cold route to make sure origin traffic stays bounded.
Worked Example: Caching a Search API
A storefront's search API was called on every keystroke debounce and every results page, all DYNAMIC. A Worker normalised the query (lowercased, trimmed, sorted parameters, removed tracking parameters), cached results for 120 seconds with the Cache API and served stale results for up to ten minutes while refreshing in the background. The hit ratio reached 64% at peak, median API latency from the browser fell from 210ms to 35ms on hits, and the origin's search cluster load dropped by over half. Typeahead INP improved too, because results arrived before users typed the next character.
Common Mistakes
- Testing on workers.dev. The Cache API is a no-op there.
- Awaiting
cache.put. Usectx.waitUntilso storage never delays the response. - Caching responses with cookies.
putwill refuse, or worse, if you strip headers, you may cache personal data. - Forgetting that the Cache API is local. Global hit rates need tiered caching or
fetch-based caching.
Edge Cases
Vary headers. The Cache API matches on URL and respects Vary in limited ways; avoid relying on it and encode variants in the key URL instead.
Large responses. There are size limits for cached objects; very large files should be cached by the CDN directly or stored in object storage.
POST requests. The Cache API only stores GET; map POST-backed queries to a synthetic GET key when caching them is safe.
Cache Reserve and tiered caching. These platform features increase hit rates for fetch-based caching; they do not apply to caches.default entries the same way.
FAQ
Is caches.default the same as the browser Cache API?
It shares the interface (match, put, delete) but stores entries in Cloudflare's edge cache for that data centre, not in the browser. Named caches via caches.open() are separate namespaces in the same edge cache.
Should I use KV instead of the Cache API?
KV is a globally replicated key-value store with eventual consistency, good for configuration and data you write deliberately. The Cache API is a per-location HTTP cache for responses. For caching origin responses, prefer the cache; use KV for state like locks or small lookups.
Do cached responses count towards Worker CPU time?
Serving from caches.default still runs the Worker, but cache lookups are I/O rather than CPU and are cheap. For assets that need no logic, avoid routing them through the Worker at all.
How do I purge entries stored with cache.put?
Purge by URL through the API or dashboard, using the same URL as the key. Tag-based purges target responses cached via fetch with Cache-Tag headers.
Can I cache personalised HTML with the Cache API?
Only per variant with a safe key, never per user, and only for content that is genuinely shared within the variant — see caching HTML for logged-in users safely.
Related
- Stale-while-revalidate implementation — the SWR model behind step 3.
- Tiered caching and origin shield — raising hit rates beyond one location.
- Caching API responses at the CDN — header-driven API caching.