How to Use the Cache API in Cloudflare Workers

This guide implements Edge Compute & Dynamic Caching on one platform, within Advanced Caching Strategies & CDN Architecture. Cloudflare Workers sit in front of the CDN cache and can control it in two ways: implicitly, by passing cf options to fetch() (cache everything, override TTLs, set a custom cache key), and explicitly, through the Service-Worker-style Cache API — caches.default.match() and put() — which lets the Worker decide exactly what to store, under which key, for how long.

The explicit Cache API is what makes advanced patterns possible: caching responses from APIs that send no cache headers, caching POST-backed search results under a GET key, implementing stale-while-revalidate yourself, or storing transformed responses (personalised-at-the-edge variants, rewritten HTML). It also has behaviours that surprise people — it is local to a data centre, it respects certain headers on put, and it does not work on workers.dev domains — that are worth knowing before you rely on it.

Worker cache flow Sequence of a Worker checking the local edge cache, fetching from origin on a miss, returning the response and storing it in the background. Worker cache flow Client Worker Edge cache Origin GET /api/products match(key) miss fetch put via waitUntil

Rapid Diagnosis

  • Check whether responses are being cached at all. Look at cf-cache-status on responses (HIT, MISS, DYNAMIC, BYPASS). DYNAMIC means the response was not eligible for the cache under default rules.
  • Check where the Worker runs. The Cache API is unavailable on *.workers.dev; test on a custom domain route.
  • Check response headers on put. Responses with Set-Cookie, Cache-Control: private or no-store are not stored by cache.put.
  • Check the key. Query strings, headers and URL variations produce different keys; inconsistent keys look like a low hit rate.

Root Cause Analysis

1. Default rules skip HTML and JSON. Without configuration, Cloudflare caches static file extensions but not HTML or API responses, which show as DYNAMIC.

2. Uncacheable origin headers. Origins often send private, no-cache or cookies; put respects some of these and silently stores nothing.

3. Per-data-centre caches. caches.default is local to the data centre handling the request; a new location starts cold. Hit rates look lower than with tiered caching.

4. Key fragmentation. Tracking parameters, varying Accept headers and inconsistent trailing slashes split the cache.

Two ways to cache from a Worker Comparison of caching through fetch with cf options and explicit caching with the Cache API. Two ways to cache from a Worker Aspect fetch() with cf options caches.default API Effort one option explicit code Custom cache key cacheKey option any Request as key Cache transformed responses no yes Tiered cache / global reach yes local data centre Purge by tag with Cache-Tag purge by URL/key

Step-by-Step Resolution

1. Use cf options for simple cases

javascript
export default {
  async fetch(request) {
    return fetch(request, {
      cf: { cacheEverything: true, cacheTtlByStatus: { '200-299': 300, '404': 30, '500-599': 0 } },
    });
  },
};
// trade-off: cacheEverything caches HTML regardless of origin headers. Only use
// it on routes you know are not personalised, or you will cache private pages.

Expected outcome: HTML or API responses on the route become cacheable at the edge with TTLs you control.

2. Use the Cache API with an explicit key for custom logic

javascript
export default {
  async fetch(request, env, ctx) {
    const url = new URL(request.url);
    url.searchParams.sort();
    ['utm_source', 'utm_medium', 'gclid', 'fbclid'].forEach((p) => url.searchParams.delete(p));
    const key = new Request(url.toString(), { method: 'GET' });
    const cache = caches.default;
    let res = await cache.match(key);
    if (res) return res;
    res = await fetch(request);
    if (res.ok) {
      res = new Response(res.body, res);
      res.headers.set('Cache-Control', 'public, max-age=300');
      res.headers.delete('Set-Cookie');
      ctx.waitUntil(cache.put(key, res.clone()));
    }
    return res;
  },
};
// trade-off: deleting Set-Cookie is correct only for responses that are truly
// shared. If the origin sets a session cookie here, this route should not be
// cached at all.

Expected outcome: normalised keys, predictable TTLs and background storage that never delays the response.

3. Implement stale-while-revalidate in the Worker

Store a timestamp header with the cached response; when the entry is older than a soft TTL but younger than a hard TTL, return it immediately and refresh it in waitUntil.

javascript
const SOFT = 60_000, HARD = 86_400_000;
const cached = await cache.match(key);
if (cached) {
  const age = Date.now() - Number(cached.headers.get('x-stored-at') ?? 0);
  if (age < HARD) {
    if (age > SOFT) ctx.waitUntil(refresh(key, request, cache));
    return cached;
  }
}
// trade-off: hand-rolled SWR is per data centre and can trigger several
// concurrent refreshes under load. Add a short lock (e.g. a KV flag) for
// expensive origins.

4. Purge correctly

Entries stored with cache.put can be purged by URL through the Cloudflare API; tag-based purge applies to responses cached by the CDN through fetch with Cache-Tag headers. Choose the caching path with your purge strategy in mind.

Hit ratio for an API route by caching approach Bar chart of edge cache hit ratio for a product API under different Worker caching approaches. Hit ratio for an API route by caching approach No Worker caching (DYNAMIC) 0% Cache API, raw URL keys 41% Cache API, normalised keys 73% cf options + tiered cache 88%

Verification

Check cf-cache-status (for fetch-based caching) or add your own x-worker-cache: hit|miss header for Cache API paths. Request the same URL with and without tracking parameters and confirm the second is a hit. In RUM, carry the status via Server-Timing and watch TTFB p75 per status. Load-test a cold route to make sure origin traffic stays bounded.

Worked Example: Caching a Search API

A storefront's search API was called on every keystroke debounce and every results page, all DYNAMIC. A Worker normalised the query (lowercased, trimmed, sorted parameters, removed tracking parameters), cached results for 120 seconds with the Cache API and served stale results for up to ten minutes while refreshing in the background. The hit ratio reached 64% at peak, median API latency from the browser fell from 210ms to 35ms on hits, and the origin's search cluster load dropped by over half. Typeahead INP improved too, because results arrived before users typed the next character.

Common Mistakes

  • Testing on workers.dev. The Cache API is a no-op there.
  • Awaiting cache.put. Use ctx.waitUntil so storage never delays the response.
  • Caching responses with cookies. put will refuse, or worse, if you strip headers, you may cache personal data.
  • Forgetting that the Cache API is local. Global hit rates need tiered caching or fetch-based caching.

Edge Cases

Vary headers. The Cache API matches on URL and respects Vary in limited ways; avoid relying on it and encode variants in the key URL instead.

Large responses. There are size limits for cached objects; very large files should be cached by the CDN directly or stored in object storage.

POST requests. The Cache API only stores GET; map POST-backed queries to a synthetic GET key when caching them is safe.

Cache Reserve and tiered caching. These platform features increase hit rates for fetch-based caching; they do not apply to caches.default entries the same way.

FAQ

Is caches.default the same as the browser Cache API?

It shares the interface (match, put, delete) but stores entries in Cloudflare's edge cache for that data centre, not in the browser. Named caches via caches.open() are separate namespaces in the same edge cache.

Should I use KV instead of the Cache API?

KV is a globally replicated key-value store with eventual consistency, good for configuration and data you write deliberately. The Cache API is a per-location HTTP cache for responses. For caching origin responses, prefer the cache; use KV for state like locks or small lookups.

Do cached responses count towards Worker CPU time?

Serving from caches.default still runs the Worker, but cache lookups are I/O rather than CPU and are cheap. For assets that need no logic, avoid routing them through the Worker at all.

How do I purge entries stored with cache.put?

Purge by URL through the API or dashboard, using the same URL as the key. Tag-based purges target responses cached via fetch with Cache-Tag headers.

Can I cache personalised HTML with the Cache API?

Only per variant with a safe key, never per user, and only for content that is genuinely shared within the variant — see caching HTML for logged-in users safely.