How to Cache API Responses at the CDN

This guide extends CDN Edge Caching Configuration to JSON APIs, within Advanced Caching Strategies & CDN Architecture. Modern frontends spend much of their loading time waiting for data: product details, search results, navigation menus, CMS content, configuration. Those requests usually go straight to the origin with Cache-Control: no-cache or no caching headers at all, even when the response is the same for every user and changes only occasionally.

Caching public GET responses at the CDN cuts their latency from origin round-trip time (often 150–500ms) to edge latency (20–60ms), which shortens client-rendered LCP, speeds up route transitions in single-page apps, and protects the origin during spikes. The challenges are deciding what is safe to cache, choosing TTLs, keeping keys tidy, and purging when underlying data changes.

Which API responses can be cached at the edge Classification of common API endpoints by whether they can be cached at the CDN and the recommended header. Which API responses can be cached at the edge Endpoint Same for all users? Recommended policy Product / article detail yes s-maxage + SWR, purge by tag Search results yes (per query) short s-maxage, normalised key Navigation / config yes long s-maxage, purge on publish Cart / account no private, no-store Prices with personal discounts no private; or split base price out

Rapid Diagnosis

  • List API calls on key routes in the Network panel with their timings and cache headers.
  • Check cf-cache-status, x-cache or equivalent on API responses; MISS/DYNAMIC/BYPASS on every request means no edge caching.
  • Classify each endpoint as public or personal, and note how often its data changes.
  • Check request methods. GraphQL APIs often use POST for queries, which CDNs do not cache by default.

Root Cause Analysis

1. Conservative default headers. API frameworks often send no-cache or nothing, which many CDNs treat as uncacheable for JSON.

2. Authentication on public data. Public endpoints that accept (and ignore) auth headers or cookies get treated as personal by CDNs that bypass on Authorization.

3. POST-based queries. GraphQL over POST and RPC-style endpoints cannot be cached without mapping them to GET.

4. No invalidation strategy. Teams avoid caching because they cannot purge reliably when data changes.

Product API latency from the browser (p75) Bar chart of p75 latency for a product detail API served from origin and from the CDN cache. Product API latency from the browser (p75) Origin (no cache) 310ms CDN hit 38ms CDN stale + background refresh 41ms

Step-by-Step Resolution

1. Send explicit shared-cache headers on public endpoints

javascript
// Express: public product endpoint.
app.get('/api/products/:id', async (req, res) => {
  const product = await getProduct(req.params.id);
  res.set('Cache-Control', 'public, max-age=0, s-maxage=300, stale-while-revalidate=3600');
  res.set('Cache-Tag', `product-${product.id}`);              // surrogate key for purging
  res.set('Vary', 'Accept-Encoding');
  res.json(product);
});
// trade-off: max-age=0 keeps browsers revalidating so they never hold stale
// data longer than the CDN; raise it only for data that tolerates staleness.

Expected outcome: the CDN caches the response for five minutes and serves stale copies while refreshing.

2. Ignore auth headers on public routes at the edge

Strip Authorization and cookies from requests to public endpoints before they reach the cache (and the origin), so the CDN does not bypass caching and the origin cannot accidentally personalise.

3. Map cacheable POST queries to GET

For GraphQL, use persisted queries sent as GET (/graphql?id=<hash>&variables=...), which CDNs can cache by URL.

javascript
// Client: persisted query as GET.
const url = `/graphql?id=${QUERY_HASH}&variables=${encodeURIComponent(JSON.stringify(vars))}`;
const data = await fetch(url).then((r) => r.json());
// trade-off: persisted queries need a registry of allowed query hashes on the
// server. That is also a security benefit, but it adds a build step.

4. Purge by tag when data changes

When a product changes, purge product-<id> so every cached response that includes it (detail, listings, search) is refreshed.

Purge-on-write for a cached API Sequence showing an admin update triggering a tag purge, followed by the next request refetching fresh data. Purge-on-write for a cached API Admin Origin CDN Client update product 42 purge tag product-42 GET /api/products/42 miss → fetch fresh

Verification

Request each public endpoint twice and confirm a cache HIT on the second request; request a personal endpoint and confirm it is never cached. Update a record and confirm the next request returns fresh data after the purge. In RUM, measure data request latency (Resource Timing for API URLs) and route transition times; both should fall.

Worked Example: Category Listing API

An e-commerce SPA fetched category listings from /api/categories/:slug/products?page=n&sort=… on every navigation, with origin latency of 280ms at p75. The team added s-maxage=120, stale-while-revalidate=600, normalised the query parameters in the cache key, tagged responses with every product ID and the category, and purged tags from the product-update pipeline. The edge hit ratio for listings reached 89%, data latency p75 fell to 45ms, and route transitions between categories became noticeably snappier; INP was unaffected, but route LCP improved by about 230ms.

Common Mistakes

  • Caching responses that vary by user. Personal discounts, inventory reserved for a user, or localisation from a profile must not be cached publicly.
  • Long browser max-age on APIs. Browsers cannot be purged; keep browser TTLs short or zero and let the CDN hold the data.
  • Too many tags. Some CDNs cap the size of tag headers; tag by entity IDs that actually drive purges.
  • Forgetting CORS. Cached responses must carry the right Access-Control-Allow-Origin; vary on Origin only if you serve different values.

Edge Cases

Personalised fields inside public data. Split them: cache the public product and fetch the user's price or wishlist status separately.

Rate limits and quotas. Cached responses reduce origin calls, but ensure rate limiting happens before the cache only for uncached routes, or legitimate cached traffic may be throttled.

Large responses. Very large JSON payloads cost transfer time even when cached; paginate and compress.

Error responses. Cache 404s briefly (seconds) and never cache 5xx, or an outage gets cached.

FAQ

Can I cache GraphQL at the CDN?

Yes, with GET-based persisted queries and responses that do not depend on the user. Client-side GraphQL caches (Apollo, urql) complement but do not replace edge caching, because they only help the same client.

How short can TTLs be and still help?

Even a few seconds helps under load — it collapses bursts of identical requests into one origin call. For user-facing latency, longer TTLs with stale-while-revalidate give the best results, with purges handling freshness.

Should API caching use the same rules as HTML?

The principles are the same — public vs private, keys, TTLs, purging — but APIs are usually more granular, so tag-based purging and parameter allow-lists matter more.

Does caching APIs help INP?

Indirectly: interactions that trigger data fetches (filters, pagination) complete their visible update sooner when data comes from the edge. INP measures the next paint after input, which may be a loading state, but users perceive the faster result.

What about server-side rendering that calls these APIs?

SSR servers benefit too if they call the API through the CDN (or a shared cache). That turns many page renders' data fetches into cache hits and lowers TTFB.

What about APIs called from other servers, not browsers?

Server-to-server calls (SSR fetching data, microservices) can use the same CDN or an internal shared cache. The headers and purge strategy are the same; the main difference is that you can often allow longer TTLs because the consumer is under your control.

How should I handle ETags on cached APIs?

Keep them. Strong or weak ETags let the CDN revalidate cheaply with the origin when a TTL expires, and let browsers revalidate with the CDN. Make sure ETags are stable across origin instances — ETags derived from server-specific values defeat revalidation.

Is it worth caching APIs with very low traffic?

Usually not for latency — each location rarely reuses the entry — but tiered caching changes that by concentrating requests at a shield. For low-traffic endpoints, focus on making the origin response itself fast and add caching mainly for spike protection.

/html>