How to Cache API Responses at the CDN
This guide extends CDN Edge Caching Configuration to JSON APIs, within Advanced Caching Strategies & CDN Architecture. Modern frontends spend much of their loading time waiting for data: product details, search results, navigation menus, CMS content, configuration. Those requests usually go straight to the origin with Cache-Control: no-cache or no caching headers at all, even when the response is the same for every user and changes only occasionally.
Caching public GET responses at the CDN cuts their latency from origin round-trip time (often 150–500ms) to edge latency (20–60ms), which shortens client-rendered LCP, speeds up route transitions in single-page apps, and protects the origin during spikes. The challenges are deciding what is safe to cache, choosing TTLs, keeping keys tidy, and purging when underlying data changes.
Rapid Diagnosis
- List API calls on key routes in the Network panel with their timings and cache headers.
- Check
cf-cache-status,x-cacheor equivalent on API responses;MISS/DYNAMIC/BYPASSon every request means no edge caching. - Classify each endpoint as public or personal, and note how often its data changes.
- Check request methods. GraphQL APIs often use POST for queries, which CDNs do not cache by default.
Root Cause Analysis
1. Conservative default headers. API frameworks often send no-cache or nothing, which many CDNs treat as uncacheable for JSON.
2. Authentication on public data. Public endpoints that accept (and ignore) auth headers or cookies get treated as personal by CDNs that bypass on Authorization.
3. POST-based queries. GraphQL over POST and RPC-style endpoints cannot be cached without mapping them to GET.
4. No invalidation strategy. Teams avoid caching because they cannot purge reliably when data changes.
Step-by-Step Resolution
1. Send explicit shared-cache headers on public endpoints
// Express: public product endpoint.
app.get('/api/products/:id', async (req, res) => {
const product = await getProduct(req.params.id);
res.set('Cache-Control', 'public, max-age=0, s-maxage=300, stale-while-revalidate=3600');
res.set('Cache-Tag', `product-${product.id}`); // surrogate key for purging
res.set('Vary', 'Accept-Encoding');
res.json(product);
});
// trade-off: max-age=0 keeps browsers revalidating so they never hold stale
// data longer than the CDN; raise it only for data that tolerates staleness.
Expected outcome: the CDN caches the response for five minutes and serves stale copies while refreshing.
2. Ignore auth headers on public routes at the edge
Strip Authorization and cookies from requests to public endpoints before they reach the cache (and the origin), so the CDN does not bypass caching and the origin cannot accidentally personalise.
3. Map cacheable POST queries to GET
For GraphQL, use persisted queries sent as GET (/graphql?id=<hash>&variables=...), which CDNs can cache by URL.
// Client: persisted query as GET.
const url = `/graphql?id=${QUERY_HASH}&variables=${encodeURIComponent(JSON.stringify(vars))}`;
const data = await fetch(url).then((r) => r.json());
// trade-off: persisted queries need a registry of allowed query hashes on the
// server. That is also a security benefit, but it adds a build step.
4. Purge by tag when data changes
When a product changes, purge product-<id> so every cached response that includes it (detail, listings, search) is refreshed.
Verification
Request each public endpoint twice and confirm a cache HIT on the second request; request a personal endpoint and confirm it is never cached. Update a record and confirm the next request returns fresh data after the purge. In RUM, measure data request latency (Resource Timing for API URLs) and route transition times; both should fall.
Worked Example: Category Listing API
An e-commerce SPA fetched category listings from /api/categories/:slug/products?page=n&sort=… on every navigation, with origin latency of 280ms at p75. The team added s-maxage=120, stale-while-revalidate=600, normalised the query parameters in the cache key, tagged responses with every product ID and the category, and purged tags from the product-update pipeline. The edge hit ratio for listings reached 89%, data latency p75 fell to 45ms, and route transitions between categories became noticeably snappier; INP was unaffected, but route LCP improved by about 230ms.
Common Mistakes
- Caching responses that vary by user. Personal discounts, inventory reserved for a user, or localisation from a profile must not be cached publicly.
- Long browser
max-ageon APIs. Browsers cannot be purged; keep browser TTLs short or zero and let the CDN hold the data. - Too many tags. Some CDNs cap the size of tag headers; tag by entity IDs that actually drive purges.
- Forgetting CORS. Cached responses must carry the right
Access-Control-Allow-Origin; vary onOriginonly if you serve different values.
Edge Cases
Personalised fields inside public data. Split them: cache the public product and fetch the user's price or wishlist status separately.
Rate limits and quotas. Cached responses reduce origin calls, but ensure rate limiting happens before the cache only for uncached routes, or legitimate cached traffic may be throttled.
Large responses. Very large JSON payloads cost transfer time even when cached; paginate and compress.
Error responses. Cache 404s briefly (seconds) and never cache 5xx, or an outage gets cached.
FAQ
Can I cache GraphQL at the CDN?
Yes, with GET-based persisted queries and responses that do not depend on the user. Client-side GraphQL caches (Apollo, urql) complement but do not replace edge caching, because they only help the same client.
How short can TTLs be and still help?
Even a few seconds helps under load — it collapses bursts of identical requests into one origin call. For user-facing latency, longer TTLs with stale-while-revalidate give the best results, with purges handling freshness.
Should API caching use the same rules as HTML?
The principles are the same — public vs private, keys, TTLs, purging — but APIs are usually more granular, so tag-based purging and parameter allow-lists matter more.
Does caching APIs help INP?
Indirectly: interactions that trigger data fetches (filters, pagination) complete their visible update sooner when data comes from the edge. INP measures the next paint after input, which may be a loading state, but users perceive the faster result.
What about server-side rendering that calls these APIs?
SSR servers benefit too if they call the API through the CDN (or a shared cache). That turns many page renders' data fetches into cache hits and lowers TTFB.
What about APIs called from other servers, not browsers?
Server-to-server calls (SSR fetching data, microservices) can use the same CDN or an internal shared cache. The headers and purge strategy are the same; the main difference is that you can often allow longer TTLs because the consumer is under your control.
How should I handle ETags on cached APIs?
Keep them. Strong or weak ETags let the CDN revalidate cheaply with the origin when a TTL expires, and let browsers revalidate with the CDN. Make sure ETags are stable across origin instances — ETags derived from server-specific values defeat revalidation.
Is it worth caching APIs with very low traffic?
Usually not for latency — each location rarely reuses the entry — but tiered caching changes that by concentrating requests at a shield. For low-traffic endpoints, focus on making the origin response itself fast and add caching mainly for spike protection.
Related
- Stale-while-revalidate data fetching with SWR and TanStack Query — the client-side cache on top.
- Purging CDN cache by tag on deploy — tag purging in depth.
- Versioning API responses for cache busting — an alternative to purging.