How to Fix TTFB from Serverless Cold Starts
This guide is part of Time to First Byte Optimization, within Network & Server Response Optimization. Serverless functions — AWS Lambda, Google Cloud Run functions, Azure Functions, and the server rendering functions behind Vercel and Netlify deployments — scale by creating new instances on demand. A new instance must start a runtime, load your code, and run initialisation (importing modules, creating clients, connecting to databases) before it can handle the request. This cold start can add 200ms to several seconds to TTFB.
For high-traffic routes, most requests hit warm instances and cold starts are rare. But low-traffic pages, traffic spikes, new deployments and regions with little traffic see many cold starts, and they show up as a long tail in TTFB: p75 looks acceptable, p90 and p95 do not. On sites where every page is rendered by functions, cold starts can be the main reason for poor field LCP.
Rapid Diagnosis
- Report init time as a Server-Timing entry (or log it) to identify cold requests.
- Compare TTFB distributions: a long tail at p90+ with a fast median suggests cold starts.
- Check platform metrics: Lambda's
Init Durationin logs, or equivalent on other platforms. - Check function bundle size: large bundles take longer to load.
Root Cause Analysis
1. Large bundles. Loading and parsing megabytes of JavaScript at startup.
2. Heavy initialisation. Creating SDK clients, loading configuration, warming caches on startup.
3. Database connections. New TCP and TLS connections from each new instance, sometimes with connection limits.
4. Low or spiky traffic. More requests land on new instances.
Step-by-Step Resolution
1. Measure cold start share and cost
// Module scope runs once per instance.
const bootStart = performance.now();
let coldStart = true;
const initDone = (async () => { /* create clients, load config */ })();
export async function handler(req, res) {
await initDone;
const init = coldStart ? (performance.now() - bootStart).toFixed(1) : '0';
res.setHeader('Server-Timing', `cold;desc="${coldStart}", init;dur=${init}`);
coldStart = false;
// ... handle request
}
// trade-off: performance.now() at module scope misses runtime startup before your
// code loads; combine with the platform's reported init duration.
2. Shrink the function bundle
Bundle and tree-shake server code, mark large unused dependencies as external or remove them, and split routes into separate functions so each loads only its code. Avoid importing whole SDKs when a modular client exists.
3. Defer and reuse initialisation
Create clients lazily on first use, reuse them across invocations (module-scope variables), and avoid work at startup that is not needed for most requests.
4. Use connection pooling and HTTP keep-alive
Use a database proxy or HTTP-based data API designed for serverless (connection poolers), and enable keep-alive for outgoing HTTP calls so warm instances reuse connections.
Verification
Track the share of requests with cold=true and their TTFB in RUM or logs. After changes, init duration should drop and the TTFB tail (p90, p95) should tighten. Check that warm request latency did not regress (lazy initialisation can shift work into the first request that needs it).
Worked Example: A Marketing Site on Serverless SSR
A B2B company's marketing site rendered every page with serverless SSR. Traffic was modest and spread across hundreds of pages, so about 18% of requests hit cold instances. Cold TTFB was 1.6–2.4 seconds; warm TTFB 220ms; overall TTFB p75 was 780ms and p90 2.1 seconds. The team cached rendered HTML at the CDN with s-maxage=600, stale-while-revalidate=86400 (most pages changed rarely), split the single large function into per-route functions, and replaced a full cloud SDK import with modular clients. Cache hit ratio for HTML reached 96%, so cold starts affected under 1% of navigations; cold init time itself fell from 1.3 seconds to 380ms. TTFB p75 dropped to 190ms and p90 to 420ms.
Avoiding Cold Starts Altogether
The most effective fix is often to avoid running a function for most requests. For content that is the same for many users, cache HTML at the CDN or pre-render at build time (static generation or incremental static regeneration), so functions only run for cache misses and revalidations. For dynamic pages, platform features can help: provisioned or minimum instances keep a number of instances warm (at a cost), and some platforms offer faster-starting runtimes or snapshotting that reduces init time. Edge functions on V8 isolates start in milliseconds, but have limited APIs and should stay close to their data.
Common Mistakes
- Scheduled "warming" pings. Keep one instance warm but do nothing for concurrent requests or traffic spikes.
- Importing everything at module scope. Pays the full cost on every cold start.
- New database connections per request. Slow and may exhaust database limits.
- Ignoring the tail. Medians hide cold starts; look at p90 and p95.
Edge Cases
Deployments. Every deploy starts new instances; staggered rollouts and pre-warming reduce the spike.
Multi-region. Low-traffic regions see more cold starts; consider fewer regions with CDN caching in front.
Memory settings. On some platforms, higher memory allocation also gives more CPU, speeding up initialisation.
Language runtimes. Runtimes differ in startup cost; JVM-based functions benefit most from snapshotting features.
FAQ
How long is a typical cold start?
From about 100ms for small functions on fast runtimes to several seconds for large bundles with heavy initialisation. Measure your own functions.
Do warming pings help?
Only slightly. They keep one instance warm, but concurrent requests and spikes still create new instances.
Is provisioned concurrency worth it?
For latency-sensitive, steady-traffic routes, often yes. It costs money for idle capacity, so combine it with caching to keep the number of needed instances small.
Do edge functions have cold starts?
Isolate-based edge functions start very quickly (milliseconds), but calling distant databases from the edge can add more latency than the cold start saves.
How do I know if a request was cold?
Track a module-level flag as in the example, or read the platform's init duration from logs, and report it via Server-Timing.
Does bundle size really matter on the server?
Yes. Loading and evaluating large bundles is a major part of cold start time. Server code deserves the same tree-shaking attention as client code.
Should every page be server-rendered on request?
No. Pages that are the same for all users should be static or cached at the CDN. Reserve per-request rendering for genuinely dynamic responses.
Do cold starts affect Core Web Vitals?
Yes, through TTFB, which delays FCP and LCP. Because they hit a minority of requests, they mostly affect the upper percentiles, but on low-traffic sites they can move p75.
Can I pre-warm functions after a deploy?
Some platforms support it; otherwise, a script that requests key routes in each region after deployment reduces the cold-start spike for early visitors.
Does the runtime language matter?
Yes. Lightweight runtimes start faster; heavier ones benefit more from snapshot features. Within one language, bundle size and initialisation work usually matter most.
Related
- Diagnosing slow TTFB with Server-Timing — measure init and handler time.
- Cache-Control for HTML documents — cache HTML to avoid running functions.
- Nuxt route rules for hybrid rendering — choose static, cached or dynamic per route.