How to Diagnose Head-of-Line Blocking

This guide is part of HTTP/2, HTTP/3 & Connection Management, within Network & Server Response Optimization. Head-of-line (HOL) blocking happens when something at the front of a queue holds up everything behind it. On the web, it occurs at two layers. At the HTTP layer, HTTP/1.1 processes one response at a time per connection, so a slow or large response blocks the next. HTTP/2 fixed this with multiplexing — but at the transport layer, HTTP/2 still runs over TCP, which delivers bytes strictly in order. If a single packet is lost, every stream on the connection waits for its retransmission, even streams whose data arrived fine. HTTP/3 over QUIC removes transport-level HOL blocking because each stream is delivered independently.

A third, related problem is prioritisation blocking: even without packet loss, a server that sends large, low-priority responses (images, video) before small critical ones (CSS, the LCP image) delays rendering. The symptoms look similar in waterfalls — critical resources finishing late despite starting early — so diagnosing which cause applies is the key step.

One lost packet on an HTTP/2 connection Timeline showing how a lost TCP packet stalls all HTTP/2 streams on the connection while HTTP/3 streams continue independently. One lost packet on an HTTP/2 connection H2 CSS stream data stalled (loss on other stream) done H2 image stream data retransmit wait H3 CSS stream data + done 0ms 100ms 200ms 300ms 400ms 500ms 600ms

Rapid Diagnosis

  • Look for long "Content Download" phases on small critical resources in the waterfall while other resources are downloading.
  • Check the protocol per request: HTTP/2 over TCP is susceptible to transport HOL blocking.
  • Check network conditions: problems cluster on lossy mobile networks and congested Wi-Fi.
  • Check request priority in DevTools (Priority column) and the order the server delivers responses.

Root Cause Analysis

1. Packet loss on TCP. All HTTP/2 streams stall until retransmission.

2. Poor server prioritisation. The server sends low-priority bytes first or interleaves them evenly with critical ones.

3. Large responses early in the queue. Big images or JSON scheduled before critical CSS.

4. HTTP/1.1 fallback. Some clients or proxies still use HTTP/1.1, with classic HTTP-level HOL blocking.

Step-by-Step Resolution

1. Distinguish loss from prioritisation

Reproduce with network emulation that includes packet loss (for example, WebPageTest custom profiles or tc netem), and compare with a lossless run. If critical resources are only late with loss, it is transport HOL blocking; if they are late even without loss, it is prioritisation.

bash
# Linux: add 2% packet loss and 100ms delay on an interface for testing.
sudo tc qdisc add dev eth0 root netem delay 100ms loss 2%
# ... run tests ...
sudo tc qdisc del dev eth0 root
# trade-off: system-wide emulation affects everything on the machine; use a
# dedicated test box or container network namespace.

2. Enable HTTP/3

For transport HOL blocking, HTTP/3 is the structural fix. Enable it at the CDN and verify the share of users on h3.

3. Fix prioritisation

Use a CDN or server with good HTTP/2 and HTTP/3 prioritisation (Extensible Priorities support). Mark the LCP image with fetchpriority="high", avoid preloading low-priority resources, and lazy-load offscreen images so they do not compete.

4. Reduce early contention

Keep early responses small: critical CSS inlined or small, the LCP image right-sized, and large non-critical resources deferred until after first render.

Which kind of blocking is it? Decision sequence for identifying the cause of critical resources finishing late. Which kind of blocking is it? Is the connection HTTP/1.1? HTTP-level HOL blocking — enable HTTP/2 or HTTP/3 yes no Are resources late only under packet loss? Transport HOL blocking — enable HTTP/3 yes no Are large low-priority responses sent first? Prioritisation — fix server priorities and hints yes no Check bandwidth and resource size

Verification

Under emulated packet loss, critical CSS and the LCP image should finish closer to their lossless times with HTTP/3 enabled. Without loss, the waterfall should show critical resources completing before large images. In the field, compare LCP p75 for users on high-loss networks (approximated by high-RTT segments or specific countries) before and after.

Worked Example: A Video Streaming Site's Homepage

A streaming service's homepage loaded 40 large poster images alongside critical CSS and the hero image over one HTTP/2 connection. On mobile networks in several markets, LCP p75 was 4.2 seconds, and waterfalls showed the critical CSS (12KB) taking 900ms to download. Lossless tests showed CSS finishing in 150ms, but with 2% loss it took 700–1,000ms. The team enabled HTTP/3 at the CDN, lazy-loaded posters below the first row, and added fetchpriority="high" to the hero image. With HTTP/3, CSS finished in 220ms under the same loss. LCP p75 in the affected markets fell to 2.6 seconds.

Prioritisation in Practice

Browsers assign priorities based on resource type and position: render-blocking CSS and fonts are highest, the LCP image is raised once detected in the viewport (or immediately with fetchpriority="high"), scripts depend on async/defer, and images default to low. Servers should use these signals to schedule bytes. Some servers and CDNs historically ignored HTTP/2 priorities or implemented them poorly; modern CDNs generally support the Extensible Priorities scheme. If your waterfall shows a high-priority request starting early but downloading slowly while low-priority images stream in parallel, the server is likely not prioritising well.

Critical CSS download time (12KB) under 2% packet loss Bar chart comparing critical CSS download time under packet loss for HTTP/2 and HTTP/3, with and without lazy-loaded posters. Critical CSS download time (12KB) under 2% packet loss H2 + 40 eager posters 900ms H2 + lazy posters 520ms H3 + lazy posters 220ms

Common Mistakes

  • Blaming HTTP/2 for prioritisation problems. The server's scheduling may be the cause.
  • Testing only lossless networks. Transport HOL blocking never appears.
  • Preloading many resources. Turns them all into high-priority competitors.
  • Assuming HTTP/3 fixes everything. It removes transport HOL blocking, not bad prioritisation or oversized resources.

Edge Cases

Single huge responses. One very large HTML document or JSON blocks nothing else on HTTP/2, but its own content arrives in order; streaming helps.

Corporate proxies. May downgrade to HTTP/1.1, reintroducing HTTP-level HOL blocking for those users.

Server push remnants. Old configurations pushing resources can compete with critical requests; push is removed in browsers but may still waste server effort.

Video segments. Media streams on the same connection compete with page resources; consider separate delivery for heavy media.

FAQ

Does HTTP/2 have head-of-line blocking?

Not at the HTTP layer, thanks to multiplexing. It still suffers TCP-level head-of-line blocking when packets are lost.

Does HTTP/3 eliminate head-of-line blocking?

It eliminates transport-level HOL blocking between streams. Data within a single stream is still delivered in order.

How common is packet loss?

On mobile and congested Wi-Fi, loss of around 1–2% is not unusual, and enough to trigger noticeable stalls on TCP.

How do I test with packet loss?

Use WebPageTest custom connectivity profiles, Linux tc netem, or dedicated network emulators. DevTools throttling does not simulate loss.

What is prioritisation blocking?

When a server sends low-priority data before or alongside high-priority data, delaying critical resources even without packet loss.

Does fetchpriority help?

It raises the browser's priority signal for a resource, which a well-behaved server uses to schedule it earlier. It cannot fix transport HOL blocking.

Can too many lazy images still compete?

Lazy images near the viewport load early by design. Native lazy loading uses distance thresholds, so only nearby images compete with critical resources.

Should critical resources use a separate connection?

No. Separate connections lose prioritisation and add setup costs. Fix prioritisation and enable HTTP/3 instead.

Is head-of-line blocking visible in Core Web Vitals?

Indirectly. It delays critical CSS and the LCP image, so it shows up as worse FCP and LCP for users on lossy networks, usually in the upper percentiles and in specific regions.

Can a service worker help?

A service worker serving critical resources from cache avoids the network entirely on repeat visits, which sidesteps head-of-line blocking for those resources.

Does HTTP/1.1 suffer more?

At the HTTP layer, yes: each connection handles one response at a time, so slow responses block the queue.