Rendering & CSS Performance: Making Style, Layout and Paint Cheap

Most performance work focuses on the network: fewer bytes, earlier requests, better caching. But once the bytes arrive, the browser still has to turn HTML and CSS into pixels, and it has to do that again every time something changes. That work — recalculating styles, computing layout, painting and compositing layers — runs largely on the main thread, the same thread that handles user input. When it is expensive, pages render late (hurting Largest Contentful Paint), respond slowly to clicks and typing (hurting Interaction to Next Paint), and shift around as late layout happens (hurting Cumulative Layout Shift).

Rendering costs are driven by decisions developers make every day: how CSS is delivered and how much of it there is, how large and deep the DOM grows, whether JavaScript reads layout in the middle of writing styles, which properties are animated, and whether independent parts of the page are isolated from each other. None of these show up as large network requests, so they are easy to miss in a waterfall. They show up instead as purple and green blocks in the Performance panel — Recalculate Style, Layout, Paint — and as long tasks that delay input.

This section is organised around the rendering pipeline. Critical CSS and render-blocking resources covers getting the first frame painted. CSS containment and content-visibility covers limiting how much work each change triggers. Layout thrashing and forced reflow covers the JavaScript patterns that make layout run repeatedly. Animation and compositing performance covers moving motion off the main thread. DOM size and list virtualization covers the cost multiplier behind all of the above: the number of elements.

The browser rendering pipeline The stages a browser runs to turn DOM and CSS into pixels, from style calculation through layout, paint and composite. The browser rendering pipeline DOM + CSSOM parse HTML and CSS Style match selectors Layout geometry of boxes Paint draw records Composite layers to GPU

Diagnostic Overview: Where Rendering Time Goes

Start from a Performance panel recording on a realistic device (or with 4x CPU throttling) covering page load and a few typical interactions. In the summary, compare time spent in Rendering (style and layout) and Painting against Scripting. On many content sites, rendering is a modest share; on complex applications, dashboards and long feeds, it can rival or exceed JavaScript execution.

Then look for the four signatures that map to the topics in this section:

  • Late first paint with an idle main thread. The browser is waiting for render-blocking CSS (or synchronous scripts) to download. The fix is in delivery: inline critical CSS, defer the rest, remove unused rules. This is a network-shaped rendering problem and directly delays FCP and LCP.
  • Large single Layout or Recalculate Style blocks. A change touched many elements, or the DOM is very large. The fix is to reduce the scope of each change (containment, content-visibility) or the number of elements (virtualization, simpler markup).
  • Many small Layout blocks interleaved with script, often with a warning triangle labelled "Forced reflow". JavaScript is reading geometry (offsetHeight, getBoundingClientRect) after writing styles, forcing synchronous layout again and again. The fix is batching reads and writes or using observers.
  • Long Paint or frequent Composite work during animation, with dropped frames. Animated properties trigger layout or paint on every frame. The fix is animating only transform and opacity, and keeping layer counts sensible.

Each signature has a characteristic metric impact. Render-blocking CSS hurts FCP and LCP for every visitor. Expensive style and layout hurt INP, because the browser must render the next frame after an event handler runs, and that frame includes the style and layout work the handler triggered. Unbounded layout cost on large DOMs hurts both. Poorly implemented animations mostly hurt smoothness, but layout-triggering animations also cause layout shifts that count towards CLS.

Rendering symptoms and the metrics they move Mapping of common rendering problems to the Core Web Vitals they affect most. Rendering symptoms and the metrics they move Problem LCP INP CLS Render-blocking CSS delays first paint minor FOUC shifts if misused Large DOM and style cost slower first layout slow next paint minor Forced reflow in handlers minor long tasks on input minor Animating layout properties minor busy main thread shifts counted

A useful habit when profiling is to record the same interaction twice: once as a typical user would perform it, and once with the CPU throttled four times. Rendering costs scale with device speed, so a 40ms layout on a developer laptop can become 160ms on a budget phone, which is exactly where INP and LCP fail in the field.

Architecture 1: Getting the First Frame Painted

The browser will not paint anything until it has the CSS it considers render-blocking: every stylesheet linked in the head without a non-matching media attribute. On a fast connection, that is a few tens of milliseconds; on a slow mobile connection with a 200KB framework stylesheet served from another origin, it can be over a second of blank screen, during which the main thread is mostly idle.

Three techniques address this. Critical CSS inlines the small set of rules needed to render above-the-fold content directly in the HTML, so the first paint does not wait for any stylesheet download. Asynchronous loading of the remaining CSS (for example with media="print" swapped to all on load, or rel="preload" with an onload handler) lets it arrive without blocking. Removing unused CSS shrinks everything — many sites ship 80–90% unused rules from frameworks and old features. A fourth consideration is the choice of styling architecture: runtime CSS-in-JS libraries generate and inject styles during rendering, which adds scripting and style recalculation cost on the client; zero-runtime alternatives extract static CSS at build time.

The trade-offs are real. Inlined CSS is not cached separately, so it is re-sent on every page view; it works best when small (under roughly 14KB compressed so it fits early in the response). Asynchronously loaded CSS can cause a flash of unstyled content or layout shifts if the critical subset misses rules for visible elements. The guides in Critical CSS and render-blocking resources cover extraction, safe async loading and measuring the effect on LCP.

Architecture 2: Limiting the Scope of Rendering Work

By default, a change anywhere in the document can, in principle, affect layout anywhere else, so browsers must be conservative about what they recompute. CSS containment tells the browser that a subtree is independent. contain: layout means the inside of an element does not affect layout outside it; contain: paint means its contents do not paint outside its bounds; contain: size means its size does not depend on its contents. With these guarantees, the browser can limit style, layout and paint work to the subtree that changed.

content-visibility: auto builds on containment to skip rendering work entirely for offscreen content. A long article or feed with content-visibility: auto on each section lets the browser skip style, layout and paint for sections far from the viewport, then render them as the user scrolls. Initial rendering cost can drop dramatically on long pages. The companion property contain-intrinsic-size provides a placeholder size for skipped content, so the scrollbar and page height stay stable.

Containment is a hint about independence, not a magic speed switch. Applied to widgets that really are independent — a chat panel, an ad slot, a comment thread — it reduces the blast radius of updates. Applied blindly, it can clip content (paint containment), break position: sticky in some layouts, or cause size jumps when intrinsic sizes are wrong. The guides in CSS containment and content-visibility explain where it pays off and how to verify the savings.

Initial rendering time for a 12,000-word article page Bar chart comparing rendering time on initial load for a long article with and without content-visibility auto. Initial rendering time for a 12,000-word article page No containment 410ms contain: layout paint on sections 330ms content-visibility: auto on sections 95ms Measured as style + layout + paint on a mid-range phone profile.

Architecture 3: Avoiding Layout Thrashing

Browsers batch style and layout work: when JavaScript changes styles, the browser normally waits until the next frame to recompute layout once. That batching breaks when script asks for layout information — offsetWidth, getBoundingClientRect(), scrollTop, getComputedStyle() — after making a change. To answer accurately, the browser must run style and layout synchronously, right now. Do this in a loop (write, read, write, read) and layout runs once per iteration: layout thrashing.

Thrashing is a classic cause of long tasks in event handlers and therefore poor INP. Common sources include measuring elements in a loop to position tooltips, animating with JavaScript that reads current positions every frame, third-party widgets that measure their container repeatedly, and resize or scroll handlers that both read and write geometry. The cure is structural: read all the geometry you need first, then write all the changes; defer writes to requestAnimationFrame; and replace polling handlers with ResizeObserver and IntersectionObserver, which deliver geometry at points where layout is already up to date.

The Performance panel flags forced reflows explicitly, with a stack trace pointing to the line of code that read layout. The guides in Layout thrashing and forced reflow walk through finding them, batching reads and writes, and migrating handlers to observers.

Write-read loop vs batched reads and writes (100 items) Bar chart comparing main-thread time for a loop that forces layout on every iteration with a batched approach that runs layout once. Write-read loop vs batched reads and writes (100 items) Interleaved write/read (100 layouts) 110ms Batched reads then writes (1 layout) 18ms

Architecture 4: Animating on the Compositor

Each frame at 60Hz has about 16.7ms; on 120Hz displays, about 8.3ms. If an animation changes a property that affects layout (width, height, top, left, margin), each frame needs style, layout, paint and composite — on the main thread, competing with scripts. If it changes a paint-only property (background-color, box-shadow), layout is skipped but paint still runs. If it changes only transform or opacity on an element with its own compositing layer, the compositor thread can run the animation alone, even while the main thread is busy.

This is why "animate only transform and opacity" is the most important rule of animation performance. Movement becomes translate, size changes become scale, fades become opacity. Layout-affecting animations also cause layout shifts that may count towards CLS when they are not triggered by user input.

Layers have a cost too. will-change: transform promotes an element to its own layer ahead of time, which helps animations start smoothly, but each layer consumes GPU memory, and promoting hundreds of elements ("layer explosion") can make performance worse, especially on low-memory phones. Newer platform features shift more work to the compositor: scroll-driven animations (animation-timeline: scroll() and view()) replace scroll listeners with declarative CSS, and the View Transitions API animates between page states with snapshots. Both have their own performance characteristics. The guides in Animation and compositing performance cover each.

Architecture 5: Controlling DOM Size

Almost every rendering cost scales with the number of elements: style recalculation must match selectors against elements, layout must compute boxes, paint must record them, and memory grows with each node. Lighthouse warns at around 800 body elements and flags pages above roughly 1,400 elements; it also reports maximum depth and the largest number of children under one parent. Pages with 5,000–20,000 elements are common on e-commerce category pages, admin dashboards and long feeds, and they show consistently worse INP.

The main techniques are: virtualization (render only the rows of a long list that are near the viewport, recycling elements as the user scrolls), pagination or "load more" instead of infinite DOM growth, deferring hidden UI (menus, modals and tabs rendered only when opened), and simplifying markup (removing wrapper divs that exist only for styling, which modern layout with grid and flex rarely needs). Selector complexity matters less than element count in modern engines, but style invalidation patterns — such as toggling a class on body that affects thousands of descendants — can still create very large recalculations.

The guides in DOM size and list virtualization cover virtualizing lists in React, fixing the excessive DOM size warning, and reducing style recalculation cost.

INP p75 by DOM size band (one retail site, field data) Bar chart of field INP at the 75th percentile across pages grouped by number of DOM elements. INP p75 by DOM size band (one retail site, field data) < 1,000 elements 140ms 1,000-3,000 190ms 3,000-8,000 270ms > 8,000 410ms INP good threshold

Monitoring and CI: Keeping Rendering Costs in Budget

Rendering regressions arrive quietly: a new component library adds 40KB of CSS, a product grid doubles its DOM, a carousel starts animating left. Guard against them with a few automated checks.

In CI, track CSS bytes (total and render-blocking) with a bundle size budget, and run Lighthouse on key templates with assertions on DOM size, render-blocking resources and total blocking time. Lab rendering times are noisy, but DOM element counts and CSS sizes are deterministic and make excellent budget metrics. In the field, attribute INP to interaction targets with the Long Animation Frames API: its script attribution and styleAndLayoutStart timing reveal how much of a slow frame was style and layout rather than script. Segment INP by template and DOM size band to see whether a rendering change moved the metric.

For animations, record smoothness in the lab with the Performance panel's frames track, and watch for dropped frames during key interactions such as opening menus or scrolling feeds. For CLS, layout shift attribution identifies elements that moved; animations of layout properties often appear here.

Reference Implementations

Inlining critical CSS and loading the rest asynchronously

html
<head>
  <style>/* critical rules for header, hero and above-the-fold layout (~8KB) */</style>
  <link rel="preload" href="/css/site.4f2a1c.css" as="style">
  <link rel="stylesheet" href="/css/site.4f2a1c.css" media="print" onload="this.media='all'">
  <noscript><link rel="stylesheet" href="/css/site.4f2a1c.css"></noscript>
</head>
<!-- trade-off: if the critical subset misses rules for visible content, users see
     a flash of unstyled content and layout shifts when the full CSS applies. -->

Skipping offscreen rendering on long pages

css
.article-section {
  content-visibility: auto;
  contain-intrinsic-size: auto 900px;   /* remembered after first render */
}
/* trade-off: in-page find and anchor links still work, but rendering on scroll
   costs a little time per section; avoid on sections just below the fold. */

Batching DOM reads and writes

javascript
function layoutTooltips(tooltips) {
  // Read phase: gather geometry once.
  const rects = tooltips.map((t) => t.anchor.getBoundingClientRect());
  // Write phase: apply all positions together, in the next frame.
  requestAnimationFrame(() => {
    tooltips.forEach((t, i) => {
      t.el.style.transform = `translate(${rects[i].left}px, ${rects[i].bottom + 8}px)`;
    });
  });
}
// trade-off: deferring writes to the next frame adds up to one frame of latency;
// for tooltips this is invisible, for drag handles it may not be.

Compositor-friendly animation

css
.drawer { transform: translateX(-100%); transition: transform 240ms ease-out; }
.drawer.open { transform: translateX(0); }
/* Not: transition: left 240ms — that runs layout on every frame. */
@media (prefers-reduced-motion: reduce) { .drawer { transition: none; } }
/* trade-off: transformed elements keep their original layout box, so content
   behind the drawer does not reflow; design the layout with that in mind. */

Virtualizing a long list

javascript
import { useVirtualizer } from '@tanstack/react-virtual';
function OrderList({ orders }) {
  const parentRef = React.useRef(null);
  const v = useVirtualizer({ count: orders.length, getScrollElement: () => parentRef.current, estimateSize: () => 56, overscan: 8 });
  return (
    <div ref={parentRef} style={{ height: 600, overflow: 'auto' }}>
      <div style={{ height: v.getTotalSize(), position: 'relative' }}>
        {v.getVirtualItems().map((row) => (
          <div key={row.key} style={{ position: 'absolute', top: 0, transform: `translateY(${row.start}px)`, height: 56, width: '100%' }}>
            {orders[row.index].number}
          </div>
        ))}
      </div>
    </div>
  );
}
// trade-off: virtualized rows are not in the DOM, so browser find-in-page and some
// assistive technologies cannot see them; provide search and pagination as well.

Common Pitfalls

  • Inlining all CSS. Critical CSS should be a small subset; inlining a 150KB stylesheet delays the HTML instead of the CSS.
  • Applying content-visibility: auto to above-the-fold content. It adds work at the moment it is first rendered without any skipped work to gain.
  • Measuring layout inside loops. Every read after a write forces layout.
  • will-change on everything. Layer explosion consumes memory and slows compositing.
  • Animating height for accordions. Use transforms, or accept layout cost and keep the affected subtree small with containment.
  • Rendering every list item. Thousands of rows multiply the cost of every style and layout operation.
  • Toggling theme or state classes on body without containment. Very large style invalidations on every toggle.
  • Ignoring rendering in INP investigations. Handlers may be fast while the rendering they trigger is slow.

FAQ

Is CSS really a performance problem if it is small?

Delivery is mostly about size and blocking; once CSS is small and non-blocking, its runtime cost depends on how styles are applied — how many elements match, how often classes change, and what properties change. A small stylesheet on a huge DOM can still be expensive to recalculate.

Do CSS selectors still matter for performance?

Much less than they used to. Modern engines match selectors very efficiently. Element count and the scope of style invalidations matter far more than whether you use descendant or class selectors.

How does rendering affect INP?

INP measures until the next frame is painted after an interaction. If the event handler changes styles that require recalculating style and layout for thousands of elements, that work is part of the interaction's presentation delay, even if the handler's JavaScript is fast.

Should I use content-visibility everywhere?

No. It helps on long pages with substantial offscreen content. For short pages or above-the-fold sections it adds overhead without benefit, and wrong intrinsic sizes can cause scrollbar jumps.

Is JavaScript animation always slower than CSS animation?

Not inherently. The Web Animations API animating transform and opacity can run on the compositor just like CSS. JavaScript animation that updates styles every frame with requestAnimationFrame runs on the main thread and can stutter when the thread is busy.

What DOM size should I aim for?

Below about 1,500 elements on initial load is a reasonable target for most pages; Lighthouse warns above roughly 800 and flags pages above about 1,400. Applications can exceed that if they virtualize long lists and keep updates scoped.

Does CSS-in-JS hurt performance?

Runtime CSS-in-JS adds scripting to generate styles and style recalculation when it injects them, especially during hydration and re-renders. Zero-runtime or compile-time approaches avoid most of the cost while keeping a component-oriented authoring style.

Topics in This Section