How to Choose AVIF and WebP Quality Settings

This guide refines Serving AVIF and WebP with Fallbacks, within Image & Media Optimization. Switching to AVIF or WebP saves bytes only if the quality settings are chosen well. Too high, and modern formats lose much of their advantage over JPEG; too low, and artefacts appear — blurred textures, banding in gradients, smeared text. Most teams pick a number once ("AVIF 50, WebP 75") and apply it to everything, which over-compresses detailed photos and under-compresses simple graphics.

Quality numbers are not comparable across formats or even encoders, and the same number produces very different visual quality on a product shot, a screenshot and a sunset. The reliable approach is to target a perceptual quality level, measured with metrics designed to match human judgement, and find the lowest setting per format and content type that reaches it.

Bytes at equal perceptual quality (1200px product photo) Bar chart of file sizes for one product photo in JPEG, WebP and AVIF when each is tuned to the same perceptual quality score. Bytes at equal perceptual quality (1200px product photo) JPEG (mozjpeg q80) 168KB WebP (q78) 121KB AVIF (q55, speed 4) 84KB AVIF (one-size-fits-all q40) 61KB The last bar is smaller but falls below the quality target, with visible texture loss.

Rapid Diagnosis

  • Check your current settings in the build pipeline or image CDN configuration; note whether one value applies to everything.
  • Spot-check outputs at 100% zoom on a high-DPR display: textures (fabric, foliage), gradients (skies) and text (screenshots, labels).
  • Compare sizes across content types. If screenshots and illustrations are as large as photos, settings are not tuned per type.
  • Check encoder speed settings. Very fast AVIF settings produce larger files at the same quality.

Root Cause Analysis

1. One global quality value. Different content needs different settings to reach the same perceived quality.

2. Quality numbers are encoder-specific. "Quality 60" in one AVIF encoder is not the same as in another, or as WebP 60.

3. Eyeballing on low-DPI screens. Artefacts invisible on a laptop at 1x can be obvious on phones at 3x zoom levels.

4. Ignoring encoder effort. Faster AVIF speed settings trade compression efficiency for build time.

Starting quality settings by content type Starting AVIF and WebP quality settings for different kinds of images, to be validated with perceptual metrics. Starting quality settings by content type Content AVIF (sharp quality) WebP quality Watch for Product photos 50-60 75-82 texture loss Hero photography 45-55 72-80 banding in skies Screenshots / UI 60-70 or lossless 85+ or lossless smeared text Illustrations / flat colour 40-55 75-80 edge ringing

Step-by-Step Resolution

1. Pick a perceptual metric and target

Use SSIMULACRA2 or Butteraugli (both designed to correlate with human perception). A SSIMULACRA2 score around 70–80 is often considered high quality for web photos; choose a target with design stakeholders by reviewing samples.

2. Search for the lowest setting that meets the target

javascript
// scripts/tune-quality.mjs — binary search AVIF quality to reach a SSIMULACRA2 target.
import sharp from 'sharp';
import { execFileSync } from 'node:child_process';
async function score(refPath, buf) {
  await sharp(buf).png().toFile('/tmp/candidate.png');
  return Number(execFileSync('ssimulacra2', [refPath, '/tmp/candidate.png']).toString().trim());
}
export async function tune(refPath, target = 75) {
  let lo = 20, hi = 90, best = hi;
  while (lo <= hi) {
    const q = Math.floor((lo + hi) / 2);
    const buf = await sharp(refPath).avif({ quality: q, effort: 4 }).toBuffer();
    if (await score(refPath, buf) >= target) { best = q; hi = q - 1; } else { lo = q + 1; }
  }
  return best;
}
// trade-off: per-image tuning is expensive; run it on a representative sample
// per content type to choose settings, rather than for every image in the build.

3. Apply settings per content type

Classify images (by folder, CMS field or tag) and apply the tuned setting per class in the pipeline or as image CDN presets.

4. Choose encoder effort for your build budget

For AVIF in sharp, effort 4–6 is a common balance; higher effort shrinks files a few percent more at much longer encode times. Cache outputs so the cost is paid once per image.

Quality tuning workflow Workflow from sampling images per content type through perceptual scoring to per-class encoder settings. Quality tuning workflow Sample 20-50 images per type Target perceptual score agreed Search lowest passing quality Apply per-class presets Re-check quarterly

Verification

Compare total image bytes per page before and after, review a sample visually at high zoom with stakeholders, and keep the perceptual score distribution for each class above the target. In RUM, LCP image resourceLoadDuration should fall in proportion to the byte savings.

Worked Example: An Online Furniture Store

An online furniture store used AVIF quality 40 for everything. Fabric textures on sofas looked blurry on high-DPR phones and returns mentioning "colour/texture not as pictured" rose. A tuning pass with SSIMULACRA2 found product photos needed quality 58 to meet the agreed target, while lifestyle backgrounds passed at 46 and line-drawing diagrams at 35. Applying per-class presets increased product image bytes by 18% but reduced backgrounds and diagrams by 22% and 30%; total image weight per product page was roughly unchanged, while texture complaints dropped noticeably.

Lossless and Near-Lossless Cases

Some content should not be lossy at all. Screenshots of UI, diagrams with thin lines, and images containing small text often look worse with lossy AVIF than with lossless WebP or PNG at a similar size. WebP's lossless mode and AVIF's lossless or high-quality modes with 4:4:4 chroma subsampling preserve sharp coloured edges that 4:2:0 subsampling blurs. Classify such images separately and test lossless output — it is often surprisingly competitive in size for flat-colour content.

Getting Sign-Off on a Quality Target

Choosing quality settings is partly a people problem: designers and merchandisers own how images look, while engineers own how much they weigh. A short review session settles it. Pick ten representative images per content type, encode each at four or five settings, and show them side by side at real display size on the devices your customers use, with file sizes hidden. Ask reviewers to mark the lowest version they would accept. Then translate their choice into a perceptual score, and make that score — not the encoder number — the agreed standard, written into the pipeline's configuration with a comment linking to the review. When encoders change or a new content type appears, re-run the search against the same score instead of reopening the debate. This turns a subjective argument into a repeatable check that a CI job can enforce on a sample of new uploads.

Common Mistakes

  • Comparing quality numbers across formats. WebP 75 and AVIF 75 are unrelated scales.
  • Judging on a 1x display. Review on high-DPR screens where artefacts are visible.
  • Using fastest AVIF settings in production builds. Larger files for the same quality.
  • Never revisiting settings. Encoders improve; re-tune occasionally.

Edge Cases

Chroma subsampling. 4:2:0 saves bytes but blurs saturated edges; use 4:4:4 for graphics with coloured text.

Gradients and banding. Smooth gradients band at low quality; slightly higher quality or dithering helps.

Image CDNs with auto quality. Some CDNs choose quality per image automatically using perceptual heuristics; validate their output against your target.

Thumbnails vs full size. Small thumbnails tolerate lower quality than zoomable product images; tune per usage.

FAQ

What AVIF quality should I use?

There is no single answer; it depends on content and encoder. As a starting point with sharp, 50–60 for photos and higher (or lossless) for screenshots, then validate with a perceptual metric.

Is SSIM or PSNR good enough?

They correlate poorly with human judgement for modern codecs. SSIMULACRA2 and Butteraugli were designed for this purpose and are more reliable for choosing settings.

Should WebP be tuned too if AVIF is served first?

Yes — WebP is the fallback for browsers or contexts without AVIF and may be served to a meaningful share of users. Tune it to the same perceptual target.

Does higher encoder effort affect decode time?

Generally no; effort affects encoding search, not decoder complexity. Decode cost depends mainly on resolution and format features.

How often should I re-tune?

When you change encoders or versions, add new content types, or change design standards — and otherwise perhaps yearly.

Can quality differ for mobile and desktop?

You can serve slightly lower quality at higher DPRs (pixels are smaller, artefacts less visible) — some CDNs do this automatically. Validate on real devices.