How to Choose AVIF and WebP Quality Settings
This guide refines Serving AVIF and WebP with Fallbacks, within Image & Media Optimization. Switching to AVIF or WebP saves bytes only if the quality settings are chosen well. Too high, and modern formats lose much of their advantage over JPEG; too low, and artefacts appear — blurred textures, banding in gradients, smeared text. Most teams pick a number once ("AVIF 50, WebP 75") and apply it to everything, which over-compresses detailed photos and under-compresses simple graphics.
Quality numbers are not comparable across formats or even encoders, and the same number produces very different visual quality on a product shot, a screenshot and a sunset. The reliable approach is to target a perceptual quality level, measured with metrics designed to match human judgement, and find the lowest setting per format and content type that reaches it.
Rapid Diagnosis
- Check your current settings in the build pipeline or image CDN configuration; note whether one value applies to everything.
- Spot-check outputs at 100% zoom on a high-DPR display: textures (fabric, foliage), gradients (skies) and text (screenshots, labels).
- Compare sizes across content types. If screenshots and illustrations are as large as photos, settings are not tuned per type.
- Check encoder speed settings. Very fast AVIF settings produce larger files at the same quality.
Root Cause Analysis
1. One global quality value. Different content needs different settings to reach the same perceived quality.
2. Quality numbers are encoder-specific. "Quality 60" in one AVIF encoder is not the same as in another, or as WebP 60.
3. Eyeballing on low-DPI screens. Artefacts invisible on a laptop at 1x can be obvious on phones at 3x zoom levels.
4. Ignoring encoder effort. Faster AVIF speed settings trade compression efficiency for build time.
Step-by-Step Resolution
1. Pick a perceptual metric and target
Use SSIMULACRA2 or Butteraugli (both designed to correlate with human perception). A SSIMULACRA2 score around 70–80 is often considered high quality for web photos; choose a target with design stakeholders by reviewing samples.
2. Search for the lowest setting that meets the target
// scripts/tune-quality.mjs — binary search AVIF quality to reach a SSIMULACRA2 target.
import sharp from 'sharp';
import { execFileSync } from 'node:child_process';
async function score(refPath, buf) {
await sharp(buf).png().toFile('/tmp/candidate.png');
return Number(execFileSync('ssimulacra2', [refPath, '/tmp/candidate.png']).toString().trim());
}
export async function tune(refPath, target = 75) {
let lo = 20, hi = 90, best = hi;
while (lo <= hi) {
const q = Math.floor((lo + hi) / 2);
const buf = await sharp(refPath).avif({ quality: q, effort: 4 }).toBuffer();
if (await score(refPath, buf) >= target) { best = q; hi = q - 1; } else { lo = q + 1; }
}
return best;
}
// trade-off: per-image tuning is expensive; run it on a representative sample
// per content type to choose settings, rather than for every image in the build.
3. Apply settings per content type
Classify images (by folder, CMS field or tag) and apply the tuned setting per class in the pipeline or as image CDN presets.
4. Choose encoder effort for your build budget
For AVIF in sharp, effort 4–6 is a common balance; higher effort shrinks files a few percent more at much longer encode times. Cache outputs so the cost is paid once per image.
Verification
Compare total image bytes per page before and after, review a sample visually at high zoom with stakeholders, and keep the perceptual score distribution for each class above the target. In RUM, LCP image resourceLoadDuration should fall in proportion to the byte savings.
Worked Example: An Online Furniture Store
An online furniture store used AVIF quality 40 for everything. Fabric textures on sofas looked blurry on high-DPR phones and returns mentioning "colour/texture not as pictured" rose. A tuning pass with SSIMULACRA2 found product photos needed quality 58 to meet the agreed target, while lifestyle backgrounds passed at 46 and line-drawing diagrams at 35. Applying per-class presets increased product image bytes by 18% but reduced backgrounds and diagrams by 22% and 30%; total image weight per product page was roughly unchanged, while texture complaints dropped noticeably.
Lossless and Near-Lossless Cases
Some content should not be lossy at all. Screenshots of UI, diagrams with thin lines, and images containing small text often look worse with lossy AVIF than with lossless WebP or PNG at a similar size. WebP's lossless mode and AVIF's lossless or high-quality modes with 4:4:4 chroma subsampling preserve sharp coloured edges that 4:2:0 subsampling blurs. Classify such images separately and test lossless output — it is often surprisingly competitive in size for flat-colour content.
Getting Sign-Off on a Quality Target
Choosing quality settings is partly a people problem: designers and merchandisers own how images look, while engineers own how much they weigh. A short review session settles it. Pick ten representative images per content type, encode each at four or five settings, and show them side by side at real display size on the devices your customers use, with file sizes hidden. Ask reviewers to mark the lowest version they would accept. Then translate their choice into a perceptual score, and make that score — not the encoder number — the agreed standard, written into the pipeline's configuration with a comment linking to the review. When encoders change or a new content type appears, re-run the search against the same score instead of reopening the debate. This turns a subjective argument into a repeatable check that a CI job can enforce on a sample of new uploads.
Common Mistakes
- Comparing quality numbers across formats. WebP 75 and AVIF 75 are unrelated scales.
- Judging on a 1x display. Review on high-DPR screens where artefacts are visible.
- Using fastest AVIF settings in production builds. Larger files for the same quality.
- Never revisiting settings. Encoders improve; re-tune occasionally.
Edge Cases
Chroma subsampling. 4:2:0 saves bytes but blurs saturated edges; use 4:4:4 for graphics with coloured text.
Gradients and banding. Smooth gradients band at low quality; slightly higher quality or dithering helps.
Image CDNs with auto quality. Some CDNs choose quality per image automatically using perceptual heuristics; validate their output against your target.
Thumbnails vs full size. Small thumbnails tolerate lower quality than zoomable product images; tune per usage.
FAQ
What AVIF quality should I use?
There is no single answer; it depends on content and encoder. As a starting point with sharp, 50–60 for photos and higher (or lossless) for screenshots, then validate with a perceptual metric.
Is SSIM or PSNR good enough?
They correlate poorly with human judgement for modern codecs. SSIMULACRA2 and Butteraugli were designed for this purpose and are more reliable for choosing settings.
Should WebP be tuned too if AVIF is served first?
Yes — WebP is the fallback for browsers or contexts without AVIF and may be served to a meaningful share of users. Tune it to the same perceptual target.
Does higher encoder effort affect decode time?
Generally no; effort affects encoding search, not decoder complexity. Decode cost depends mainly on resolution and format features.
How often should I re-tune?
When you change encoders or versions, add new content types, or change design standards — and otherwise perhaps yearly.
Can quality differ for mobile and desktop?
You can serve slightly lower quality at higher DPRs (pixels are smaller, artefacts less visible) — some CDNs do this automatically. Validate on real devices.
Related
- AVIF vs WebP: which format to serve — format choice before quality choice.
- Batch-converting images with sharp — applying settings at scale.
- Fixing blurry images on high-DPI displays — resolution vs quality.