How to Batch-Convert Images to AVIF and WebP with sharp

This guide is the hands-on companion to Serving AVIF and WebP with Fallbacks, within Image & Media Optimization. Most sites that want modern formats already have hundreds or thousands of JPEGs and PNGs in a repository, a CMS export or object storage. Converting them is mechanical but has traps: running out of memory with large batches, losing colour accuracy when stripping profiles, rotating photos wrongly because EXIF orientation was ignored, regenerating unchanged files on every run, and switching references before every variant exists.

sharp — the Node.js binding to libvips — is the standard tool: fast, memory-efficient, and able to resize and encode AVIF, WebP and JPEG in one pipeline. A good conversion script processes images with bounded concurrency, normalises orientation and colour, writes outputs named by content hash, skips work already done, and produces a manifest that templates can use.

Safe batch conversion process Process from scanning source images through bounded conversion and verification to switching references. Safe batch conversion process Scan list sources + hashes Convert bounded concurrency Verify counts, sizes, samples Manifest dims + URLs Switch refs templates use manifest

Rapid Diagnosis

  • Inventory sources. Count images by format and size; very large originals (over 20MP) need care with memory.
  • Check orientation handling. Phone photos often rely on EXIF orientation; outputs must be rotated correctly.
  • Check colour profiles. Images in Adobe RGB or Display P3 can shift colour when converted carelessly.
  • Plan storage. Widths × formats multiplies files; estimate storage and upload costs.

Root Cause Analysis: What Goes Wrong in Naive Conversions

1. Unbounded parallelism. Launching thousands of conversions at once exhausts memory and crashes.

2. Orientation loss. Stripping metadata without applying EXIF rotation produces sideways images.

3. Colour shifts. Dropping embedded profiles without converting to sRGB changes colours.

4. Reprocessing everything. Without caching, every run re-encodes all images, making AVIF conversion prohibitively slow.

Library conversion results (12,000 product images) Bar chart of total storage for an image library before and after conversion to AVIF and WebP variants. Library conversion results (12,000 product images) Original JPEGs (one size) 18.4GB AVIF variants (5 widths) 4.1GB WebP variants (5 widths) 6.3GB JPEG fallbacks (5 widths) 9.8GB

Step-by-Step Resolution

1. Write a bounded, cached conversion script

javascript
// scripts/convert.mjs
import sharp from 'sharp';
import pLimit from 'p-limit';
import { glob } from 'node:fs/promises';
import { readFile, mkdir, access, writeFile } from 'node:fs/promises';
import { createHash } from 'node:crypto';

sharp.concurrency(1);                       // libvips threads per image; parallelism comes from pLimit
const limit = pLimit(4);                    // ~ number of CPU cores
const WIDTHS = [480, 800, 1200, 1600];
const manifest = {};

async function convert(src) {
  const buf = await readFile(src);
  const hash = createHash('sha256').update(buf).digest('hex').slice(0, 12);
  const base = sharp(buf, { failOn: 'truncated' }).rotate().toColourspace('srgb');   // apply EXIF orientation
  const meta = await base.metadata();
  const outputs = [];
  for (const w of WIDTHS.filter((w) => w <= meta.width)) {
    for (const [ext, fmt] of [['avif', { quality: 55, effort: 4 }], ['webp', { quality: 78 }]]) {
      const out = `out/${hash}-${w}.${ext}`;
      try { await access(out); } catch { await base.clone().resize(w).toFormat(ext, fmt).toFile(out); }
      outputs.push(out);
    }
  }
  manifest[src] = { hash, width: meta.width, height: meta.height, outputs };
}
for await (const f of glob('src/images/**/*.{jpg,jpeg,png}')) await limit(() => convert(f));
await writeFile('out/manifest.json', JSON.stringify(manifest, null, 2));
// trade-off: the loop awaits each limit() call, serialising scheduling; for very
// large libraries collect promises and await Promise.all in batches instead.

2. Verify before switching references

Check that every source has the expected outputs, compare total sizes, and visually spot-check a random sample (including portrait phone photos and images with embedded profiles).

3. Switch references through the manifest

Templates read the manifest to emit <picture> markup; images missing from the manifest keep their original references. Never rename or delete originals until all references have moved.

4. Upload with immutable caching

Upload outputs with Cache-Control: public, max-age=31536000, immutable and correct Content-Type (image/avif, image/webp).

Conversion safety checklist Checks that make a batch image conversion safe to roll out. Conversion safety checklist Bounded concurrency Avoid memory exhaustion on large libraries rotate() and sRGB conversion Correct orientation and colour Content-hash names and skip-if-exists Fast re-runs; immutable caching Verify then switch via manifest No broken references mid-rollout 1 2 3 4

Verification

Run the script twice: the second run should finish in seconds (everything cached). Compare a sample of outputs side by side with originals at 100% on a high-DPR display. After switching references, crawl key templates and check for 404s on image URLs, then compare image bytes per page view in RUM.

Worked Example: A Marketplace Catalogue

A marketplace converted 12,000 product images (18GB of single-size JPEGs). A first naive script with Promise.all over all files crashed after a few hundred images. The bounded script above (four concurrent images on an 8-core runner) converted the library in about 70 minutes; re-runs on new uploads took seconds. Sixty phone photos had relied on EXIF orientation — rotate() fixed them automatically. After switching templates to the manifest, average product page image weight dropped by 72%, and LCP p75 on mobile improved by 650ms.

Converting at the Edge Instead

For very large or constantly changing libraries (user uploads, marketplaces with millions of images), batch conversion may be impractical. Image CDNs and self-hosted proxies such as imgproxy convert on first request and cache the result, removing the batch step entirely. A hybrid is common: batch-convert the editorial and site images you control, and let an image proxy handle user content. See self-hosting an image proxy with imgproxy.

Running the Conversion Incrementally

Large libraries rarely convert cleanly in one pass. Plan for interruptions: because outputs are named by content hash and skipped when they already exist, a crashed or cancelled run can simply be restarted and continues where it stopped. Write failures to a separate log (corrupt files, unsupported colour spaces, images over the pixel limit) rather than aborting the batch, then fix those files by hand. Roll references out section by section — one category or template at a time — and watch error rates and image 404s before moving on. Keep the manifest under version control or in object storage with history, so a bad batch can be rolled back by restoring the previous manifest rather than deleting files. Finally, wire the same script into the upload path or a nightly job, so new images are converted as they arrive and the library never drifts back to unoptimised originals.

Common Mistakes

  • Unbounded Promise.all. Memory blow-ups on big libraries.
  • Forgetting rotate(). Sideways photos.
  • Stripping profiles without converting. Colour shifts.
  • Switching references before conversion completes. Broken images.

Edge Cases

Animated GIFs. Handle separately (convert to video or animated formats); still-image settings will flatten them.

Transparent PNGs. Keep alpha in AVIF/WebP; use PNG, not JPEG, for fallbacks.

Huge source images. Set limitInputPixels appropriately; extremely large images may need more memory or pre-scaling.

HEIC sources. sharp's HEIC support depends on build options; convert HEIC uploads server-side with a supported toolchain.

FAQ

How fast is sharp for AVIF?

Much slower than JPEG or WebP — often hundreds of milliseconds to seconds per large image depending on effort. Caching by content hash and parallelism across cores make large conversions practical.

Should I keep the originals?

Yes. Keep originals as the source of truth so you can regenerate with new settings or formats later.

Does sharp preserve metadata?

By default it strips most metadata, which saves bytes. Use withMetadata() (or newer per-field options) if you must keep copyright or colour profile data.

Can I run this in CI?

Yes, with output caching between runs (object storage or a CI cache keyed by content hashes). Without caching, AVIF conversion of a large library on every build is too slow.

What concurrency should I use?

Roughly the number of CPU cores, with sharp.concurrency(1) so libvips does not also multithread each image. Monitor memory for very large sources.

How do I handle images referenced in Markdown or a CMS?

Rewrite references at render time using the manifest (a Markdown plugin or CMS field formatter), so content authors keep referencing originals and templates emit optimised markup.