Pipeline
source bytes (local file or HTTP origin) → format sniff → decode JPEG: mozjpeg streaming decode at full size (shrink-on-load is opt-in via OXIMG_DCT_MARGIN; CMYK/YCCK keep it at 1.7x, where it caps a staged frame instead of costing quality) PNG: png crate (palette/gray/16-bit normalized to RGB(A)8) WebP: libwebp AVIF: dav1d (8/10/12-bit, all subsamplings, alpha, bilinear chroma upsampling) GIF: gif crate, frames composited onto the logical screen (every frame for an animated GIF into a WebP target, else the first) → linear-light resize: sRGB u8 → linear u16 → Lanczos3 → sRGB u8 (alpha is premultiplied before resampling, unpremultiplied after; JPEG rows stream through in-tree ring-scheduled f32 row kernels — AVX2 on x86-64, NEON on aarch64, both verified against an f64 reference — optionally fused with the decode on a second thread; other formats resize full-frame: pic-scale on x86-64, the same in-tree kernel on aarch64) → encode in the source format (GIF sources: WebP, the default target) JPEG: jpegli, progressive (PRESET=fast / PRESET=small select mozjpeg profiles) PNG: png crate | WebP: libwebp | AVIF: SVT-AV1 (10-bit 4:2:0, tune=ssim) animated GIF: each frame composited, resized and streamed into libwebp's animation encoder (one canvas in memory, not N)Concurrent identical requests are coalesced and share one result. CPU concurrency is pinned to the core count with a semaphore; the HTTP layer (axum/tokio) only does queueing and IO.