Skip to content

Configuration

Everything is environment variables, read once at startup. The shared rule is fail-closed: any variable that is set but unparseable or out of range refuses to boot with a message naming it — a typo’d limit never silently falls back to a default.

Variable Default Meaning
PORT 8081 Listen port (0 = OS-assigned, printed on stderr)
OXIMG_BIND 0.0.0.0 Listen address. oximg-ctl auto-spawn sets 127.0.0.1 so a local proof does not publish IMAGES_DIR on every interface
IMAGES_DIR ./images Local source directory (when no source URL is set)
OXIMG_OPTIONS_PREFIX unset Mounts the Cloudflare-style options route at this prefix (e.g. /image, /cdn-cgi/image)
OXIMG_KEY / OXIMG_SALT unset Hex HMAC key/salt; setting both requires signed URLs
OXIMG_WORKERS observed parallelism Pins the CPU permit count (1-512). The default is right almost everywhere — on quota-scheduled platforms like Cloud Run, “pinning to the billed number” measured 17-36% slower (issue #10). The knob exists for noisy-neighbor hosts, tail-latency-over-throughput shapes, and platforms where observed parallelism is unrelated to what is available. The remote-source reason to raise it is gone (issue #22): permits are no longer held across the origin fetch, so fetch/process no longer names throughput that extra permits would recover — the earlier guidance (“above the CPU count is often right for remote sources”, with its permits x (1 - fetch/process) saturation arithmetic) applied to 0.8.x and earlier. Permits are still bounded by memory, not just CPU: (memory limit - idle RSS) / decoded-bytes p99 (see OXIMG_MAX_DECODED_BYTES) is the other ceiling, and it is the binding one on small pods with a heavy decode tail. Verify with the oximg_cpu_workers gauge, which is the only way to know what a given deployment actually got. What “observed” observes is worth knowing — see below
OXIMG_FETCH_CONCURRENCY 4 x permits, max 256 Bounds concurrent origin downloads (1-1024). Fetches hold no CPU permit (issue #22), so they need their own bound: the buffered-source memory hazard is this knob times OXIMG_MAX_SOURCE_BYTES at worst case. The default absorbs an 8-wide srcset burst per permit at production-like fetch shares; raise it when the origin RTT is large relative to per-request CPU work (many fetches must overlap to keep one core fed) and the sources are known-small
OXIMG_LOG error error = one stderr line per failure; request also logs successes. The RUST_LOG names work too, case-insensitively: warn = error; info/debug/trace = request. Like OXIMG_AUTO_FORMAT’s unknown tokens, an unknown value warns rather than refusing to boot, and logs failures only, since verbosity cannot make output wrong
OXIMG_METRICS 0 1 serves Prometheus text at /metrics: requests by status class and resolved format, upstream outcomes (timeout distinct from fault — and note that rejected reading zero is itself the signal in gs:// mode: an over-length key is refused locally, so the store is never asked and the request lands in not_found. If rejected ever moves there, the store refused something, which is a different event), duration histograms split into remote-source fetch (everything between “ready to fetch” and “source in hand” — fetch-slot wait plus the whole download — none of it holding a CPU permit since issue #22), CPU-permit queue wait, and processing (the permit’s actual hold). fetch/process therefore no longer names recoverable throughput; it names the wait the permit no longer pays for. Read fetch numbers from warm traffic, since a fresh process pays connection and TLS setup and reads high for its first requests. Permit/coalescing gauges included. Outside the signing scheme — expose it to your scrape network only
Variable Default Meaning
OXIMG_SOURCE_BASE_URL unset https://… or gs://bucket[/prefix] (see Serving)
OXIMG_GCS_ENDPOINT https://storage.googleapis.com Override for Private Service Connect or emulators; GCE_METADATA_HOST is honored the same way for the token source
OXIMG_UPSTREAM_TIMEOUT 30 Seconds for the whole origin fetch — bounds how long a stalled upstream can hold a fetch slot (and its buffer); timeouts answer 504, distinct from other upstream failures’ 502
OXIMG_UPSTREAM_CONNECT_TIMEOUT 5 Seconds to establish the origin connection
OXIMG_MAX_SOURCE_BYTES 64 MiB Compressed-size cap; over-limit remote sources answer 413
OXIMG_MAX_SRC_PIXELS 64,000,000 Cheap sanity guard on source dimensions, enforced after each format’s header parse; over-cap sources answer 413. Not a memory budget — see the next row
OXIMG_MAX_DECODED_BYTES unset Cap on what a single decode is estimated to allocate, in bytes — the unit a container limit is in. Source pixels cannot be mapped to memory here: cost per pixel varies ~16x with the encoding, because baseline JPEG decodes through DCT shrink-on-load (cost tracks the output) while progressive JPEG buffers whole-image coefficients and PNG/AVIF decode full frames (cost tracks the source), and CMYK stages four channels. The estimate models the buffers the code actually holds at once: the decoder’s frame, the linear-light resize input (the same frame as u16), the output-side dst16+out8, progressive JPEG’s coefficient arrays, and the compressed source where a format needs it whole. Field-validated at 1.2-1.8x above measured peaks across four real sources — deliberately conservative, since under-estimating is what gets a container OOM-killed while the cap reports itself satisfied. Encode-side buffers are still excluded. Over-cap sources answer 413. The response body is deliberately generic across all three source caps (it would otherwise hand clients the configured limits); the stderr line names which limit was hit and the estimated figure, so that is where to look when calibrating. Unset (the default) still computes and exposes the estimate as the oximg_decoded_bytes_estimate histogram, so a cap can be read off a real corpus before being enforced
OXIMG_LOG_DECODED_BYTES_ABOVE unset Report any decode whose estimate exceeds this — filename and per-term breakdown to stderr — and serve it normally. Orthogonal to the cap: the cap refuses and names what it refused, this names without refusing. That distinction is what makes a cap settable: a cap high enough to be safe names nothing, and one set at the tail buys names by refusing live traffic. The histogram tells you that a request cost 512 MiB; this tells you which image. Setting only this is the natural first step for a new deployment — learn the corpus, then choose the cap. Applies to the CLI too
Variable Default Meaning
QUALITY 80 JPEG quality
PRESET jpegli fast = mozjpeg baseline, small = mozjpeg trellis+progressive
OXIMG_JPEG_PROGRESSIVE 1 0 = baseline jpegli: a few percent larger output for lower latency; with OXIMG_OVERLAP this is the speed profile (~-13% single-request latency, ~+9% saturated throughput)
OXIMG_WEBP_QUALITY 75 WebP quality
OXIMG_WEBP_EFFORT 2 libwebp method
OXIMG_AVIF_QUALITY 55 AVIF quality (libavif semantics; chosen by operating point, see bench/quality/QUALITY.md)
OXIMG_AVIF_ALPHA_QUALITY color quality Alpha-plane quality
OXIMG_AVIF_SPEED 8 SVT preset; 9 trades ~-0.6 SSIMULACRA2 at unchanged bytes for ~28% less encode CPU
OXIMG_PNG_EFFORT path-dependent fastest/fast/balanced/high, or a zlib-style 0-9: 0-1 = fastest, 2-5 = fast, 6-8 = balanced (zlib’s default 6), 9 = high (zlib’s best). Unset resolves to fast for lossless output and balanced for quantized output, where effort matters ~2x more; setting it pins one level for both. Like OXIMG_LOG, an unknown value warns rather than refusing to boot, and encodes as if unset, since effort changes size and speed, never what is produced
OXIMG_PNG_QUANTIZE 0 1 palette-quantizes opaque PNG output (Wu + Floyd–Steinberg): typically ~3x smaller photographic PNGs at the quantized balanced default (about half that if effort is forced fast), near-exact on flat graphics. Opt-in because quality loss on a lossless format must be deliberate; alpha sources always encode lossless RGBA and ignore this knob
OXIMG_PNG_QUANTIZE_COLORS 256 Palette size, 2-256; 64 trades visible-on-inspection banding for another ~15%
OXIMG_AUTO_FORMAT unset Comma-separated Accept-negotiation preference list (e.g. avif,webp); see the ordering guidance under Supported formats
OXIMG_FLATTEN_BG ffffff Background for alpha → JPEG flattening

On a Linux container the count comes from the cgroup CPU quota and the process’s CPU affinity, whichever is smaller, floored at 1 — measured on cgroup v2 (workers is the oximg_cpu_workers gauge):

container CPU config cgroup workers
no limit cpu.max: max host CPU count
1 CPU quota 1.0 1
1.5 CPU (1500m) quota 1.5 1
1.9 CPU (1900m) quota 1.9 1
2 CPU quota 2.0 2
2.5 CPU (2500m) quota 2.5 2
0.5 CPU (500m) quota 0.5 1 (the floor)
CPU shares only, no quota cpu.weight set, cpu.max: max host CPU count
pinned to 2 cores affinity 0-1, no quota 2
pinned to 3 cores + quota 1 both 1 (the smaller)

Three consequences that catch people out:

  • On Kubernetes this is limits.cpu. requests.cpu has no effect: it becomes cpu.weight, a scheduling share with no count in it, so there is nothing there to observe. A limits.cpu set as a blast-radius guard silently becomes a concurrency decision.
  • Fractional limits round down. limits.cpu: 1500m yields the same single permit as 1000m while costing 50% more, and 1900m is still
    1. Whole numbers are the only way to buy concurrency — the second permit arrives at 2, not at 1001m.
  • A pod with only requests.cpu and no limit is the dangerous shape: it sizes itself to the node, so on a 64-core node it will admit 64 concurrent decodes while being scheduled for a fraction of one core — and since peak memory is permits x per-request decode cost (OXIMG_MAX_DECODED_BYTES), that presents as unexplained memory pressure rather than as queue latency, with nothing in the pod spec looking wrong. oximg prints a startup note when it takes its permit count from full host parallelism with no CPU quota visible; a deliberately restricted cpuset (Kubernetes’ static CPU-manager policy) is not warned about, because that count is correct.
  • Platforms without a hard quota fall back to host parallelism, which is why an equivalently-sized container reports a different number on Cloud Run (cpu: "1" there observes 2 — see the Cloud Run guide, where pinning it down measured slower).

An animated GIF into a WebP target is served as an animated WebP; every other target renders its first frame. The knobs below bound what one such request may cost — and unlike the source caps above, exceeding one is not an error: the request degrades to the still first frame and still answers 200. That is the point. An animation is not one image but N, so a single request can cost hundreds of still requests’ CPU (measured: 3.2 s for the worst file in a 15-file corpus, against 5 ms for a still — docs/gif-evaluation.md §5), and a service that answers a smaller thing beats one that answers 413 to an image a browser will happily display.

Variable Default Meaning
OXIMG_GIF_ANIMATION 1 0 renders every animated GIF as its still first frame, i.e. 0.11.0 behaviour
OXIMG_MAX_ANIM_FRAMES 200 Source frames an animation may carry before it degrades to a still. Bounds the decode+composite half of the work, which shrinking the output does not touch
OXIMG_MAX_ANIM_WORK 8,000,000 Encoded frames x post-resize frame area, in pixels — the product that predicts encode time, which is what dominates an animation (3018 ms of a 3213 ms worst case). The default admits the measured corpus’ worst in-budget file (26 frames of 1280x720 into a 512 box, ~6.8 Mpx, 517 ms) and refuses its worst overall (~34 Mpx, 3.2 s). Because it is measured after the resize, a small output box is what buys an expensive source back: the same source that is refused at native size fits into a thumbnail
OXIMG_ANIM_FRAME_STEP 1 Encode every Nth frame. Total play time is preserved (a skipped frame extends the previous frame’s duration), so this costs smoothness, not fidelity — which is why it is off by default. 2 roughly halves encode cost

One caveat on bytes: animated WebP wins hugely on photographic and video-like GIFs (4–30% of the source across the corpus) but can lose on small flat-graphics ones, where GIF’s tiny palette is exactly what LZW compresses best — a measured 25 KB, 8-frame cartoon comes back at 41 KB. oximg re-encodes whatever it is asked to, in every format; if a caller has such sources, serving them at their own size gains nothing and OXIMG_GIF_ANIMATION=0 is the cheaper answer.

The estimated-memory cap (OXIMG_MAX_DECODED_BYTES) applies here too, and degrades rather than refusing for the same reason: a still of the same GIF fits under any cap that admits one frame. Frames are composited, resized and handed to the encoder one at a time, so our staging is a function of the canvas alone — but libwebp’s animation encoder retains every frame it has compressed until the container is assembled, so peak memory does grow with the animation. The estimate prices that retained output at one byte per encoded pixel (~9x the 0.11 B/px measured across the corpus, i.e. deliberately pessimistic), which makes it OXIMG_MAX_ANIM_WORK, not the canvas, that bounds it.

Variable Default Meaning
OXIMG_AUTO_ROTATE 1 0 serves the stored orientation
OXIMG_ICC 1 0 strips source ICC profiles and converts CMYK naively instead of through their profile
OXIMG_RESIZE linear srgb resizes in sRGB space instead of linear light
OXIMG_RESIZE_BACKEND kernel fir selects the portable fast_image_resize convolution instead of the platform SIMD kernel
OXIMG_OVERLAP auto JPEG decode fused with resize+encode on a second thread (~-20% single-request latency); auto fuses while 2 x active requests <= visible CPUs. Bytes are identical either way
OXIMG_PAR 1 Resize threads per request
OXIMG_DCT_MARGIN unset Decode-size headroom over the target: shrink-on-load, off by default. It is a speed knob. libjpeg’s reduced IDCT is erratic per scale, so the cost is not graded: on a 5.3x downscale the 1.7 that used to be the default selects libjpeg’s 3/8 scale, which measured 13.4 SSIMULACRA2 points below a full decode for the same output size and the same bytes, while 5/8 on the same image was optimal — and no single value avoids the bad scales at every ratio (dct_sweep.py). Set it to trade quality for decode time on large sources. The buffered paths ignore the default and keep 1.7 — CMYK/YCCK JPEG and WebP stage a whole frame at the decode size, so there the shrink caps peak RSS (WebP 21.7 MB against 133.9 MB full-size) rather than costing throughput; an explicit value still applies to them
OXIMG_WEBP_DECODE_THREADS 1 0 disables libwebp’s two-thread decode pipelining
OXIMG_AVIF_DECODE_THREADS arch-dependent dav1d workers: 2 on x86-64 (SMT absorbs the second thread), 1 on aarch64
OXIMG_TIMING unset Print per-stage timing lines to stderr