← PIXL EngineDeep dive · Codecs

Encode and decode

Nine input formats, eleven ways out, and not one default. PIXL Engine decodes with proven libraries and its own readers, encodes with exactly the parameters it is given, and reports what it wrote. Where the bytes and the request disagree, it says so instead of guessing.

The format matrix

Every pixel format converts to every other. RAW converts to all of them and to DNG. Nothing outputs RAW: a raw DNG is written from a RAW source only, and a linear DNG from any rendered frame in linear light.

SourceDecoderWhat it keeps
JPEGlibjpeg-turbo (SIMD)greyscale stays one channel
PNGpng 0.18native colour type, 8 or 16 bits, cICP
HEIF / HEIC / AVIFlibheif; subsampled colour made by PIXL8/10/12 bits, RGB/RGBA, the nclx box
JPEG XLlibjxl through PIXL’s own bindings8/16-bit integer and 32-bit float, 1–4 channels
TIFFtiff 0.118/16-bit integer and 32-bit float; EXIF lifted out of the directory
WebPlibwebp8 bits; an animation is refused, not cut to its first frame
RAW (CR2, CR3, ARW, NEF, RAF, RW2, ORF …)LibRaw 0.22.2 with PIXL’s shimthe mosaic, or developed on request
DNGPIXL’s own reader, DNG 1.7.1every compression DNG defines
Master cachethe JPEG XL reader and a pixl boxa render resumed where it stopped (HDR)
SinkEncoderAccepts
JPEG XL lossy / losslesslibjxl1–4 channels, 8 or 16 bits; lossless also 32-bit float, read back bit for bit
JPEG XL repacklibjxl JPEG reconstructiona JPEG source, no pixel work
JPEG from JPEG XLlibjxl JPEG reconstructiona JXL that carries reconstruction data
JPEGlibjpeg-turbo8 bits, 1 or 3 channels
PNGpng 0.188 or 16 bits, 1–4 channels
AVIFlibheif and aom8/10/12 bits, 3 or 4 channels
TIFFtiff with LZW / Deflate8 or 16 bits or 32-bit float, 1–4 channels, EXIF as a real sub-IFD
WebPlibwebp8 bits, 3 or 4 channels, up to 16 383 px a side
DNGPIXL’s own writera RAW source
Linear DNGPIXL’s own writerthe rendered frame, 16-bit or float, in linear light
Raw pixelsnone: little-endian samples, rows packedU8, U16, F16 or F32, with the colour description in the report

HEIC is read, through libde265, and never written: AVIF is the one HEIF-family export, so no GPL encoder ships in a host application. A buffer an encoder cannot take is refused with the pixel field that would fix it. The engine will not drop an alpha channel or halve a bit depth to make a request succeed.

Every parameter is the caller’s

The request names the input format, and the decoder it names is the one that runs. If the bytes say otherwise the answer is a FormatMismatch, not a quiet reroute by file extension. On the way out, every encoder parameter is stated by the caller, with no default and no “auto”. A missing or out-of-range value is refused with the field’s name.

Quality and effort are different axes. Quality (or distance) decides how much detail survives; effort and speed decide only how long the encoder works to reach that quality. The best detail is distance 0, quality 100 or lossless — never effort 1.

CodecQuality axisEffort / speed axis
JPEG XLdistance 0–25 (0 lossless, about 1 visually lossless)effort 1–9
JPEGquality 1–100optimize (Huffman)
AVIFquality 1–100 under a named tune, or losslessspeed 0–9, encoder threads 1–64
WebPquality 1–100, or losslessmethod 0–6
PNG, TIFFlossless alwayscompression
DNGlossless alwayslossless-JPEG predictor 1–7

Combinations that cannot be what they claim are refused rather than produced. Lossless AVIF with subsampled chroma is refused, because subsampling discards chroma. Lossless AVIF is lossless of the decoded pixels only with the identity matrix; under any other matrix RGB is rounded to Y/Cb/Cr before the codec, and the report’sloss.lossy_encoder says true. Every report also counts how many times the pixels were rounded between decode and encode.

Matching a source’s own encoding is a decision too, so it lives outside convert. A separate pure call,suggest_encode, returns an ordinary encode request from what probe read. It matches every parameter that affects the pixels and never invents one the file does not state. PNG and TIFF match exactly; a JPEG matches when its quantisation tables are the standard tables scaled for some quality. HEIC, AVIF, lossy WebP and lossy JPEG XL do not record their quality, so the call refuses, naming the field to set.

AVIF, measured

AVIF is written through libheif’s C API with the aom encoder. Every knob the AV1 encoder has is a required field: quality or lossless, bit depth, chroma, speed, the YCbCr matrix, the encoder’s threads, the tune and the tiling. The YCbCr matrix is the caller’s, and one structure is both the matrix libheif converts with and thenclx box a decoder converts back with, so the two cannot disagree.

Three tunings, three scales

Psnr, Ssim and Iq (aom’s still-image tuning) put quality on different scales, so a preset is a pair of tune and quality. The tunings were compared over 800 encodes, scored with SSIMULACRA2 (higher is better) and Butteraugli’s 3-norm (lower is better). Two metrics are used because aom tuned Iq against SSIMULACRA2 itself.

Size for quality, 24 MP photograph, 4:2:0File size against SSIMULACRA2 for AVIF's three tunings and x265 HEIC on the 24 MP photograph. AVIF Iq reaches each score up to about 88 with the smallest file; x265 HEIC reaches higher scores at larger sizes.Size for quality, 24 MP photograph, 4:2:04050607080900246810121416file size, MBSSIMULACRA2AVIF PsnrAVIF SsimAVIF Iqx265 HEIC
Quality 30–95 in steps of 5 (points below a score of 40 left out). The 24 MP road photograph (6000 × 4000, P3, 8-bit JPEG source), 4:2:0, speed 6, 8 encoder threads; libheif 1.21.2 with aom 3.13.3, and x265 4.1 through the same libheif (preset slow, tune ssim), measured before the HEIC sink was removed. Scores from libjxl 0.11.2’s ssimulacra2 on PIXL’s decode against the source PNG. i7-14700F, encodes pinned to eight performance cores. Rows: docs/avif/avif-quality.csv, heic-x265.csv.

Read at equal quality, the same rows give the size each encoder needs:

Input, SSIMULACRA2PsnrSsimIqx265 HEIC
24 MP photo, 702.71 MB2.54 MB2.22 MB2.33 MB
24 MP photo, 804.37 MB4.33 MB3.79 MB4.04 MB
24 MP photo, 855.57 MB5.53 MB5.14 MB5.50 MB
24 MP, 16-bit → 10-bit, 856.01 MB5.88 MB5.16 MB5.76 MB
4 MP screen capture, 850.22 MB0.21 MB0.18 MB0.22 MB
PQ fixture, 8049 kB48 kB41 kB49 kB
HLG fixture, 8044 kB42 kB37 kB45 kB
Input, Butteraugli 3-normPsnrSsimIqx265 HEIC
24 MP photo, 1.22.64 MB2.47 MB2.40 MB2.25 MB
24 MP photo, 0.84.34 MB4.27 MB3.98 MB3.85 MB
24 MP photo, 0.47.11 MB6.90 MB6.61 MB6.65 MB
4 MP screen capture, 0.80.10 MB92 kB96 kB0.12 MB
  • Iq is the smallest at a given quality on photographs, by 7–13% on SSIMULACRA2 or 3–7% on Butteraugli. Psnr is never the smallest.
  • Against x265 HEIC the two metrics disagree on photographs. At equal SSIMULACRA2, AVIF Iq is 4–10% smaller; at equal Butteraugli it is level, from 7% larger at low quality to 1% smaller at high. On the screen capture and the HDR fixtures it is 10–20% smaller. The fair summary is “about the size of HEVC on photographs, smaller on graphics”, in 1.3–2.3 s against 2.0–4.0 s and about 640 MB of memory against 820 MB.
  • x265’s scale reaches higher: HEIC at quality 80–95 scores 90.4–91.5 in 9.7–15.6 MB; AVIF at 95 stops at 89.7 (90.7 at 4:4:4).
  • 4:4:4 costs 2–3% more bytes and 44% more memory on the photograph for about one point.

Speed

Speed (Iq, quality 70)SizeEncodeSSIMULACRA2
04.15 MB71.49 s82.5
34.16 MB18.56 s82.3
64.30 MB1.90 s82.0
94.50 MB0.45 s81.3

At equal score speed 9 is 9–10% larger than speed 6; speed 3 is 3–7% smaller for 5–10 times the time, and speeds 0–2 write nearly the same file as 3. The encoder’s own thread count changes no byte from 2 threads up. At 1 thread libheif leaves aom’s row multithreading off, and the file is a different encode of the same quality.

Grids and 102 megapixels

Grid writes a HEIF grid item over square AV1 tiles — the layout an iPhone’s own HEIC uses, 8 × 6 tiles of 512. Each tile gets an encoder of its own and drops it before the next: one encoder over every tile kept every tile’s codec context alive, and a 24 MP grid then peaked at 914 MB, above the single image’s 650. A single image is refused past 16 384² pixels or 32 768 a side, the limits AV1 readers apply by default. A grid needs an even width (and height, at 4:2:0) under subsampled chroma, as MIAF requires, and the engine refuses to write one that readers would reject.

102 MP GFX100S frame, Iq q80Speed 6Speed 9
one image, 4:4:412.39 MB · 4.72 s · 3858 MB12.73 MB · 1.44 s · 3838 MB
one image, 4:2:012.28 MB · 3.93 s · 2659 MB12.59 MB · 1.09 s · 2640 MB
tiles of 4096, 4:4:4 (9)12.69 MB · 6.46 s · 1196 MB13.17 MB · 2.22 s · 1202 MB
tiles of 2048, 4:4:4 (30)12.62 MB · 7.02 s · 775 MB12.98 MB · 2.13 s · 784 MB
tiles of 512, 4:4:4 (414)12.79 MB · 14.11 s · 692 MB12.93 MB · 2.29 s · 678 MB
Peak memory, 102 MP frame, by layoutPeak resident memory of the AVIF encode of the 102 MP GFX100S frame at 4:4:4, quality 80, speed 6: 3858 MB as one image, 1196 MB in tiles of 4096, 775 MB in tiles of 2048, 732 MB in tiles of 1024 and 692 MB in tiles of 512.Peak memory, 102 MP frame, by layout01000200030004000peak resident memory, MBone image3858 MBtiles of 40969 tiles1196 MBtiles of 204830 tiles775 MBtiles of 1024108 tiles732 MBtiles of 512414 tiles692 MB
The 102 MP GFX100S frame (11 648 × 8735) from an 8-bit PNG, 4:4:4, Iq quality 80, speed 6, 8 encoder threads; i7-14700F, pinned to eight performance cores. From docs/avif/README.md.

Size, encode time, peak memory; 11 648 × 8735 from an 8-bit PNG, 8 encoder threads. A grid is 0.2–2.5% larger and slower, scores the same, and at tiles of 2048 takes a fifth of the memory. The frame’s odd height rules out a 4:2:0 grid, so it is written 4:4:4; as a grid of 4096 at quality 90, speed 9, it reads back at 44.2 dB.

Reading grids has a cost too. libheif keeps a decoder per tile, each starting a thread per core, and past the process’s thread limit its default returns the missing tiles green. PIXL reads strictly: a grid it cannot decode whole is a decode error naming the grid and its tile count, never a different picture.

Read by others

Every file kind the sink writes — sRGB, Display P3 by ICC and by nclx alone, 10 and 12 bits, 4:4:4 and 4:2:0, lossless identity, a grid, alpha, PQ, HLG, one- and three-channel gain maps, 102 MP, 16 384², 32 768 wide — was read back by libavif 1.3.0, libheif 1.23.5 and exiftool 13.50:

  • every file opens in both libraries at their default limits;
  • libheif’s decode is PIXL’s sample for sample at 8 bits, within one code at 10 and 12;
  • libavif’s is within one code on 4:4:4 (on 4:2:0 its own chroma upsampling differs);
  • EXIF, XMP and a preserved ICC profile come out as the source’s bytes;
  • libavif tone maps PIXL’s gain maps as PIXL does: MaxCLL 795 against 797 cd/m² at 2 stops.

Two writer bugs were found this way and fixed, each with a test. On Apple’s devices (macOS 27 Preview, ImageIO and Core Image on an M2 Pro; Safari, Chrome and an iPhone the same) every file opens at the size written, each colour description reads as written, and Core Image’s SDR renders are PIXL’s decodes within 0.3 of an 8-bit code in every channel mean.

JPEG XL

The engine wraps libjxl’s C API directly, with bindings generated once and committed, because the order of calls matters: libjxl wants distance 0 set before the lossless flag, and a box buffer must be released before it is grown. Setting uses_original_profile only for lossless frames matters too: on a lossy frame it costs roughly twice the file size.

  • 8.0 → 6.5 MBa JPEG repacked into JPEG XL
  • byte for bytethe original JPEG reconstructed from it
  • bit for bit32-bit float masters read back

The two bit-exact paths in the engine are inverses. A JPEG repacked into JPEG XL keeps its DCT coefficients and every APP segment inside the reconstruction data, and comes back as the original file. The way back works only when the JXL carries reconstruction data; otherwise it is refused, because re-encoding would be lossy. Everywhere else, “lossless” means lossless of the decoded pixels.

The 19% here is one camera JPEG. The JPEG XL authors measured a set: for JPEGs a camera typically produces, “the gains are around 20%”, with libjpeg-turbo output at quality 70–95 saving 19–21% and JPEGs from mozjpeg and jpegli, which squeeze harder already, about 10–19% (Sneyers et al., 2025, Fig. 53). The engine runs the same libjxl reconstruction.

Lossless JPEG XL also holds IEEE single-precision floats. The engine’s master cache is one: the linear master in PixlRGB with alpha, read back bit for bit — negative zero, subnormals and the largest finite float included — and decoded identically at 1, 7 and 16 threads. A JXL’s own ICC profile is carried only when the file holds one; an enumerated colour encoding (every lossy JXL, every PQ or HLG one) is read as CICP code points, as a PNG’scICP is.

HEIC, read as Apple reads it

A phone’s HEIC is 4:2:0. By default libheif upsamples its chroma by nearest neighbour, and its bilinear option indexes chroma at cx / 2 in its border loops, up to 154 codes off on the iPhone corpus. Core Image upsamples bilinearly from centre-sited chroma with the edge repeated. So PIXL asks libheif for the coded planes and makes the colour itself:

  • Siting is what the stream states (an HEVC VUI chroma location, AV1’schroma_sample_position), else the centre. On every iPhone file the centre lands within 1.27 codes of Core Image; H.265’s default, left-sited, lands 13–28 codes away. probe reports which siting was used and whether it was stated or assumed.
  • The arithmetic is libheif’s own: its matrix coefficients and range factors, rounded once at the coded depth. Run with nearest sampling, PIXL matches libheif’s RGB decode within one code at 8 bits and exactly at 10, full and limited range.
  • The transformations (irot, imir, clap) run in the file’s order with libheif’s rounding, alpha included.
A tree against a sunset sky, cropped from an iPhone 17 HEIC as Apple's Core Image renders it.Core Image · the reference
libheif's default decode against Core Image's, amplified: bright outlines along every leaf and branch.libheif default · up to 25 codes
PIXL's decode against Core Image's, amplified: an even dark field, no edges.PIXL · at most 1 code
An iPhone 17 HEIC (4:2:0) where branches cross a sunset. First, Apple's own render; then how far libheif's default decode and PIXL's land from it, per pixel, the largest channel's difference ×10. libheif's nearest-neighbour chroma misses along every colour edge, by up to 25 of 255; PIXL's centre-sited bilinear chroma is never more than one code away. Over the whole 12 MP frame: libheif 2.0% of pixels more than one code off, PIXL none. libheif 1.23.5's heif-dec; Core Image on macOS 27.0.

The cost is small. On a 48 MP file at 6 threads, the conversion takes 80 ms beside libheif’s 310 ms HEVC decode, against 474 ms for libheif’s own RGB decode. It is bit-identical at any thread count.

Against Apple’s own renders of eleven iPhone HEICs (iPhone 17 and iPhone 13 Pro, iOS 17.5 to 27.0, 12 and 48 MP), the base image is within 1.27 8-bit codes at most. The HDR reading of the same files is covered on theHDR page.

RAW and DNG

A sensor mosaic is not an image, so developing one is always explicit: raw is required whenever a RAW source targets anything but DNG, and refused otherwise. Camera formats are read by LibRaw 0.22.2, built from its unmodified source with PIXL’s own C++ shim. A DNG is read by PIXL’s own reader, written from Adobe’s DNG 1.7.1 specification. Both hand over one view of the sensor, and everything after the decode — scaling, white balance, highlight handling, demosaic, colour — is PIXL’s, on the caller’s threads.

  • 0.31 sCR2 35.6 MB → DNG 28.3 MB, whole sensor, mosaic bit-identical through two readers
  • 1490 MBpeak above the start for a 102 MP GFX100S develop, in 4.9 s
  • every sampleequal to LibRaw’s on the five DNGs both read

Developing in bands

The develop never holds a second full frame. It works in bands of 256 rows, each read with the margin its demosaic reaches, and writes every band straight to its upright place. The result is bit for bit the whole-frame develop. Peak memory above the start, at 8 threads:

FileFrameScene developTime
Canon R5 CR317.3 MP284 MB0.49 s
Sony A7R IV ARW60.2 MP920 MB1.0 s
Fujifilm GFX100S RAF101.7 MP1490 MB4.9 s

The finished 102 MP frame alone is 1221 MB of 32-bit floats. The test holds every develop under 2 GB at 100 MP.

DNG, both ways

The reader handles every compression DNG defines: uncompressed and bit-packed, lossless JPEG (all seven predictors), deflate for integers and for 16-, 24- and 32-bit floats, lossy JPEG and JPEG XL tiles. Floats stay floats. Levels, black-level deltas, up to three calibration illuminants, forward matrices and camera calibrations follow the specification’s chapters 5 and 6. On the five DNGs LibRaw can also read, every sample is equal and the matrices agree to LibRaw’s single precision. The reader is as fast or faster: an iPhone ProRAW decodes in 118 ms against 405, because its 64 tiles decode in parallel.

The opcode lists run where the specification places them. OpcodeList1 and 2 run on the mosaic when the request saysApply, and the plan for both lists is fixed before a sample moves. On a Pixel 6 Pro’s DNG, its four gain maps lift the sky’s top corners from 0.36 of the top centre to 0.71–0.73; a CHDK Canon’s 9 591 dead photosites are patched and no other sample moves. OpcodeList3 — the lens corrections — is never run by itself. The engine reports it with its parameters and offers it, translated exactly, as a lens correction the caller can choose to apply.

The writer stores 256 × 256 tiles, uncompressed or lossless JPEG with the caller’s predictor. It writes the camera’s matrices as exact rationals, carries the MakerNote as DNGPrivateData, and can embed the original file. That original comes back byte for byte; its digest is the MD5 the specification defines. A Fujifilm X-E1 RAF converted to DNG develops as the RAF does, within 1e-5. Everything the engine will not ask LibRaw to do is refused by name: Foveon and SuperCCD layouts, multi-frame files (a dual-pixel CR3 is never reduced silently to its first frame), and files cut short.

Metadata, verbatim

EXIF, ICC, XMP and IPTC are normalised to their bare payloads and copied verbatim — never parsed, rewritten or “fixed”. Container wrappers are added at each boundary: JPEG’s Exif header, the TIFF offset JPEG XL prepends, the Photoshop resource block around IPTC.

SinkEXIFICCXMPIPTCCICP
JPEG XLExif boxprofilexml box—colour encoding
JPEGAPP1APP2, splitAPP1APP13via ICC
PNGeXIfiCCPiTXt—cICP + cLLI
AVIFExif itemcolrXMP itemiptc itemnclx
TIFFIFD0 + Exif/GPS sub-IFDstag 34675tag 700tag 33723via ICC
WebPEXIFICCPXMP —via ICC

A sink with no place for a block reports it false in metadata_written rather than pretending. TIFF is the one container where EXIF is the file’s own directory, so it is re-serialised; the bytes differ but every value compares equal, type and bytes, and libtiff and exiv2 read the result. A RAW’s MakerNote stays with the RAW, because its offsets point into that file; a developed frame’s EXIF carries no orientation tag, because the pixels are already upright. Profiles the engine builds itself are reproducible: the creation time and platform Little CMS stamps in are fixed, so the same conversion writes the same bytes on any run and any OS.

Hostile input

Every parser runs inside the host’s process, on bytes the host did not write. Two measures keep a hostile file from taking that process down.

Limits from the header

A PNG claiming 100 000 × 100 000 pixels asks for tens of gigabytes, and the operating system kills the host. A request’s limits — a pixel count and a longest side — are checked from each format’s header before a pixel buffer is allocated: libjpeg-turbo’s header, PNG’s IHDR, libheif’s handle, WebP’s bitstream features, TIFF’s dimensions, JPEG XL’s basic info, a RAW’s dimensions before unpacking, a DNG’s raw directory and first tile. Every other picture a request reads — a gain map, a mask, an overlay, a flat-field, a model’s grid — is checked where its size is first known and refused by its own field.

Measured: seven forged headers (PNG, JPEG, WebP, AVIF, JPEG XL, TIFF, DNG), each claiming gigabytes, are refused through convert and analyze with the process’s peak resident memory rising under 50 MB in every case. With no limit, a generous one, or one set exactly at the picture’s size, the engine writes the same bytes.

Every parser fuzzed

Thirteen cargo-fuzz targets feed arbitrary bytes to every parser under AddressSanitizer, with libFuzzer’s coverage compiled into the C and C++ as well as the Rust: LibRaw and its shim, libjpeg-turbo, libheif with libde265 and aom, libjxl, libwebp, the PNG and TIFF readers, the gain-map and box readers, PIXL’s EXIF directory reader, Little CMS on any ICC profile, the .cube parser, the prompt-embedding format and the JPEG reconstruction’s coefficient reader. The two RAW targets run for five minutes each on every pull request; every target runs for about 30 minutes before each release.

The same bytes at any thread count

The engine owns no thread pool. Every request names its threads (1–1024), and the work is cut into contiguous runs of whole pixels or rows on scoped threads that end before the call returns. The output is bit-identical whatever the count. Per-element arithmetic never changes, integer counts are summed in any order, and float reductions use a fixed chunking that does not depend on the count. Resampling is cut along the lines its arithmetic already follows, so every output sample is one sum in tap order.

The codecs keep their own threads. The subsampled-HEIF conversion, the DNG tile decode and the JPEG XL decoder are each bit-identical at any count. The AV1 encoder is byte-identical from 2 threads up; at 1 thread it is a different encode of the same quality, and the documentation says so.

Sources and method

Unless stated, figures were measured on an Intel i7-14700F under Linux, with 8 threads. The AVIF encodes were pinned to eight performance cores and run one at a time; the scorers ran one at a time in a container capped at 10 GB. Every claim is held by a test in the engine’s repository; the by-hand measurements are named as such.

  • AVIF tunings, speed and tiling: ci/avif-measure.py → docs/avif/avif-quality.csv, avif-speed.csv, avif-tiling.csv, heic-x265.csv (800 encodes); the 102 MP frame from docs/avif/README.md.
  • AVIF behaviour: tests/avif.rs — the_thread_count_changes_no_byte_from_two_up, one_thread_is_another_encode_of_the_same_quality, a_grid_of_many_tiles_is_read_whole_or_refused_never_wrong; lossless_identity_round_trips_byte_exact, the_nclx_box_states_the_matrix_the_caller_named.
  • Interoperability: ci/avif-interop/ (libavif 1.3.0, libheif 1.23.5, exiftool 13.50), run by hand; Apple devices checked by hand on macOS 27.
  • HEIC chroma: the_arithmetic_is_libheifs_but_for_the_sampling, the_rows_are_the_per_pixel_conversion_at_any_thread_count; Apple’s renders via tests/heic.rs (by hand, --ignored).
  • Sneyers, Alakuijala, Versari, Szabadka et al., “The JPEG XL Image Coding System: History, Features, Coding Tools, Design Rationale, and Future”, arXiv 2506.05987 (2025), CC BY-SA 4.0: §10.1 and Fig. 53 (libjxl 0.11.1, the 8-bit colour photographs of imagecompression.info).
  • JPEG XL repack and reconstruction, and the format matrix: tests/conversions.rs; float JXL: tests/master_cache.rs.
  • RAW and DNG: tests/libraw.rs, tests/dng.rs, tests/raw_memory.rs (peak resident memory via VmHWM), tests/raw_sensors.rs; banding: a_frame_run_in_small_bands_is_the_frame_run_whole.
  • Limits: tests/limits.rs. Fuzzing: fuzz/ (13 targets) and CI’s fuzz job.
  • Threads: a_grade_is_byte_identical_whatever_the_thread_count, resize_f32_in_bands_is_the_single_call_bit_for_bit.