Benchmark
How HazeHash compares with BlurHash and ThumbHash, and why it is built the way it is. Every number comes from the scripts in the repository's bench/ folder and can be reproduced locally:
node scripts/fetch-test-images.mjs # one-off: the 513-image set into test-images/
pnpm bench # report.html, results.csv, summary.csv in bench/out/
pnpm ablate # the ablation table in bench/out/ablation.md
pnpm --filter hazehash-bench perf # timingsMethod
- The set — 513 images (photos, portraits, architecture, products, screenshots, graphics, 91 with transparency, food, documents, night scenes, panoramas), at most 1280 px on the long side.
- The reference — the source image area-averaged in linear light to the decoder's output size. Images with alpha are scored over white and over black, and the errors are averaged.
- The error — ΔE is 100 × the Euclidean distance in OKLab per pixel. The score of an image is the mean over its pixels, and the tables report the mean over images. SSIM uses the OKLab L channel with a 7×7 window.
- Baselines at the same byte count — BlurHash uses the largest grid whose string carries at most the budget in information (length × log₂ 83 / 8) and sees the image flattened on white. ThumbHash uses its native output, about 21 bytes on this set.
Results
Profile default. Lower ΔE is better, higher SSIM is better.
| Codec | Budget | Median bytes | Mean ΔE | SSIM(L) |
|---|---|---|---|---|
| HazeHash | 16 | 16 | 8.41 | 0.495 |
| HazeHash | 20 | 20 | 7.89 | 0.549 |
| HazeHash | 24 | 24 | 7.54 | 0.588 |
| HazeHash | 28 | 28 | 7.29 | 0.614 |
| HazeHash | 36 | 36 | 7.00 | 0.644 |
| HazeHash | 48 | 43 | 6.91 | 0.651 |
| BlurHash | 16 | 16 | 15.78 | 0.243 |
| BlurHash | 28 | 28 | 14.32 | 0.335 |
| BlurHash | 48 | 48 | 13.62 | 0.405 |
| ThumbHash | native | 21 | 9.13 | 0.449 |
At 28 bytes HazeHash has a mean ΔE 49.1% lower than BlurHash and 20.2% lower than ThumbHash. The 95th percentile over images is 14.55 against 15.74 for ThumbHash, so the gain is not bought with outliers. Even at 20 bytes, the size of ThumbHash's output, HazeHash is 13.6% better.
The default budget
| Step | Change in mean ΔE | Bytes |
|---|---|---|
| 28 → 24 | +3.4% | −14% |
| 24 → 20 | +4.6% | −17% |
| 28 → 36 | −4.0% | +29% |
The default of 28 bytes (38 characters) keeps a wide margin over ThumbHash and sits between the steep part of the curve below 24 bytes and the flat part above 36. Use budget: 24, which is 32 characters and still 17% better than ThumbHash, when storage dominates.
Ablation
Each row changes one design choice of the production encoder, using a laboratory codec whose baseline reproduces the production one bit for bit. 120 opaque images, mean ΔE and the change against the baseline.
| Variant | 20 B ΔE | vs base | 28 B ΔE | vs base | 36 B ΔE | vs base |
|---|---|---|---|---|---|---|
| baseline (production) | 6.894 | 6.216 | 5.899 | |||
narrow mask, i/nx + j/ny < 1 | 6.889 | −0.1% | 6.445 | +3.7% | 6.422 | +8.9% |
| rectangular mask | 6.906 | +0.2% | 6.243 | +0.4% | 5.894 | −0.1% |
| frequency-dependent Qmax bands | 6.838 | −0.8% | 6.219 | +0.0% | 5.939 | +0.7% |
| fixed 4 bits instead of Rice | 6.878 | −0.2% | 6.268 | +0.8% | 5.968 | +1.2% |
| no rate-distortion optimisation | 6.904 | +0.1% | 6.244 | +0.5% | 5.931 | +0.5% |
| one shared grid for L, a and b | 7.368 | +6.9% | 6.666 | +7.2% | 6.213 | +5.3% |
| gamma sRGB (Y'CbCr) instead of OKLab | 7.031 | +2.0% | 6.358 | +2.3% | 6.030 | +2.2% |
What this says:
- Separate luma and chroma grids matter most: without them the error is 5–7% higher.
- OKLab is worth about 2% over a gamma-encoded Y'CbCr basis.
- The mask matters at larger budgets. The original, narrower triangle costs 3.7% at 28 bytes and 8.9% at 36 bytes, because the rate-distortion search can only choose among the coefficients the mask allows. The production mask keeps
i/nx + j/ny < 1.5, and a full rectangle brings nothing more while storing more coefficients. - Rice coding saves about 1% over a fixed 4-bit code and rate-distortion optimisation about 0.5%; both are cheap enough to keep.
- Frequency-dependent quantisation brings nothing (±1%), so the format uses one flat value per channel type: 7 levels for luma, 3 for chroma and alpha.
Speed
Approximate timings for a 64×48 input on a desktop machine, the order of magnitude to expect. The repository's perf script measures them on your own hardware.
- Decoding a preview at 32 px takes about 0.3 ms.
- Encoding with
fasttakes about 1 ms, withdefaultabout 5 ms and withhighabout 14 ms. Images with alpha are slower, around 100 ms. A larger input only adds the time of the area-average downscale.