Skip to content

Hash Format ​

How a HazeHash is laid out, for people who want to understand a hash or write another decoder. This is the format version 1 as implemented by hazehash 0.1. Conformance is defined by the frozen test vectors in the repository (packages/core/test/vectors/v1.json): a decoder must produce the same RGBA, within ±1 per channel.

Overview ​

An image becomes a short byte string: a fixed header followed by Golomb–Rice coded DCT coefficients of the image in OKLab. The default budget is 28 bytes (38 base64url characters), and the supported range is 16–48 bytes. The header alone is 7 bytes, or 9 with alpha, and a hash can be at most 1024 bytes.

Bits are written most significant first, and multi-bit fields are big-endian. The tail of the last byte is padded with zeros. A bit read past the end of the data is 0, so a truncated string is valid and decodes with less detail.

The header has 56 bits:

FieldBitsOffsetMeaning
ver200 for version 1; 1–3 are reserved and must be rejected
alpha121 if the alpha block follows the header
aspect63r = 2^((code − 32) / 8), code 32 is 1:1
Lx − 139luma grid width, 1–8
Ly − 1312luma grid height, 1–8
Cx − 1215chroma grid width, 1–4
Cy − 1217chroma grid height, 1–4
DC L619L = q / 63
DC a, DC b6+625, 31v = q · 0.64 / 63 − 0.32
scale L/a/b4×337AC amplitude limit code per channel
k L/a/b2×349Rice parameter per channel
reserved155written as 0, ignored when reading

The alpha block has 16 bits and is present only when alpha = 1: DC A (5 bits, A = q / 31), Ax − 1 (2), Ay − 1 (2), scale A (4), k A (2) and a reserved bit (1).

A reader must reject input shorter than the header (InvalidLength), longer than 1024 bytes, and an unknown version (UnsupportedVersion).

Aspect ratio and output size ​

The ratio is stored as code = clamp(round(8 · log2(W / H)) + 32, 0, 63), rounding half away from zero, so one step is about 9% and the range is about 1:16 to 15:1. A decoder asked for a preview with the long side size produces size × max(1, round(size / r)) when r ≥ 1, and max(1, round(size · r)) × size otherwise. The ratio is for drawing the placeholder; page layout should use the real image dimensions.

Coefficients ​

The image is expanded in a DCT-II basis with pixel-centre sampling on the unit square:

text
f(u, v) = Σ c[i][j] · cos(π i u) · cos(π j v)

c[0][0] is the mean of the channel and lives in the header. The other coefficients of a grid nx × ny are stored only when 2 · (i · ny + j · nx) < 3 · nx · ny, that is i / nx + j / ny < 1.5, which cuts only the far corner of the grid. The number of stored AC coefficients of an n × n grid is 0, 3, 8, 14, 23, 32, 45 and 57 for n = 1…8.

Stream order ​

All AC coefficients of all channels form one sequence, sorted by:

  1. ρ² = (i² · ny² + j² · nx²) / (nx² · ny²), compared by cross-multiplication in integers,
  2. then the channel in the order L, a, b, A,
  3. then the smaller j, then the smaller i.

Cutting the tail therefore removes the highest frequencies of every channel at once.

Quantisation ​

Each channel has a scale s = smax · 2^(−(15 − code) / 3), with smax = 0.64 for L and A and 0.32 for a and b. A coefficient is coded as an integer q in [−Qmax, Qmax]:

text
q = round(sign(c) · sqrt(min(|c| / s, 1)) · Qmax)
ĉ = sign(q) · s · (|q| / Qmax)²

Qmax is 7 for luma and 3 for a, b and alpha, at every frequency. A decoder must reconstruct with exactly this formula, with no reconstruction offset, and clamp the decoded q to ±Qmax.

Entropy coding ​

Every q is mapped to n = 2q (for q ≥ 0) or −2q − 1 (for q < 0) and written as n >> k one bits, a zero bit, then the k low bits of n. k (0–3) is stored per channel. A zero coefficient costs k + 1 zero bits, so the encoder may drop trailing zero bytes of the data, which is why a simple image gives a hash shorter than the budget.

Decoding to pixels ​

  1. Rebuild each channel by separable synthesis at the output size.
  2. Clamp L to [0, 1] and convert OKLab to linear RGB with the Ottosson matrices.
  3. Gamut mapping: if any component leaves [−0.0005, 1.0005], scale the chroma by the largest t ∈ [0, 1] found with 8 bisection steps, then clamp to [0, 1].
  4. Convert to sRGB and apply interleaved-gradient-noise dithering: u = fract(52.9829189 · fract(0.06711056 x + 0.00583715 y)) and out = floor(v · 255 + u).
  5. Alpha, when present, is clamp(A, 0, 1) · 255 rounded. Colours are not premultiplied.

String form ​

The string form is base64url (RFC 4648 §5) without padding, so 28 bytes are 38 characters. A length that is 1 modulo 4, or any character outside the alphabet, is invalid.