Guide

Checksums

CRC-32 and Adler-32 and CRC-64 and XXH32 and XXH64. Every one checked on the way out and a wrong one is an error

Checked, not trusted

Every container that stores a checksum gets it checked on the way out. A wrong one is a ChecksumError, never a shrug and a warning:

ts
decompress("deflate", flipped, { container: "gzip" });
// ChecksumError: deflate: gzip CRC-32 does not match at byte 24

The bytes did come out, though. They're on the error as partial, so a puzzle with a deliberately broken trailer can still be read. Limits and partial reads has the partial option that returns them instead of throwing.

Which one where

ContainerChecksumOver
zlibAdler-32the output
gzipCRC-32, and the size modulo 2³²the output of each member
bzip2CRC-32/BZIP2each block's output, then all blocks combined
xzCRC-32, CRC-64 or SHA-256, as the stream sayseach block's output
xzCRC-32the stream header, each block header, the index, the footer
zstdthe low 32 bits of XXH64the frame's output, when the flag is set
lz4 frameXXH32the header, each block when flagged, the content when flagged

Raw deflate, .lzma, brotli, lz4 legacy and Unix compress carry none. A damaged stream in one of those can decode into garbage without a word. That's the format, not this package. identify scores them lower for exactly that reason.

Where the checksums come from

From @agntn/hashes, not from this package. The same code that computes crc32, crc64, adler32 and xxhash there checks the streams here, so a fix lands in one place. A few of them came in for this package: CRC-32/BZIP2 is crc32 with variant: "bzip2", CRC-64/XZ is crc64, XXH32 is xxhash with bits: 32.

Sizes count too

gzip stores the size modulo 2³², zstd and lz4 can store the content size, xz stores sizes in the block headers and the index. Each one is held against what actually came out, and a mismatch is a ChecksumError as well. A stream that decodes cleanly but to the wrong length is still a broken stream.