Compressing and decompressing
Two calls
import { compress, decompress } from "@agntn/compressions";
const packed = compress("lzma", "hello world!", { container: "xz", level: 9 });
const { bytes, details } = decompress("lzma", packed, { container: "xz" });
details; // { check: "crc64", blocks: 1 }
compress(name, input, options) takes a string, read as UTF-8, or a Uint8Array. decompress(name, data, options) takes a Uint8Array and hands back { bytes, details }. That's the whole API you need on day one.
The name forgives you. BZIP2, bz2 and .bz2 all find bzip2, zst finds zstd, br finds brotli, compress and z find lzw. What it won't do is guess a container for you: gzip is a container, so decompress("gzip", …) throws and tells you to write deflate with container: "gzip".
Options
Each format declares what it takes, and info() lists it. Anything else is an error, never silently dropped.
| Format | Options |
|---|---|
deflate | container (raw, zlib, gzip), level 0 to 9, name and mtime for gzip |
bzip2 | level 1 to 9, the block size in hundreds of kilobytes |
lzma | container (alone, xz), level 0 to 9, check for xz |
lz4 | container (frame, legacy), checksum for the frame |
lzw | bits 10 to 16 |
zstd, brotli | none, this package reads them |
compress("deflate", "x", { name: "a.txt" });
// InvalidOptionError: Invalid option name=a.txt: only gzip takes it, not raw
compress("deflate", "x", { levl: 9 });
// InvalidOptionError: Invalid option levl=9: deflate takes container, level, name, mtime
decompress takes only the options marked to read with, which is container and nothing else. Plus limit and partial, which every format takes, see limits.
Levels
level means what the format's own tool means by it. deflate follows zlib's ten levels, from 0 (stored, nothing compressed) to 9 (longest search). The default is 6, same as gzip. lzma follows the xz presets: the level picks the dictionary size, 256 KiB at 0 up to 64 MiB at 9, and how hard it searches. bzip2's level is the block size, and 9 is the default, same as bzip2.
On a kilobyte of repeated text deflate gives:
| Level | Bytes |
|---|---|
| 0 | 1025, stored blocks |
| 1 | 32 |
| 6 | 26 |
| 9 | 26 |
Higher levels mostly pay off on bigger, messier input. Small text tops out fast.
What decompress gives back
Bytes, always. And details, the facts the stream carried about itself:
decompress("deflate", gz, { container: "gzip" }).details;
// { name: "notes.txt", mtime: 1700000000, os: 255, members: 1 }
Every format reads every stream in a row, the way its own tool does. Two gzip members come out as one, and members says 2. Bytes after the last stream that aren't another stream show up as trailing, with how many. Zero padding after gzip is fine and isn't counted, since tapes did that.
Format objects and subpaths
Every format is an object with name, info(), compress and decompress. The registry functions just look the object up by name. Import the object straight from its own subpath and plain Node loads that format and nothing else:
import { deflate } from "@agntn/compressions/deflate";
deflate.compress("hi", { container: "zlib" });
| Subpath | Exports |
|---|---|
@agntn/compressions/deflate | deflate, DEFLATE_CONTAINERS |
@agntn/compressions/bzip2 | bzip2 |
@agntn/compressions/lzma | lzma, LZMA_CONTAINERS, XZ_CHECKS |
@agntn/compressions/zstd | zstd |
@agntn/compressions/brotli | brotli |
@agntn/compressions/lz4 | lz4, LZ4_CONTAINERS |
@agntn/compressions/lzw | lzw |
They're the same objects the root exports, not copies. The brotli one is the heavy one, since it carries the 122,784 byte dictionary of RFC 7932, deflated. Nothing else loads it.
Errors
Every error is a CompressionError. A broken stream throws DecompressError, with the format, the offset in the input when one byte is at fault, and partial, everything that came out before the damage. Its subclasses say more: ChecksumError for a stream that decoded but doesn't check, LimitError for output past the limit, UnsupportedError for a valid stream using something this package doesn't do, like a zstd dictionary or the xz ARM filter.
decompress("deflate", flipped, { container: "gzip" });
// ChecksumError: deflate: gzip CRC-32 does not match at byte 24