Formats

LZMA

LZ77 with a range coder that learns every bit. The xz container with its checks and filters and the older .lzma
IDlzma03 / 72 containers · reads and writes
Format / Range coding

LZMA

A range coder over bit models that learn. Small and patient.

Ratio
8% of the input, on text
Checksum
CRC-64 by default; CRC-32, SHA-256 or none
Standard
LZMA SDK lzma-specification.txt (Igor Pavlov)
Used by
xz, 7-Zip, kernel images
group 1 / 7

Containers

alone
.lzma.lzma · magic 5d · no checksum
34 bytes · 5d0000800033000000000000…reads back · properties lc=3 lp=0 pb=2, dictionary 8388608
xz
xz.xz .txz · magic fd377a585a00 · CRC-64 by default; CRC-32, SHA-256 or none
80 bytes · fd377a585a000004e6d6b446…reads back · check crc64, blocks 1

Options

containeralone, xz
default alone · compress and decompress
alone for a .lzma file, xz for an .xz file
level0 to 9
default 6 · compress
Dictionary size and search depth, as xz's presets: 0 is 256 KiB, 9 is 64 MiB
checknone, crc32, crc64, sha256
default crc64 · xz only
What xz stores after the block to check it

Access

Importimport { lzma } from "@agntn/compressions/lzma"
CLIcompressions compress lzma notes.txt --container alone > file.lzma
Tryplayground with the sample above

LZMA doesn't write Huffman codes. It keeps a probability for every decision, a literal or a match, which length, which distance, and a range coder spends a fraction of a bit on each one it guessed right. Those probabilities learn as the stream goes, and that's where the size comes from.

ts
compress("lzma", "hello world!");                       // 30 bytes, .lzma
compress("lzma", "hello world!", { container: "xz" });  // 68 bytes, fd377a585a00…

.lzma

The alone container, the default: thirteen header bytes with the lc, lp, pb properties, the dictionary size and the unpacked size, then one LZMA stream. No checksum. It's what LZMA Utils wrote before xz existed, and what plenty of embedded firmware still ships.

ts
decompress("lzma", compress("lzma", "hello world!")).details;
// { properties: "lc=3 lp=0 pb=2", dictionary: 8388608 }

xz

LZMA2 inside blocks, an index of every block, CRC-32 on every header and a check of your choosing on the data. check takes none, crc32, crc64 (the default) or sha256. On "hello world!" that's 60, 64, 68 and 92 bytes.

LZMA2 cuts the stream into chunks, so a chunk that won't compress gets stored as it is instead of growing. Concatenated streams and the padding between them decompress into one output, as xz -d reads them.

Filters on the read side: delta and x86 BCJ, mirrored from xz 5.8. ARM, PowerPC, SPARC, IA-64 and RISC-V filters are an UnsupportedError that names the filter. The writer uses no filters.

Levels

level from 0 to 9, 6 by default, sets the dictionary and how hard the match finder looks, the same way the xz presets do. The writer is honest and not fast: about 10% bigger than xz on binaries, close on text.