Unix compress
Unix compress (LZW)
A dictionary that grows as it reads. The .Z of Unix compress.
- Ratio
- 27% of the input, on text
- Checksum
- none
- Standard
- ncompress 5.0 (compress 4.2)
- Used by
- compress, GIF, old tarballs
Containers
- compress
- compress (.Z).Z .taz · magic 1f9d · no checksum
- 37 bytes · 1f9d9068cab061f306c49d37…reads back · bits 16, blockMode true
Options
Access
- Import
import { lzw } from "@agntn/compressions/lzw" - CLI
compressions compress lzw notes.txt > file.Z - Tryplayground with the sample above
Older than gzip, older than most of the internet. LZW builds a dictionary as it reads: every new code is an old code plus one byte. Codes start at 9 bits and widen as the dictionary fills, up to bits. In block mode a clear code throws the dictionary away and starts over. compress sends one when its ratio starts slipping. The writer here never does, it keeps the full table, and the reader takes clear codes wherever they come.
compress("lzw", "hello world!"); // 17 bytes, 1f9d90…
The third byte is 90: block mode on, 16 bits. That's what compress writes and what gzip -d and uncompress read.
bits
From 10 to 16, 16 by default. Fewer bits means a smaller dictionary, which is what old machines with little memory needed. The reader takes 9 too, and widens to 10 the way gzip and ncompress do, so a 9-bit file reads the same here as there.
The quirk
When codes widen, compress skips to the end of the current group of eight codes. Those padding bits are in every .Z file ever written, and a reader that doesn't skip them reads garbage after the first width change. This one skips them, and the writer writes them, so uncompress reads what it makes.
No checksum, no stored size. A damaged .Z file decodes into the wrong bytes without a word.