Formats

Unix compress

Lempel-Ziv-Welch with codes from 9 to 16 bits. The .Z files compress wrote and gzip still reads
IDlzw07 / 71 container · reads and writes
Format / Dictionary codes

Unix compress (LZW)

A dictionary that grows as it reads. The .Z of Unix compress.

Ratio
27% of the input, on text
Checksum
none
Standard
ncompress 5.0 (compress 4.2)
Used by
compress, GIF, old tarballs
group 1 / 7

Containers

compress
compress (.Z).Z .taz · magic 1f9d · no checksum
37 bytes · 1f9d9068cab061f306c49d37…reads back · bits 16, blockMode true

Options

bits10 to 16
default 16 · compress
Widest code, 10 to 16, as compress -b takes it; 9 reads differently in gzip and ncompress

Access

Importimport { lzw } from "@agntn/compressions/lzw"
CLIcompressions compress lzw notes.txt > file.Z
Tryplayground with the sample above

Older than gzip, older than most of the internet. LZW builds a dictionary as it reads: every new code is an old code plus one byte. Codes start at 9 bits and widen as the dictionary fills, up to bits. In block mode a clear code throws the dictionary away and starts over. compress sends one when its ratio starts slipping. The writer here never does, it keeps the full table, and the reader takes clear codes wherever they come.

ts
compress("lzw", "hello world!");  // 17 bytes, 1f9d90…

The third byte is 90: block mode on, 16 bits. That's what compress writes and what gzip -d and uncompress read.

bits

From 10 to 16, 16 by default. Fewer bits means a smaller dictionary, which is what old machines with little memory needed. The reader takes 9 too, and widens to 10 the way gzip and ncompress do, so a 9-bit file reads the same here as there.

The quirk

When codes widen, compress skips to the end of the current group of eight codes. Those padding bits are in every .Z file ever written, and a reader that doesn't skip them reads garbage after the first width change. This one skips them, and the writer writes them, so uncompress reads what it makes.

No checksum, no stored size. A damaged .Z file decodes into the wrong bytes without a word.