deflate
deflate
LZ77 and Huffman, raw or wrapped in zlib or gzip. Still everywhere.
- Ratio
- 9% of the input, on text
- Checksum
- Adler-32, CRC-32
- Standard
- RFC 1951
- Used by
- gzip, PNG, ZIP, HTTP, PDF
Containers
- raw
- raw deflateno file extension · no magic number · no checksum
- 17 bytes · cb48cdc9c95728cf2fca4951…reads back
- zlib
- zlib.zz .zlib · no magic number · Adler-32
- 23 bytes · 789ccb48cdc9c95728cf2fca…reads back · level default
- gzip
- gzip.gz .tgz · magic 1f8b08 · CRC-32
- 35 bytes · 1f8b08000000000000ffcb48…reads back · os 255, members 1
Options
Access
- Import
import { deflate } from "@agntn/compressions/deflate" - CLI
compressions compress deflate notes.txt --container raw > file - Tryplayground with the sample above
- Kinzstd, brotli
The one everybody has. gzip files, zlib streams, PNG images, ZIP entries, HTTP bodies, PDF streams. All of it is deflate inside a different envelope, and here the envelope is an option:
compress("deflate", "hello world!"); // 14 bytes, raw
compress("deflate", "hello world!", { container: "zlib" }); // 20 bytes, 789c…
compress("deflate", "hello world!", { container: "gzip" }); // 32 bytes, 1f8b08…
The writer is LZ77 with lazy matching, the way zlib does it, and picks fixed, dynamic or stored blocks by whichever comes out smaller. On text it lands at zlib's size or under.
Levels
level goes from 0 to 9, 6 by default. 0 stores: no compression, just the bytes in blocks of up to 64 KiB, so "hello world!" comes out as 17 bytes with the text sitting in plain sight. 1 is fastest, 9 searches longest. The gaps are small on short input and grow with it.
zlib
Two header bytes, raw deflate, then the Adler-32 of the output. The header names the level the writer used, which comes back in details.level as fastest, fast, default or best. A zlib stream that asks for a preset dictionary is an UnsupportedError, since nobody hands one over.
gzip
Ten header bytes starting 1f 8b 08, raw deflate, then a CRC-32 and the size. name and mtime go in the header when you set them, and come back in details:
const gz = compress("deflate", "hello world!", { container: "gzip", name: "notes.txt", mtime: 1700000000 });
decompress("deflate", gz, { container: "gzip" }).details;
// { name: "notes.txt", mtime: 1700000000, os: 255, members: 1 }
os is 255, unknown, because this package doesn't know what machine you're on and isn't going to guess. Several members in a row decompress into one output, the way gunzip reads cat a.gz b.gz, and details.members counts them. Zeros after the last member are tape padding and get skipped. Anything else after it is reported as trailing.