Identify and peel
Try it, don't guess it
identify doesn't pattern match and hope. It tries every container whose start fits, decompresses for real, and ranks what worked:
import { identify } from "@agntn/compressions";
identify(gz);
// [{ format: "deflate", container: "gzip", confidence: 100, confirmed: true,
// reasons: ["starts with its magic number", "checksum matches",
// "decodes to the last byte", "decompresses to readable text"], bytes, details }]
Each candidate carries the decompressed bytes, the details and the reasons behind its score. The confidence ranks, it isn't a probability.
What counts
| Evidence | Points |
|---|---|
| starts with its magic number | 50 |
| header fields are valid (zlib, .lzma) | 25 |
| checksum matches | 30 |
| decodes to the last byte | 10 |
| bytes follow the stream | minus 15 |
| decompresses to readable text | 10 |
| comes out longer than it went in | 5 |
A container with a magic number is tried only when the bytes start with it. zlib and .lzma have no real magic, so their headers get checked instead: zlib's two bytes carry their own check, .lzma's properties and sizes have to make sense. Raw deflate and brotli have neither. They get tried on everything, and they only count when they decode every single byte and give something back, since random bytes decode as deflate more often than you'd think.
confirmed marks a candidate backed by more than luck: a magic number, or a checked header with a checksum or readable text behind it.
When nothing fits
identify returns an empty list. That's an answer, not a failure. Plain text isn't compressed. And when the bytes start like an archive this package doesn't open, archiveOf names it:
archiveOf(zipBytes); // "a ZIP archive"
ZIP, 7-Zip, RAR, tar, cabinet files, lzip and PNG are on that list. A ZIP entry's data is usually raw deflate, so once you've cut it out, decompress("deflate", entry) reads it.
Peel
Puzzles love nesting. peel takes off one layer after another while each is confirmed:
import { peel } from "@agntn/compressions";
peel(compress("bzip2", compress("lzma", compress("deflate", text, { container: "gzip" }), { container: "xz" })));
// bzip2 · bzip2, 90, confirmed
// lzma · xz, 90, confirmed
// deflate · gzip, 100, confirmed
The list runs outermost first, and the last layer's bytes is what's at the bottom. An unconfirmed best guess ends the list, marked confirmed: false, so you see it without the next layer building on it. depth caps the layers, 10 by default.
The tool does the same with peel: true, and hands the innermost bytes back in base64. On the CLI, compressions identify --peel file lists the layers on stderr and writes the bottom to stdout.