How LZMA2 Handles Uncompressible Data in 7-Zip
This article explains how the LZMA2 compression algorithm manages incompressible data streams within 7-Zip. While traditional compression algorithms can cause file bloat and waste processing cycles when encountering already compressed or random data, LZMA2 uses a chunk-based container format that identifies non-compressible blocks and stores them raw. Below is an overview of the mechanics, chunk structure, and performance benefits behind LZMA2’s handling of incompressible chunks.
The Problem with Traditional LZMA
The original LZMA algorithm treats an input stream as a continuous sequence. When fed data that cannot be compressed—such as encrypted files, JPEGs, or already compressed archives—the algorithm continues to update its probability models and emit encoded symbols. This results in two major drawbacks:
- Data Expansion: The output size becomes larger than the input size due to encoding overhead.
- Wasted CPU Cycles: The processor expends resources calculating complex range encoding and dictionary matches for zero gain.
The LZMA2 Chunking Mechanism
LZMA2 is an envelope format built on top of the raw LZMA compression engine. Instead of compressing the entire stream continuously, LZMA2 divides the input into variable-sized chunks:
- Compressed Chunks: Can contain up to 2 MB of uncompressed data.
- Uncompressed Chunks: Can contain up to 64 KB (65,536 bytes) of raw data.
Each chunk is preceded by a control byte that dictates how the decoder must interpret the following payload.
Detecting Incompressible Data
During compression, 7-Zip monitors the compression ratio of the current block. If the compressed output exceeds the original input size, the compressor aborts the compression of that specific block.
Instead of writing the expanded encoded data, the compressor discards it and marks the data as an uncompressed chunk.
Storing Raw Data Chunks
When an uncompressible chunk is written, LZMA2 encapsulates it using a minimal header format:
- Control Byte: LZMA2 uses a specific control byte
(values
0x01or0x02) to signal an uncompressed chunk.0x01indicates that the LZMA state dictionary should be reset.0x02indicates that the dictionary state should be preserved.
- Data Size: A two-byte integer follows the control byte, storing the uncompressed size minus one (supporting up to 65,535 bytes).
- Payload: The original, raw bytes are copied directly to the archive without any encoding.
Because the maximum uncompressed chunk size is 64 KB, large
incompressible files are simply split into sequential 64 KB chunks, each
adding only a 3-byte header overhead
(1 control byte + 2 size bytes). This limits data expansion
to less than 0.005%.
Advantages of the LZMA2 Approach
- Zero File Bloat: Overhead on random or pre-compressed data is virtually negligible compared to standard LZMA.
- CPU Efficiency: The encoder quickly moves through non-compressible sections without running deep dictionary searches.
- State Management: When transitions occur between compressible and incompressible data, LZMA2 can preserve or reset its dictionary states, allowing it to resume optimal compression immediately when compressible patterns reappear.
- Multi-threading Stability: Because data is split into independent or semi-independent blocks, multiple threads can process different chunks concurrently without corrupting global compression states.