LZMA vs LZMA2: 7-Zip Compression Differences
The LZMA2 compression algorithm is an updated container format built on top of the original Lempel-Ziv-Markov chain algorithm (LZMA), designed specifically to resolve critical performance bottlenecks in 7-Zip. While both algorithms use the same core dictionary-based compression engine, LZMA2 introduces major improvements in multi-threaded CPU utilization, the handling of uncompressible data, and dynamic stream management. These updates make LZMA2 significantly faster and more resource-efficient on modern multi-core systems without sacrificing compression ratio.
True Multi-Threading Support
The most significant limitation of the original LZMA is its restricted multi-threading capability. When compressing data using original LZMA, 7-Zip can effectively utilize at most two threads—one for the match finder and one for the main compression process. This creates a severe bottleneck on modern multi-core processors.
LZMA2 solves this by splitting the input data into multiple independent chunks. Each chunk is processed in parallel across available CPU cores. As a result, LZMA2 can scale across 4, 8, 16, or more threads, dramatically reducing overall compression time on multi-core systems.
Efficient Handling of Incompressible Data
The original LZMA attempts to compress all incoming data, regardless of its content. When processing data that is already compressed or naturally random (such as JPEG images, MP3 audio, or encrypted files), LZMA wastes CPU cycles and can actually cause "data expansion," resulting in an output file larger than the original input.
LZMA2 dynamically inspects the data during compression. If a block of data is determined to be incompressible, LZMA2 stops attempting to compress it and stores it as uncompressed raw data. This eliminates data inflation and prevents the processor from wasting cycles on unshrinkable files.
Dynamic State Resets
In original LZMA, the entire compressed stream relies on a continuous dictionary state established from the start of the file. If the nature of the incoming data changes (for example, transitioning from text to executable binary), the algorithm cannot easily adapt its internal probabilities.
LZMA2 permits dynamic dictionary and property resets between data chunks. The compressor can change encoding parameters or reset the LZMA state mid-stream to optimize for different types of data bundled within the same archive.
Compression Ratio vs. Speed
At its core, LZMA2 uses the same entropy encoding and dictionary mechanisms as LZMA. Consequently, the compression ratio between the two algorithms is virtually identical on single-threaded runs. However, because LZMA2 divides data into discrete chunks for multi-threading, each thread maintains its own context, which can cause a negligible decrease in compression ratio (typically under 1%). In practical use, the dramatic increase in compression speed far outweighs this minor difference.