How LZMA2 Multi-Threading Improves 7-Zip Speed
This article explains how the LZMA2 algorithm enhances data compression performance in 7-Zip through multi-threading. While the original LZMA algorithm was fundamentally restricted to single-threaded or dual-threaded operations, LZMA2 was designed to divide data into independent chunks, allowing modern multi-core processors to compress data streams concurrently. Understanding this architecture clarifies why LZMA2 scales efficiently across high-core systems, how it minimizes throughput bottlenecks, and what trade-offs it introduces regarding memory usage and compression ratios.
The Single-Thread Bottleneck of Original LZMA
The original LZMA (Lempel-Ziv-Markov chain-Algorithm) provides exceptionally high compression ratios, but its internal mechanics make parallelization difficult. LZMA relies on a single continuous sliding window and dynamic probability models that depend directly on the sequence of preceding bytes. Because the encoder must know the state of previous data to encode the next byte, the workload cannot easily be split across multiple CPU cores without breaking the compression context. As a result, standard LZMA in 7-Zip is typically limited to at most two threads (one for finding matches and one for entropy encoding), leaving modern multi-core CPUs largely underutilized.
How LZMA2 Solves the Problem: Block-Based Parallelism
LZMA2 resolves this architectural limitation by packaging the LZMA stream into discrete, independent chunks. Instead of treating an entire file or archive as one uninterrupted data stream, LZMA2 segments the uncompressed input into separate blocks.
- Independent Chunk Processing: 7-Zip assigns individual data chunks to separate worker threads. Each thread compresses its designated block independently of the others.
- Concurrent Execution: On an 8-core, 16-thread CPU, 7-Zip can assign 16 chunks simultaneously, utilizing the full processing capacity of the hardware.
- Stream Reconstruction: Once the threads finish
compressing their individual chunks, 7-Zip assembles the resulting LZMA2
blocks into the final
.7zcontainer sequentially.
Because each thread operates on an isolated chunk, there is no need for cross-thread synchronization during the computationally heavy compression phase. This eliminates lock contention and allows near-linear scaling with core counts up to the thread limit configured by the user.
Handling Incompressible Data Efficiently
Another way LZMA2 improves speed through multi-threading is its handling of incompressible data. If a thread encounters a block containing data that does not compress effectively (such as already-compressed media files or encrypted archives), LZMA2 detects this condition quickly. The thread can immediately store the chunk uncompressed instead of expending intensive CPU cycles attempting to optimize an LZMA model. This dynamic switching prevents single threads from stalling the entire multi-threaded pipeline on difficult data blocks.
The Impact on Compression Ratio and Memory
While LZMA2 multi-threading dramatically reduces compression time, it introduces two practical trade-offs:
- Slight Compression Ratio Loss: Because the sliding dictionary state must periodically reset between independent blocks, the encoder cannot reference patterns across chunk boundaries. This typically results in a negligible decrease in compression ratio compared to a single-stream LZMA archive, but the speed gains usually outweigh the minimal increase in file size.
- Increased Memory Consumption: Each active thread requires its own dictionary buffer in RAM. If 7-Zip is configured to use a 64 MB dictionary with 8 threads, the compression process will require significantly more system memory than a single-threaded job. If system RAM is insufficient, 7-Zip will automatically reduce the number of active threads to prevent swapping.