7-Zip Memory Consumption: Dictionary and Threads
Memory consumption in 7-Zip is primarily dictated by the compression algorithm, the selected dictionary size, the match finder method, and the number of concurrent threads. While decompression memory scales almost exclusively with the dictionary size, compression memory scales dramatically as the number of CPU threads increases under multi-threaded algorithms like LZMA2. Understanding the underlying formulas allows users to optimize compression speed without exhausting system RAM.
Decompression Memory Formula
Decompression requires significantly less memory than compression and is largely independent of thread count for a single archive stream:
\[\text{Decompression Memory} \approx D + 2 \text{ MB}\]
- \(D\): The dictionary size used during compression (e.g., 16 MB, 64 MB, 128 MB).
- Overhead: A small buffer allocation, typically between 2 MB and 4 MB.
Because the decompressor only needs to maintain the sliding dictionary window to rebuild the data, memory consumption directly mirrors the dictionary size.
Compression Memory Formula
Compression memory usage depends heavily on the match finder algorithm and the number of threads (\(T\)). In modern 7-Zip archives, the default algorithm is LZMA2, which divides data into blocks and compresses them concurrently across multiple threads.
The General Formula (LZMA2)
\[\text{Total Compression Memory} \approx T \times (M \times D) + \text{System Buffers}\]
Where:
- \(T\): Number of active compression threads.
- \(D\): Dictionary size.
- \(M\): Match finder multiplier (determined by the compression level).
- System Buffers: Additional chunk and I/O buffers (typically \(D \times 2\) to \(D \times 4\) per stream, plus base program overhead of roughly 10–30 MB).
Match Finder Multipliers (\(M\))
The multiplier depends on the match finder (MF) selected by the compression level:
- BT4 (Normal, Maximum, Ultra levels): \(M \approx 11.5\)
- BT3 (Fast level): \(M \approx 9.5\)
- BT2 / HC4 (Fastest level): \(M \approx 7.5\)
For default, high-compression workflows using BT4, each thread requires roughly \(11.5 \times D\) for the sliding window and hash chains.
Practical Simplified Formula for BT4 (Default/Ultra)
For standard high-compression presets with \(T\) threads:
\[\text{Memory} \approx T \times (11.5 \times D) + (T \times 2 \times D)\]
In practice, 7-Zip simplifies its GUI memory estimation to approximately:
\[\text{Memory} \approx T \times (12 \text{ to } 13) \times D\]
(Note: If \(T\) exceeds the number of available chunks or data size is smaller than the block size, memory usage caps at the number of active blocks.)
Legacy LZMA (LZMA1) Memory Formula
Unlike LZMA2, the original LZMA algorithm does not support multi-block parallelism. It supports a maximum of 2 threads (one for compression and one for match finding):
\[\text{LZMA Compression Memory} \approx 11.5 \times D + 6 \text{ MB}\]
Threads set beyond 2 in LZMA mode do not increase memory consumption or compression speed.
Example Calculation
Compressing an archive using LZMA2, Ultra (BT4 match finder), a 64 MB dictionary, and 8 threads:
- Per-thread memory: \(64 \text{ MB} \times 11.5 = 736 \text{ MB}\)
- Total thread allocation: \(8 \times 736 \text{ MB} = 5,888 \text{ MB}\)
- Buffer and overhead allocation: \(\approx 8 \times 128 \text{ MB} \approx 1,024 \text{ MB}\)
- Estimated total memory: \(\approx 6.9 \text{ GB}\) to \(7.2 \text{ GB}\) of RAM.