7-Zip Memory Usage vs Dictionary Size Explained
When compressing files with 7-Zip using the default LZMA or LZMA2 algorithms, memory consumption is directly tied to the selected dictionary size. The dictionary acts as a sliding window of recent data used to find repetitive patterns; expanding this window improves the compression ratio on repetitive data but exponentially increases RAM requirements. During compression, memory usage scales at roughly ten to twelve times the dictionary size per thread, whereas decompression requires only slightly more RAM than the dictionary size itself.
The Compression Memory Formula
For standard 7-Zip compression using LZMA, memory usage does not simply equal the dictionary size. The algorithm requires additional data structures to index and search for matches within that window.
The primary driver of this overhead is the Match Finder algorithm:
- BT4 (Binary Tree 4): The default for "Normal," "Maximum," and "Ultra" compression levels. It requires approximately 10.5 to 11.5 times the dictionary size in RAM for a single thread. This overhead accommodates the dictionary buffer itself, cyclic hash chains, and binary tree nodes used to track 2-byte, 3-byte, and 4-byte hashes.
- HC4 (Hash Chain 4): Used in "Fast" mode. It demands roughly 6.5 to 7.5 times the dictionary size in RAM, trading compression efficiency and tree structures for lower memory usage and faster processing.
Multi-Threading and LZMA2 Scaling
Modern multi-core systems typically use the LZMA2 algorithm, which breaks input data into independent chunks to distribute work across multiple CPU threads.
Because each active thread must maintain its own search buffers and hash structures, memory usage scales with both dictionary size and thread count:
- Thread Allocation: 7-Zip typically allocates memory
equal to
(Match Finder Multiplier × Dictionary Size) × Number of Active Threads. - Chunk Overlap: To maintain compression ratios across split blocks, LZMA2 threads keep overlapping references, which can add modest additional memory buffers per thread.
- Core Saturation: On a system with 8 or 16 threads, selecting a large dictionary can exhaust system RAM rapidly. For example, a 64 MB dictionary running on 8 threads with the BT4 match finder requires approximately 5.5 GB to 6 GB of RAM, whereas the same dictionary size on a single thread requires around 700 MB.
Real-World Compression Scaling Examples
Assuming standard "Ultra" settings (BT4 match finder, LZMA2):
- 16 MB Dictionary: ~180 MB per thread (approx. 1.4 GB on an 8-thread CPU)
- 64 MB Dictionary: ~700 MB per thread (approx. 5.6 GB on an 8-thread CPU)
- 128 MB Dictionary: ~1.4 GB per thread (approx. 11.2 GB on an 8-thread CPU)
- 512 MB Dictionary: ~5.5 GB per thread (approx. 44 GB on an 8-thread CPU)
- 1024 MB Dictionary: ~11 GB per thread (exceeds typical desktop RAM when multi-threaded)
If total required memory exceeds physical RAM, 7-Zip automatically scales down the number of threads or the dictionary size to prevent disk thrashing via the system pagefile.
Decompression Memory Scaling
Unlike compression, decompression memory usage scales linearly at an almost 1:1 ratio with the dictionary size, regardless of thread count or the match finder used during creation.
The decompressor does not need to search for repeated patterns; it only needs to retain the sliding window to look up previous references specified by the compressed stream. Consequently, decompressing an archive created with a 64 MB dictionary requires roughly 65 MB to 70 MB of RAM, and a 1024 MB dictionary requires just over 1 GB of RAM.