Impact of Word Size on 7-Zip Compression Ratio
Word size in 7-Zip defines the maximum length of identical byte patterns the LZMA and LZMA2 algorithms search for when compressing data. Choosing a larger word size generally improves the overall compression ratio by allowing the algorithm to replace longer repeating sequences with compact references, though it requires significantly more processing time. This article explains how word size functions within 7-Zip, how it influences archive size across different file types, and how to balance compression efficiency with performance.
What Word Size Means in 7-Zip
In 7-Zip, the "Word size" setting (often referred to as "fast bytes") controls the maximum match length evaluated by the match finder. For LZMA and LZMA2 compression methods, this parameter typically ranges from 5 to 273 bytes, with standard presets defaulting to values like 32 or 64.
When the compressor scans input data, it looks for sequences of bytes that match previously encountered data within the designated dictionary window. If a match is found, the word size setting determines how far the algorithm will continue comparing bytes to find the longest possible match before writing a back-reference token.
How Increasing Word Size Improves the Compression Ratio
The primary mechanism connecting word size to the compression ratio is pattern matching efficiency:
- Longer Back-References: A single back-reference token can represent up to the maximum configured word size. Replacing a 200-byte identical string with a single distance-length pair yields a significantly smaller footprint than breaking that sequence into multiple smaller matches.
- Higher Redundancy Capture: Files with large blocks of identical or near-identical content—such as source code repositories, log files, uncompressed databases, and repetitive structured text (XML, JSON, CSV)—benefit substantially from higher values, often achieving noticeably smaller archive sizes when pushed toward the maximum limit of 273.
Diminishing Returns and Data Limitations
While a larger word size can enhance compression, its effectiveness depends entirely on the nature of the target files:
- Incompressible or Unique Data: For files with low redundancy, such as already-compressed archives (ZIP, RAR), media formats (JPEG, MP3, MP4), or encrypted data, increasing the word size yields virtually zero improvement because long recurring byte sequences do not exist.
- Marginal Gains Beyond 64: For typical mixed-data workloads, the majority of repeating patterns are relatively short. Increasing the word size from 32 to 64 often produces a measurable drop in file size, but moving from 64 to 273 frequently yields diminishing returns, often shaving off less than 1% of the final archive size.
Performance Trade-Offs
The compression ratio gains provided by higher word sizes come with distinct computational costs:
- CPU Utilization and Speed: As word size increases, the match finder must perform significantly more byte-by-byte comparisons per block. Selecting the maximum word size of 273 can increase compression time by several orders of magnitude compared to the default setting of 32 or 64.
- Memory Usage: Word size has a minimal impact on RAM usage during compression, which is predominantly governed by dictionary size.
- Decompression Impact: Word size has zero negative impact on decompression speed or decompression memory requirements. Archives compressed with a word size of 273 decompress just as quickly as those compressed with a word size of 32.
Optimal Word Size Selection
To maximize efficiency, align the word size with your specific storage and time constraints:
- Word Size 32 to 48 (Fast/Normal): Ideal for general day-to-day archiving where speed is important and data types are mixed.
- Word Size 64 (Maximum Preset): Provides an optimal balance, capturing most meaningful patterns without causing excessive CPU bottlenecks.
- Word Size 128 to 273 (Ultra Preset): Best reserved for distribution packages, long-term cold storage, or highly repetitive text and raw binary data where minimizing storage size justifies lengthy compression times.