How 7-Zip Processes Solid Block Sizes
When creating an archive in the 7z format, 7-Zip uses a technique called solid compression to bundle multiple files together into a single continuous data stream before applying compression algorithms like LZMA or LZMA2. The solid block size setting determines the maximum uncompressed volume of data allowed in one continuous stream before a new block is initiated. By configuring this parameter, users directly manipulate how 7-Zip balances overall compression efficiency against multi-threading capabilities, decompression speeds, and archive fault tolerance.
File Sorting and Concatenation
Before compression begins, 7-Zip sorts files based on metadata, prioritizing file extensions, names, and modification dates. This organization ensures that identical or closely related file types are grouped adjacent to one another.
Once sorted, 7-Zip treats these files not as individual files, but as a single contiguous stream of bytes. This allows the compression dictionary to reference redundant data patterns that appear across entirely different files, which drastically reduces the final archive size when dealing with collections of similar files (such as source code or text documents).
Solid Block Boundary Enforcement
During the creation process, 7-Zip tracks the cumulative uncompressed size of the files being read into the active stream.
- Reaching the Limit: Once the uncompressed byte count crosses the designated solid block size threshold (such as 2 GB, 4 GB, or a custom value), 7-Zip closes the current solid block.
- Flushing the Stream: Closing a block flushes the compression state and resets the dictionary window.
- Starting a New Block: The next file in the queue begins a completely separate, independent solid block, repeating the sorting-and-streaming process.
If the solid block size is set to "Solid" (infinite), 7-Zip packs all files into a single unbroken block. Conversely, setting it to "Non-solid" forces the archiver to treat each file as its own independent block.
Multi-Threading and CPU Utilization
Solid block sizes directly govern how 7-Zip utilizes modern multi-core processors:
- Independent Threads: Because a solid block relies on past data within that exact stream, a single solid block cannot be compressed across multiple threads using the legacy LZMA algorithm.
- LZMA2 Handling: Under LZMA2, 7-Zip can split data into chunks across threads within a block, but distinct solid blocks are inherently independent.
- Parallel Block Processing: When multiple solid blocks are defined, 7-Zip can assign each block to a separate CPU thread, allowing distinct groups of files to be compressed completely in parallel without thread contention.
Trade-Offs in Block Sizing
Adjusting the solid block size alters archive performance characteristics:
- Compression Ratio: Larger solid blocks provide higher compression ratios because the sliding dictionary has a continuous flow of data from which to find matching patterns.
- Extraction Speed: To extract a single file from inside a solid block, 7-Zip must decompress all preceding data within that specific block up to the target file. Smaller solid blocks significantly speed up random file access and selective extraction.
- Data Recovery: If corruption occurs within a solid block, all subsequent data within that same block may become inaccessible. Limiting the block size isolates corruption to a single block, protecting the integrity of the rest of the archive.