How 7-Zip Organizes Files in Solid Archives

This article explains how 7-Zip sorts and structures files within solid archives to maximize compression efficiency. Solid archiving combines multiple files into a single continuous data stream, allowing the compression algorithm to spot recurring patterns across separate files. To ensure that repetitive data falls within the compressor's sliding memory window, 7-Zip systematically sorts files by extension, file path, and size, dramatically increasing match redundancy and reducing the final archive size.

The Mechanics of Solid Archiving

In a non-solid archive (like a standard ZIP file), each file is compressed individually. This prevents data sharing across files, meaning identical phrases or structures in separate files are compressed repeatedly from scratch.

In a solid 7z archive, the files are concatenated into a unified stream before the compression algorithm (typically LZMA or LZMA2) processes them. LZMA utilizes a sliding dictionary—a finite block of memory (such as 32 MB, 64 MB, or larger)—to identify repeated byte sequences. If identical data appears outside this memory window, the compressor cannot reference it. Therefore, the physical order in which files enter the stream determines whether cross-file redundancy is captured.

The File Sorting Heuristic

To keep similar data within the sliding dictionary window, 7-Zip reorders the file list before compression using a specific hierarchy:

  1. File Extension: 7-Zip groups files by their file extensions (e.g., all .cpp files together, all .png files together). Because files with the same extension typically share internal syntax, headers, or structural patterns, placing them side-by-side yields the highest concentration of duplicate byte sequences.
  2. File Name and Directory Path: Within the same extension group, files are sorted alphabetically by path and name. Files that share similar names or reside in the same directories often belong to the same project or contain similar revisions, making them more likely to share substantial chunks of data.
  3. File Modification Date or Size: Secondary attributes are used as tie-breakers to ensure consistent, predictable ordering across different runs.

Sliding Window Locality

The primary objective of this sorting process is to exploit temporal and spatial locality relative to the dictionary size. If an archive contains ten versions of the same 5 MB text document scattered across different directories, sorting solely by folder structure might place them gigabytes apart in the stream, rendering the dictionary window ineffective. By sorting by extension and name first, all ten versions are fed to the LZMA engine sequentially. The engine compresses the first version fully and then compresses the subsequent nine versions as minute delta differences, achieving near-optimal compression ratios.

User Control Over Sorting

7-Zip allows users to modify this behavior using command-line parameters. By default, 7-Zip prioritizes sorting by extension. However, using the -mqs (sort by name) switch forces the archiver to preserve the directory and filename order regardless of extension. This can be beneficial when an archive consists of paired, tightly coupled files with different extensions (such as pairs of .wav and .cue files, or .c and .h files) where cross-format similarity exceeds the similarity between unrelated files of the same format.