How 7-Zip Compresses Unallocated Disk Space
This article explains how 7-Zip handles unallocated disk space when creating archives of disk images and storage volumes. It covers how 7-Zip’s primary algorithms treat raw data streams, the difference between zero-filled and "dirty" unallocated sectors, and how file system parsing determines whether free space is compressed or omitted entirely.
When 7-Zip compresses a raw disk image file (such as
.img, .dd, or .vhd), it treats
the disk image as a sequential stream of bytes rather than an active
file system. Because it does not run real-time file system maintenance
during block-level compression, its ability to compress unallocated
space depends entirely on the state of the bytes residing in those
unallocated sectors.
Zero-Filled Free Space
If the unallocated space consists of clean, zeroed sectors
(represented as continuous sequences of 0x00), 7-Zip’s
default LZMA and LZMA2 compression algorithms handle it with extreme
efficiency. LZMA uses dictionary-based encoding combined with a range
coder. When the algorithm encounters repeated patterns, such as millions
of consecutive zeros, it replaces the vast sequences of identical bytes
with tiny distance-and-length references. As a result, hundreds of
gigabytes of zero-filled unallocated space can be compressed down to a
few kilobytes in the final .7z file.
"Dirty" Unallocated Space
In standard operating systems, deleting a file removes its entry from the file system table (such as the Master File Table in NTFS) without physically overwriting the underlying sectors. These sectors become unallocated to the OS, but the previous data remains. When 7-Zip compresses a raw sector-by-sector disk image containing this "dirty" free space, it cannot distinguish between active files and abandoned data. It attempts to compress the old, deleted file fragments using standard pattern matching. If the leftover data consists of compressed files, encrypted blocks, or high-entropy media formats, 7-Zip cannot reduce their size effectively, resulting in a significantly larger archive.
File-Level Archive Extraction
The behavior changes if 7-Zip is used to read an existing disk image directly rather than compressing it as a raw file. 7-Zip possesses built-in parsers for file systems like NTFS, FAT, ext2/3/4, and ISO. When you open a disk image within 7-Zip and extract or archive its contents at the file level, 7-Zip reads the partition metadata directly. In this mode, it bypasses unallocated sectors completely, copying only the files explicitly marked as active and discarding the free space entirely.
Maximizing Compression on Raw Images
To achieve minimal file sizes when capturing full raw disk images
with 7-Zip, users frequently "zero out" free drive space prior to
creating the image. Running utilities such as sdelete on
Windows or zerofree on Linux writes consecutive zeros
across all unallocated sectors. When 7-Zip subsequently archives the raw
volume, the LZMA/LZMA2 engine detects the uniform byte streams and
collapses the unallocated space nearly completely.