How 7-Zip Detects Corrupted Encrypted Archives
This article explains the underlying mechanisms 7-Zip uses to identify damaged or altered data blocks within password-protected archives. When dealing with encrypted archives, detecting corruption relies on a combination of cryptographic key verification, decompression stream validation, and pre-computed integrity checksums like CRC-32 and SHA-256. Together, these layers allow 7-Zip to distinguish between an incorrect password, a damaged archive header, and a corrupted payload block.
The Decryption and Extraction Pipeline
To understand how 7-Zip flags corrupted data, it is necessary to examine the order in which encrypted files are processed:
- Key Derivation: 7-Zip uses the PBKDF2 (Password-Based Key Derivation Function 2) algorithm with HMAC-SHA-256 over many iterations to turn the user's password into a 256-bit AES encryption key.
- Decryption: The raw, compressed data blocks are decrypted using the AES-256-CBC algorithm.
- Decompression: The decrypted stream is passed to the decompression engine (such as LZMA or LZMA2).
- Integrity Check: The uncompressed output is checked against pre-stored integrity hashes.
Corruption can be detected at several points along this pipeline.
Decompression Stream Validation
7-Zip archives are typically compressed before they are encrypted. When a corrupted block is decrypted using the correct AES key, the resulting output will contain altered bytes rather than the original compressed stream.
Compression algorithms like LZMA rely on strict syntax rules, distance-length pairs, and probability models. When the LZMA decoder encounters corrupted bits in the decrypted stream, it often encounters invalid codes, out-of-range dictionary references, or broken markers. When this occurs, the decompression engine immediately halts and reports a data error.
CRC-32 and SHA-256 Checksums
Even if a corrupted block manages to pass through the decompression engine without causing an outright syntax crash, 7-Zip maintains a second, definitive line of defense: embedded integrity checks.
- Payload CRC-32: 7-Zip stores a standard 32-bit Cyclic Redundancy Check (CRC-32) for every uncompressed file inside the archive metadata. Once decompression is complete, 7-Zip computes the CRC-32 of the newly extracted file. If this calculated value does not match the stored value, the file is marked as corrupted.
- Block and Header Checksums: The 7z format groups files into "folders" (solid blocks) and headers. These structures also contain their own internal CRC values. If any individual block inside the archive structure has flipped bits, the block-level CRC calculation will fail.
Header Encryption and Early Detection
When an archive is created with the "Encrypt file names" option enabled, 7-Zip encrypts the archive's central directory (the main header).
The header contains its own verification markers and CRC-32 checks. If data corruption occurs within an encrypted header block:
- Decryption will yield invalid header bytes.
- 7-Zip will fail to parse the archive structure.
- The software will report that the archive cannot be opened or is corrupted before the user even attempts to extract individual files.
Distinguishing Between Wrong Passwords and Data Errors
Because 7-Zip uses AES in CBC mode without an authenticated encryption mode (like AES-GCM), the decryption step itself does not intrinsically authenticate the data. Instead, 7-Zip relies on the downstream checks to determine the nature of a failure:
- If an extraction fails immediately on header parsing or the very first bytes of a stream, 7-Zip typically flags a "Wrong Password" error.
- If a significant portion of an archive extracts successfully and then fails midway through a solid block or at the final CRC check, 7-Zip flags a "Data Error" or "CRC Failed" message, indicating physical corruption in the underlying archive file rather than an authentication failure.