How Unrar Uses Recovery Records to Fix Archives

RAR archives often face file corruption during downloads, transfers, or storage failures, but they possess a built-in resilience mechanism known as recovery records. When an archive is created with this feature, redundant mathematical parity data is appended to the file. Utility tools like unrar parse this extra data to detect damaged sectors, reconstruct missing byte sequences, and generate a fully repaired, extractable archive.

The Structure of Recovery Records

A RAR recovery record relies on Reed-Solomon error-correcting codes, a standard algorithm used in optical media, RAID arrays, and data transmissions. When a user creates an archive, WinRAR or RAR calculates parity data based on the original contents and packages it into distinct RR (Recovery Record) blocks.

This recovery data typically ranges from 1% to 10% of the total archive size. It does not store duplicate copies of files; instead, it contains mathematical parity formulas derived from the original bytes.

How unrar Identifies Corruption

During standard extraction or a repair operation (using the command unrar r archive.rar), the tool reads the archive's internal structure and checks individual blocks against their stored Cyclic Redundancy Check (CRC32) or BLAKE2 checksums:

  1. Header Validation: unrar inspects file headers to verify that archive metadata is intact.
  2. Checksum Verification: It reads through the data streams and recalculates checksums. When a calculated checksum does not match the header's reference value, unrar flags those specific sectors as corrupt or missing.
  3. Recovery Chunk Mapping: unrar scans the end of the archive to locate the recovery record block and determine how many recovery sectors are available.

The Reconstruction Process

Once unrar identifies the location and volume of corrupted data, it applies Reed-Solomon erasure decoding:

  1. Treating Data as Equations: The recovery algorithm treats data blocks as variables in a system of linear equations. If an archive has \(N\) original blocks and \(M\) recovery blocks, any \(N\) valid blocks out of the total \(N + M\) blocks can be used to solve the equation.
  2. Solving for Missing Bytes: Corrupted sectors are treated as unknowns. As long as the total number of corrupted blocks does not exceed the total number of recovery blocks (\(M\)), unrar can mathematically solve for the missing values.
  3. Assembling the Fixed File: After resolving the missing bytes, unrar writes the corrected data to a new archive file (typically prefixed as fixed.archive.rar), leaving the original corrupted file untouched for safety.

Capabilities and Limitations

The success of the repair process depends entirely on the ratio of damage to recovery record size:

  • Recoverable Scenarios: If an archive has a 5% recovery record, it can tolerate up to 5% total data loss or corruption anywhere across the archive—whether it is isolated bit-rot or contiguous byte destruction.
  • Unrecoverable Scenarios: If the physical damage exceeds the size of the recovery record (for example, 8% corruption with only a 5% record), unrar cannot solve the mathematical matrix. In such cases, the tool alerts the user that the recovery record is damaged or insufficient, and the data cannot be rebuilt.