How Unrar Decompresses Intel x86 Executables

Unrar processes archives compressed with the Intel executable compression algorithm by first extracting the raw stream using standard dictionary decompression and then passing the resulting data through a dedicated inverse post-processing filter. This filter scans the byte stream for specific x86 branching instructions—specifically relative CALL and JMP opcodes—and mathematically converts previously normalized absolute memory addresses back into their original relative offsets.

The Purpose of Intel Executable Pre-Processing

Standard LZSS and Huffman-based algorithms depend on recurring byte sequences to achieve high compression ratios. In compiled Intel x86 binaries, repeated calls to the same function generate different byte sequences because the target address is encoded as a relative offset dependent on the instruction's position in memory.

To overcome this, the RAR compressor applies an executable filter prior to main compression. It translates relative displacement values into absolute virtual addresses, ensuring that calls to the same function produce identical byte patterns. Consequently, unrar must reverse this process upon extraction to recreate the original executable code.

Stream Decoding and Filter Detection

When unpacking an archive block, unrar reads the header flags to identify whether any pre-processing filters were attached to the data block:

  1. Entropy and Dictionary Decompression: The compressed data is initially decoded through standard Huffman tables and the LZSS sliding window.
  2. Filter Queueing: If an executable filter is marked, unrar records the filter type (such as FILTER_E8 or FILTER_E8E9), the starting offset within the unpacked buffer, and the block length.
  3. Buffer Management: The decompressed bytes are placed into a circular unpack buffer before being written to the output file.

Reversing the Address Transformation

Once the intermediate data sits in memory, the unpack engine invokes its executable filter function (commonly implemented in the Unpack::ExecuteFilterE8 or Unpack::UnstoreFilter routines in the unrar source code). The process operates sequentially through the buffer:

  1. Opcode Scanning: The algorithm scans byte by byte for the x86 opcodes 0xE8 (relative CALL) and 0xE9 (relative JMP).
  2. Address Extraction: When an opcode is encountered, the subsequent 4 bytes are read as a 32-bit little-endian integer representing the normalized address.
  3. Range and Validity Checks: To prevent modifying data that resembles an instruction opcode by coincidence (such as bitmap data or string constants), unrar applies specific boundary checks:
    • It checks whether the address exceeds specified thresholds (e.g., negative offsets or addresses beyond the expected memory space).
    • In standard RAR implementations, if the top byte of the address matches a specific flag byte, the engine adjusts or skips the value to handle negative jump displacements correctly.
  4. Inverse Math Calculation: The engine subtracts the current file position from the extracted absolute address: RelativeOffset = AbsoluteAddress - CurrentFileOffset
  5. Byte Substitution: The calculated 32-bit relative offset replaces the original 4 bytes in the buffer directly following the opcode.
  6. Pointer Advancement: The scanner advances past the 4-byte address to prevent re-evaluating the replacement bytes as new opcodes, continuing the scan until the designated filter length is exhausted.

Final Output

After the filter completes its pass across the target block, the modified buffer contains the exact machine code of the original binary. unrar then verifies the block's CRC or BLAKE2 checksum against the archive header and flushes the buffer to disk.