How Unrar Decompresses Intel x86 Executables
Unrar processes archives compressed with the Intel executable
compression algorithm by first extracting the raw stream using standard
dictionary decompression and then passing the resulting data through a
dedicated inverse post-processing filter. This filter scans the byte
stream for specific x86 branching instructions—specifically relative
CALL and JMP opcodes—and mathematically
converts previously normalized absolute memory addresses back into their
original relative offsets.
The Purpose of Intel Executable Pre-Processing
Standard LZSS and Huffman-based algorithms depend on recurring byte sequences to achieve high compression ratios. In compiled Intel x86 binaries, repeated calls to the same function generate different byte sequences because the target address is encoded as a relative offset dependent on the instruction's position in memory.
To overcome this, the RAR compressor applies an executable filter
prior to main compression. It translates relative displacement values
into absolute virtual addresses, ensuring that calls to the same
function produce identical byte patterns. Consequently,
unrar must reverse this process upon extraction to recreate
the original executable code.
Stream Decoding and Filter Detection
When unpacking an archive block, unrar reads the header
flags to identify whether any pre-processing filters were attached to
the data block:
- Entropy and Dictionary Decompression: The compressed data is initially decoded through standard Huffman tables and the LZSS sliding window.
- Filter Queueing: If an executable filter is marked,
unrarrecords the filter type (such asFILTER_E8orFILTER_E8E9), the starting offset within the unpacked buffer, and the block length. - Buffer Management: The decompressed bytes are placed into a circular unpack buffer before being written to the output file.
Reversing the Address Transformation
Once the intermediate data sits in memory, the unpack engine invokes
its executable filter function (commonly implemented in the
Unpack::ExecuteFilterE8 or
Unpack::UnstoreFilter routines in the unrar
source code). The process operates sequentially through the buffer:
- Opcode Scanning: The algorithm scans byte by byte
for the x86 opcodes
0xE8(relativeCALL) and0xE9(relativeJMP). - Address Extraction: When an opcode is encountered, the subsequent 4 bytes are read as a 32-bit little-endian integer representing the normalized address.
- Range and Validity Checks: To prevent modifying
data that resembles an instruction opcode by coincidence (such as bitmap
data or string constants),
unrarapplies specific boundary checks:- It checks whether the address exceeds specified thresholds (e.g., negative offsets or addresses beyond the expected memory space).
- In standard RAR implementations, if the top byte of the address matches a specific flag byte, the engine adjusts or skips the value to handle negative jump displacements correctly.
- Inverse Math Calculation: The engine subtracts the
current file position from the extracted absolute address:
RelativeOffset = AbsoluteAddress - CurrentFileOffset - Byte Substitution: The calculated 32-bit relative offset replaces the original 4 bytes in the buffer directly following the opcode.
- Pointer Advancement: The scanner advances past the 4-byte address to prevent re-evaluating the replacement bytes as new opcodes, continuing the scan until the designated filter length is exhausted.
Final Output
After the filter completes its pass across the target block, the
modified buffer contains the exact machine code of the original binary.
unrar then verifies the block's CRC or BLAKE2 checksum
against the archive header and flushes the buffer to disk.