How 7-Zip Compresses ARM and IA-64 Binaries

7-Zip achieves superior compression ratios on ARM and IA-64 executables by pairing its core LZMA and LZMA2 algorithms with specialized pre-processing filters known as BCJ (Branch, Call, Jump) converters. Rather than attempting to compress raw binary data directly, 7-Zip first normalizes relative branch target addresses across the code, transforming disparate byte sequences into repetitive patterns that dictionary-based compression can pack far more efficiently.

The Challenge of Executable Code

Standard dictionary compressors like LZMA rely on finding repeated sequences of bytes. Compiled code from architectures like ARM and Intel IA-64 (Itanium) contains numerous subroutine calls and jumps. Because these instructions typically use relative offsets rather than fixed destinations, calling the same function from multiple locations produces completely different machine code values.

For instance, two identical calls to a single utility function located at different points in memory generate different offset values in their binary representation. This variation disrupts pattern recognition, degrading the efficiency of sliding-dictionary algorithms.

How Pre-Processing Filters Work

To solve this issue, 7-Zip introduces a reversible pre-processing pass prior to compression. When processing an archive, 7-Zip detects or is instructed to use architecture-specific filters:

  1. Address Normalization: The filter scans the byte stream, detects branch opcodes, and converts relative jump targets into absolute virtual addresses.
  2. Repetition Amplification: Because absolute targets point to the exact same destination regardless of where the call originates, identical function calls across the binary are converted into identical byte strings.
  3. LZMA Compression: The normalized stream is passed to the LZMA/LZMA2 engine, which easily detects these recurring patterns and compresses them tightly.
  4. Reversible Extraction: During decompression, the inverse filter converts absolute addresses back into their original relative offsets, ensuring the restored executable is bit-identical to the source.

ARM Filter Mechanics

ARM architectures utilize both fixed-length 32-bit instructions (standard ARM) and 16-bit/32-bit variable-length instructions (Thumb and Thumb-2).

The 7-Zip ARM filter targets conditional and unconditional branch instructions, primarily B (Branch) and BL (Branch with Link). It scans 4-byte aligned boundaries to locate the distinctive bitmasks of ARM branch opcodes. Upon identifying a branch, it extracts the 24-bit immediate field that holds the relative offset, converts it to an absolute address relative to the current program counter, and rewrites the instruction in place. A dedicated Thumb filter performs an equivalent operation for Thumb-mode binary patterns.

IA-64 (Itanium) Filter Mechanics

The IA-64 architecture utilizes a Very Long Instruction Word (VLIW) format, which packages instructions into 128-bit bundles. Each bundle contains:

  • A 5-bit template specifying instruction types and execution execution stops.
  • Three 41-bit instruction slots.

Because IA-64 does not follow simple byte-aligned instruction boundaries, a generic byte filter cannot detect branch offsets. The 7-Zip IA-64 filter decodes the 128-bit bundles, evaluates the 5-bit template to find which slots contain branch-type instructions, and extracts the branch targets directly from the 41-bit bitfields. It then normalizes these targets to absolute values and repacks the bits into the bundle.

Performance and Efficiency Impact

By transforming non-repeating relative addresses into identical absolute addresses, the ARM and IA-64 filters reduce the entropy of the machine code. This targeted pre-processing often reduces the final compressed size of large binary libraries and system images by 5% to 15% compared to raw LZMA compression alone, requiring minimal CPU overhead during both compression and decompression.