7-Zip Executable Compression Filters Explained
When compressing executable binaries, 7-Zip relies on specialized preprocessing filters known collectively as Branch-Call-Jump (BCJ) converters to substantially boost compression ratios. Rather than compressing the data directly, these filters normalize machine code addresses before passing the modified stream to general-purpose compressors like LZMA or LZMA2. This article explains how 7-Zip's BCJ and BCJ2 algorithms operate, the technical challenges they address in raw binaries, and the architectural variants available.
The Challenge of Compressing Machine Code
Compiled executable files (such as .exe,
.dll, or .so) are inherently difficult for
dictionary-based algorithms like LZMA to compress efficiently. Even when
identical functions or routines appear multiple times throughout an
application, the machine instructions that call them often use relative
offset addressing rather than absolute addresses.
Because relative offsets change depending on where the calling instruction is located in memory, identical operations produce different byte sequences. This fragmentation prevents dictionary-based algorithms from identifying recurring patterns.
The BCJ Filter (x86)
The standard BCJ filter (Branch, Call, Jump) solves this issue for 32-bit x86 machine code. It acts as a bidirectional preprocessor:
- Detection: The filter scans the binary for x86
branch instructions, specifically relative
CALL(opcode0xE8) and relative unconditionalJMP(opcode0xE9) instructions. - Address Normalization: When a branch instruction is encountered, the filter converts the relative offset into an absolute target address.
- Pattern Alignment: Because repeated calls to the same function now share the exact same absolute address bytes regardless of where the calls originate, the LZMA dictionary encounters identical, repetitive sequences, drastically improving redundancy detection.
During decompression, the inverse BCJ filter automatically translates the absolute addresses back into their original relative offsets, restoring the binary to an identical state.
The BCJ2 Filter (Advanced x86)
For higher compression ratios, 7-Zip offers BCJ2, an advanced four-stream split filter designed for x86 binaries:
- Main Stream: Contains all original code and data bytes, with the call and jump address payloads stripped out.
- Call Stream: Gathers the absolute addresses
extracted from relative
CALLinstructions. - Jump Stream: Gathers the absolute addresses
extracted from relative
JMPinstructions. - Control Stream: A bitstream indicating whether each byte in the main stream corresponds to a branch instruction or regular code.
By separating continuous execution data from address pointers, 7-Zip can compress each stream using dedicated LZMA instances and optimized settings. Addresses are typically grouped together and share common higher-order bytes, allowing LZMA to compress them far more densely than it could if they remained interspersed with code.
Architecture-Specific BCJ Variants
Modern versions of 7-Zip extend the BCJ concept beyond standard x86 architectures to handle instructions for other processor designs:
- ARM: Targets 32-bit ARM instructions, converting
24-bit relative branch and branch-with-link
(
B/BL) offsets. - ARMT (ARM Thumb): Tailored for compact 16-bit and 32-bit Thumb instruction sets used widely in embedded systems and mobile binaries.
- IA-64: Optimizes Intel Itanium 128-bit instruction bundles, normalizing branch targets within multi-instruction templates.
- PPC: Designed for 32-bit PowerPC binaries,
translating relative branch operations
(
B/BL). - SPARC: Preprocesses SPARC 30-bit displacement branch instructions.
- RISC-V: Newer iterations of 7-Zip also include support for normalizing jump and branch targets within RISC-V architectures.
These architecture-specific filters ensure that modern multi-platform binaries benefit from the same structural normalization that standard x86 binaries receive during compression.