How LZMA Compression Works in 7-Zip

The Lempel-Ziv-Markov chain Algorithm (LZMA) is the flagship compression method behind 7-Zip's native .7z format, renowned for its exceptionally high compression ratios. This article explains the internal mechanics of LZMA, detailing how it combines an enhanced sliding-dictionary scheme, context-dependent probability modeling, and a precision range encoder to pack data far more tightly than traditional algorithms like Deflate.

The Two-Stage Architecture

LZMA operates as a two-stage stream compressor. When 7-Zip processes an uncompressed file, the data passes through:

  1. A modified LZ77 dictionary encoder, which detects and replaces repetitive byte sequences with pointer references.
  2. A Range Coder paired with Markov modeling, which performs context-based entropy encoding on the resulting symbols and literals.

Stage 1: The Enhanced LZ77 Dictionary

Standard LZ77 algorithms identify repeated sequences of bytes and replace them with pairs indicating length and backward distance. LZMA refines this concept in several key ways:

  • Massive Dictionary Sizes: While standard ZIP (Deflate) relies on a 32 KB sliding window, 7-Zip's LZMA implementation allows dictionary sizes ranging from 64 KB up to several gigabytes (commonly 16 MB to 64 MB by default). This massive window enables 7-Zip to find recurring patterns spaced very far apart in large archives.
  • Match Finders: To search through gigabytes of history without stalling, 7-Zip uses fast match-finding data structures, primarily Hash Chains (HC) and Binary Trees (BT). Binary tree variants (such as BT4) evaluate potential matches more exhaustively, producing smaller files at the cost of higher CPU and memory usage during compression.
  • Repeated Offsets: LZMA tracks the four most recently used match distances in a small buffer. If a pattern repeats at an offset identical to one of the last four matches, 7-Zip encodes a short reference to that recent distance instead of transmitting a full, multi-byte distance value.

Stage 2: Context Modeling (Markov Chains)

Once data is split into literals (unmatched bytes) and match descriptors (lengths and distances), LZMA uses a Markov chain mechanism to predict upcoming bits:

  • State Machine: LZMA maintains an internal state variable that tracks the sequence of recent token types (e.g., whether the last token was a literal, a match, or a repeated match). This state dictates the probability tables used for upcoming tokens.
  • Literal Context Properties (lc, lp, pb): 7-Zip optimizes literal byte encoding using three configurable parameters:
    • lc (Literal Context bits): Selects probability models based on the high bits of the previous byte.
    • lp (Literal Position bits): Selects probability models based on the absolute byte position modulo \(2^{lp}\) (useful for structured or aligned data like 32-bit executables).
    • pb (Position Bits): Regulates match probability models based on position alignment.

By maintaining separate probability distributions for distinct contexts, the algorithm avoids mixing data patterns (such as code versus text), maximizing prediction accuracy.

Stage 3: The Range Coder

The final output stage replaces classic Huffman coding with a binary Range Coder.

Unlike Huffman coding, which assigns an integer number of bits to each symbol, a range coder assigns fractional bits by narrowing a mathematical interval between 0 and 1. If a bit has a 99% probability of occurring, the range coder represents it with a minute fraction of a bit, approaching theoretical Shannon entropy limits. The probability tables constantly adapt as new data is processed, ensuring optimal bit efficiency across varying file types.

Asymmetric Performance Characteristics

A core characteristic of LZMA inside 7-Zip is its asymmetry. Compression requires substantial memory and processor cycles to navigate large binary search trees and compute optimal parsing paths. Decompression, by contrast, requires negligible CPU overhead and only enough RAM to hold the sliding dictionary, allowing 7-Zip archives to extract rapidly across virtually any hardware environment.