How PPMd Improves Text Compression in 7-Zip

This article examines how the Prediction by Partial Matching (PPMd) algorithm enhances text file compression within the 7-Zip archiver. It covers the underlying statistical context-modeling mechanics of PPMd, highlights why it significantly outperforms traditional dictionary-based algorithms like LZMA and Deflate on textual data, and outlines the practical configuration options available in 7-Zip to achieve the smallest possible file sizes.

What is the PPMd Algorithm?

PPMd is an advanced implementation of the Prediction by Partial Matching compression algorithm, originally adapted by Dmitry Shkarin. Unlike general-purpose dictionary algorithms that replace repeated character sequences with back-references, PPMd uses statistical modeling to predict incoming characters based on the context of preceding characters. It dynamically tracks letter frequencies in natural language or source code and encodes highly predictable characters using fewer bits via arithmetic coding.

Why PPMd Excels at Text Compression

Natural language and structured text files (such as source code, XML, JSON, and server logs) follow rigid grammatical, syntactic, and probabilistic rules. For example, in English text, the letter "h" frequently follows "t", and the letter "u" almost always follows "q".

While dictionary coders like LZMA or Deflate require exact recurring strings to achieve high compression, PPMd predicts individual symbols based on variable context lengths:

  • Contextual Prediction: If the full context (such as a 6-character sequence) has not been seen before, PPMd smoothly drops down to a smaller context (e.g., 5, 4, or 3 characters) using an escape mechanism to maintain high prediction accuracy.
  • Arithmetic Encoding: Characters with the highest probability given the prior context are mapped to tiny fractional bit spaces rather than full bytes, squeezing maximum entropy from plain text.
  • Absence of Dictionary Limits: PPMd does not depend on a sliding window limit in the same way LZ-based algorithms do, meaning it can model complex linguistic grammar across an entire document without needing duplicate multi-word phrases.

PPMd vs. LZMA in 7-Zip

7-Zip uses LZMA and LZMA2 as its default compression methods because they provide an optimal balance between compression ratio, high-speed decompression, and low memory usage across mixed file formats. However, on pure text payloads, PPMd consistently outperforms LZMA:

  • Compression Ratio: PPMd often achieves 10% to 30% smaller archives on natural language texts, books, and code repositories compared to standard LZMA/LZMA2.
  • Decompression Asymmetry: Unlike LZMA, which decompresses rapidly regardless of compression effort, PPMd requires symmetric computing power. Decompressing a PPMd archive requires virtually the same CPU time and memory as compressing it.

Configuring PPMd in 7-Zip

To utilize PPMd in 7-Zip, select the .7z or .zip format in the "Add to Archive" dialog and change the compression method to PPMd. Two main parameters determine its efficiency:

  • Model Order: Defines the maximum length of previous characters used to predict the next character (range: 2 to 32). An order between 8 and 16 is typically optimal for natural text and code. Setting the order too high can cause diminishing returns and increase memory usage without improving file size.
  • Memory Size: Dictates the amount of RAM allocated to hold the statistical context tree. Allocating between 64 MB and 256 MB is sufficient for most text corpora, though larger sets can use up to 1024 MB or more to prevent the algorithm from having to purge its context cache.