How 7-Zip Applies PPMd Memory Size Limits
This article examines how the 7-Zip compression utility enforces memory limits when utilizing the PPMd (Prediction by Partial Matching) algorithm. PPMd is a specialized statistical compression method primarily used for text. 7-Zip controls PPMd memory consumption through fixed upfront buffer allocation, a dedicated internal sub-allocator, and automatic model-reset mechanisms that strictly prevent the process from exceeding user-defined memory thresholds during both compression and decompression.
Pre-Execution Allocation
Unlike sliding-window algorithms such as LZMA, which dynamically
expand buffer usage based on dictionary sizes and multi-threading
requirements, PPMd operates within a strictly bounded, pre-allocated
memory block. When a user configures a memory limit—either via the 7-Zip
graphical interface or the command line (e.g.,
-m0=PPMd:mem=256m)—7-Zip reads this value prior to starting
the operation.
Upon execution, 7-Zip requests a single contiguous block of system
memory matching that exact value via its PPMd wrapper
(Ppmd7_Alloc or Ppmd8_Alloc, based on the PPMd
implementation version). If the operating system cannot provide this
contiguous block, the process halts immediately with an allocation error
rather than degrading into dynamic system heap expansion.
The Dedicated Sub-Allocator
Once the initial block of memory is assigned, PPMd ceases making
standard operating system memory calls (such as malloc or
HeapAlloc). Instead, it relies on a custom internal
sub-allocator.
PPMd models data by building a complex trie of context trees, tracking character frequencies and order histories. To store these nodes efficiently:
- The pre-allocated memory pool is partitioned into discrete memory units.
- The sub-allocator divides available memory into index-based free lists categorized by allocation unit size.
- Context nodes are created, split, and recycled strictly within this bounded space.
By shifting memory management away from the host operating system kernel to this internal sub-allocator, 7-Zip eliminates memory fragmentation issues and ensures that runtime execution cannot leak or grow beyond the initial boundary.
Memory Saturation and Model Resets
As high-order context models grow during the processing of large files, the pre-allocated memory buffer eventually fills up. When the sub-allocator detects that no free units remain, 7-Zip does not expand the buffer. Instead, it activates PPMd’s internal maintenance routines:
- Context Rescaling: The model halves or scales down context frequencies to free unused statistical counters.
- Model Pruning and Restart: If rescaling does not
recover sufficient space, the engine triggers a model restart mechanism
(
RestartModelRare). This purges stale, low-frequency context branches and resets the model to a baseline state while retaining only essential context paths.
Compression then continues using the newly freed segments inside the original memory envelope. This cycle repeats as often as necessary across large streams, locking the total runtime footprint at the chosen limit.
Decompression Memory Parity
7-Zip records the designated PPMd memory size directly into the archive stream metadata. During decompression, the extraction engine reads this parameter and pre-allocates an identical memory buffer. Because PPMd is a deterministic state-machine algorithm, the decompressor must follow the exact same sub-allocation, frequency scaling, and model restart steps as the compressor. If the extracting machine has less available RAM than specified in the archive header, the extraction process fails immediately at the initialization stage, protecting the host system from out-of-memory crashes.