How 7-Zip Multi-Threaded LZMA2 Decompression Works

Multi-threaded decompression in 7-Zip accelerates the extraction of LZMA2 archives by processing separate, independent data blocks simultaneously across multiple CPU cores. While traditional LZMA compression enforces a continuous data stream that limits decompression to a single thread, LZMA2 overcomes this restriction by dividing data into self-contained chunks. This article explains the underlying architecture of LZMA2, how 7-Zip identifies chunk boundaries, assigns tasks across worker threads, and reassembles decoded streams in the correct sequence.

The Architectural Shift from LZMA to LZMA2

Legacy LZMA uses a continuous sliding dictionary where every byte relies on the historical context of the preceding uncompressed data. Because a decoder cannot predict the state of the dictionary midway through the archive without decompressing everything prior, single-threaded execution is mandatory.

LZMA2 resolves this bottleneck by acting as a container format that segments data into a series of smaller chunks. Crucially, LZMA2 allows the encoder to reset the dictionary state at designated block boundaries. When an archive is created using multiple threads, 7-Zip writes these independent blocks into the container, laying the structural groundwork for parallel extraction.

Boundary Detection and Block Parsing

When decompressing an LZMA2 archive, 7-Zip begins by reading the metadata stored in the archive headers. This metadata outlines the structure of the compressed streams, identifying where discrete blocks start and end.

Because each thread requires a distinct segment to decode without waiting for adjacent data, 7-Zip scans for blocks that have either completely reset dictionaries or clearly defined contextual boundaries. Once these offsets are mapped, the archive reader divides the incoming data stream into discrete decompression jobs.

Worker Thread Allocation

7-Zip utilizes an internal thread pool managed by its core compression engine. The main thread acts as an orchestrator, distributing parsed data chunks to available worker threads:

  1. Chunk Distribution: Compressed blocks are fed into memory buffers assigned to idle worker threads.
  2. Isolated Decoding: Each worker thread runs its own instance of the LZMA2 decoding algorithm, maintaining an isolated dictionary state in RAM.
  3. Hardware Utilization: The number of active threads scales according to system thread limits or user-defined parameters, maximizing CPU utilization across multi-core systems.

In-Memory Buffering and Ordered Output

Although worker threads operate concurrently, the extracted files must be written to disk in their original sequential order to prevent data corruption. Because blocks vary in complexity, some threads finish decompressing faster than others regardless of their order in the original file.

To handle this discrepancy, 7-Zip uses a synchronized buffer queue:

  • Asynchronous Processing: Threads write their decompressed payloads into temporary memory buffers.
  • Sequenced Flushing: The orchestrator thread tracks the sequential index of each block. It only writes data to the target disk or file stream when the exact, next logical block is fully decoded.
  • Backpressure Management: If downstream disk write speeds lag behind CPU decoding, or if an early block takes longer to decode than subsequent ones, 7-Zip throttles the creation of new worker jobs to prevent excessive RAM consumption.

Through this combination of independent LZMA2 block structures, isolated decoding threads, and strict output sequencing, 7-Zip achieves high-throughput multi-threaded decompression without compromising data integrity.