How 7-Zip Scales on AMD EPYC and Intel Xeon

This article explores how the open-source archiver 7-Zip scales across high-core-count enterprise processors, specifically AMD EPYC and Intel Xeon platforms. While 7-Zip’s native LZMA and LZMA2 compression algorithms offer built-in multithreading that benefits greatly from parallel processing, performance gains often hit diminishing returns on ultra-dense, multi-socket architectures due to memory bandwidth limits, cache contention, and non-uniform memory access (NUMA) overhead.

The Parallel Architecture of LZMA2

7-Zip scales across multiple cores primarily through the LZMA2 compression algorithm. Unlike the original LZMA format, which was largely serial, LZMA2 divides uncompressed data streams into independent blocks and distributes them across available CPU threads.

Because each thread can compress its designated block independently, compression workloads scale near-linearly on consumer chips and mid-tier enterprise processors up to 16 to 32 cores. However, decompression remains fundamentally less parallelized than compression, often utilizing only a fraction of the available threads because it depends heavily on block boundaries and sequential dependencies within the compressed stream.

Diminishing Returns at High Core Counts

When deployed on modern server chips featuring 64, 96, or 128 cores per socket (such as AMD EPYC Genoa/Bergamo or Intel Xeon Sapphire Rapids/Emerald Rapids), 7-Zip rarely scales linearly across all threads in a single process. Several hardware-level bottlenecks emerge:

  • Memory Bandwidth Saturation: Compression requires significant memory bandwidth to continuously read input streams and write hash tables. When dozens of threads compete simultaneously for RAM access, memory channels become a major bottleneck, regardless of raw compute headroom.
  • Cache Contention: LZMA uses large dictionary sizes to achieve high compression ratios. If the working set per thread exceeds the processor's L3 cache, performance drops sharply as threads are forced to fetch data from main memory.
  • NUMA and Inter-Socket Latency: Multi-die processors and dual-socket configurations split CPU cores into multiple NUMA nodes. When a single 7-Zip instance spawns threads across different compute dies or CPU sockets, high-latency interconnect traffic across AMD Infinity Fabric or Intel UPI degrades performance.

AMD EPYC vs. Intel Xeon in 7-Zip Scaling

Both processor families handle high-thread compression workloads differently:

  • AMD EPYC: With its chiplet architecture and massive pools of L3 cache—especially on "X" variants featuring 3D V-Cache—EPYC excels at keeping dictionary data close to the execution cores. However, scheduling 128 or more threads within a single task across multiple Core Complex Dies (CCDs) introduces inter-core synchronization latency.
  • Intel Xeon: Intel's monolithic and tiled architectures typically offer strong per-core throughput and balanced memory access. Xeon handles multi-threaded tasks efficiently, but configurations with smaller L3 cache allocations per core may experience earlier memory bandwidth saturation compared to high-cache EPYC alternatives.

Optimization Strategies for Server Deployments

To achieve maximum compression efficiency on high-core-count servers, executing a single 7-Zip process with maximum threads (e.g., -mmt128) is often counterproductive. The following adjustments yield better performance:

  1. Process Parallelism: Instead of running one archive job across 64+ cores, partition workloads into multiple concurrent jobs pinned to specific NUMA nodes or CPU sockets using tools like numactl on Linux or CPU affinity settings on Windows.
  2. Cap Thread Counts: In most LZMA2 configurations, sweet-spot efficiency peaks between 16 and 32 threads per archive job. Capping threads to this range avoids excessive thread management overhead.
  3. Balance Dictionary Size with RAM: Reduce dictionary sizes when using higher thread counts to ensure each active thread fits within the local cache tier, minimizing costly trips to system memory.