How to Port 7-Zip to a New CPU Architecture

Porting 7-Zip to a new CPU architecture requires adapting its core C and C++ source code, addressing platform-specific constraints like data alignment and endianness, and handling architecture-specific assembly optimizations. This guide outlines the essential steps to prepare the codebase, configure the build toolchain, transition platform-dependent code, and validate the resulting binaries for stability and performance.

1. Obtain and Analyze the Codebase

Acquire the latest 7-Zip source code or the LZMA SDK from the official distribution. Identify the core components:

  • C Core: Contains compression and decompression algorithms such as LZMA, LZMA2, Deflate, and basic cryptographic primitives.
  • C++ Wrappers and UI: Implements archive management, archive formats (7z, ZIP, TAR), and file system abstractions.
  • Assembly Routines: Highly optimized ASM implementations for CRC calculation, AES encryption, SHA hashing, and LZMA state transitions.

2. Configure the Toolchain and Build System

Set up a cross-compiler or native compiler (such as GCC, Clang, or a vendor-specific toolchain) targeting the new architecture. 7-Zip typically uses customized makefiles (nmake on Windows or GNU make for POSIX/p7zip-derived builds).

  • Define the target architecture flags in the makefiles (e.g., -march, -mcpu, -mabi).
  • Establish appropriate macro definitions (such as _7ZIP_ST for single-thread builds or custom architecture macros) to control conditional compilation.

3. Handle Architectural Characteristics

Ensure the core algorithms respect the hardware specifications of the new architecture:

  • Endianness: The standard LZMA format is inherently little-endian. If porting to a big-endian architecture, ensure endian conversion macros (such as SetUi32 and GetUi32) correctly swap byte orders during file header processing and stream encoding.
  • Memory Alignment: Many RISC architectures enforce strict memory alignment or suffer high penalties for unaligned memory accesses. Audit and replace direct unaligned pointer casts with safe read/write macros (e.g., GetUi32 via byte-wise shifts or memcpy).
  • Word Size and Data Models: Confirm that assumptions regarding 32-bit (ILP32) vs. 64-bit (LP64 or LLP64) data sizes match variable declarations, particularly for buffer pointers and 64-bit file offsets (UInt64).

4. Manage Assembly and Hardware Acceleration

7-Zip achieves maximum throughput by leveraging architecture-specific vector extensions and assembly code (e.g., x86 SSE/AVX, ARM NEON, or RISC-V Vector).

  • Initial Bootstrap: Disable architecture-specific assembly by undefining hardware acceleration macros (e.g., compile with generic C implementations for CRC and AES first).
  • Optimization Phase: Port or rewrite critical performance bottlenecks into native assembly or compiler intrinsics. Focus on:
    • CRC-32 and CRC-64: Implement using native CRC instructions if supported by the CPU.
    • AES and SHA-256: Use native cryptographic extensions if available.
    • LZMA Match Finder: Optimize the sliding-dictionary search routines using SIMD operations.

Compile the standalone command-line client (7z or 7za) first, as it contains fewer dependencies on higher-level system APIs.

  • Verify that the runtime linking step successfully resolves standard library routines (memory allocation, threading primitives like pthreads or Win32 threads).
  • Resolve any undefined symbols resulting from missing platform-specific assembly functions.

6. Validate and Benchmark

Perform rigorous testing to guarantee data integrity and measure performance:

  • Integrity Testing: Compress and decompress various payloads (binary, text, sparse files) and verify checksums against standard references to confirm byte-for-byte compatibility.
  • Edge-Case Validation: Test maximum dictionary sizes and multi-threaded compression to detect race conditions or memory alignment crashes.
  • Built-in Benchmark: Run the native benchmark command (7za b) to evaluate compression and decompression speeds in Million Instructions Per Second (MIPS), ensuring the new build performs efficiently on the target hardware.