7-Zip Big-Endian vs Little-Endian Architecture
This article examines how 7-Zip processes data across different processor architectures, focusing on its handling of big-endian and little-endian byte ordering. 7-Zip ensures full cross-platform compatibility by strictly enforcing little-endian serialization for archive headers and metadata, while using internal preprocessor macros and byte-swapping routines to translate data seamlessly on big-endian hardware.
Endianness Standards in the 7z Format
The native 7z archive specification defines all multi-byte integers—including file sizes, timestamps, property IDs, and CRC32 checksums—in little-endian byte order. Because the format specification does not vary based on the host platform, an archive created on an x86 little-endian machine must produce the exact same binary structures when created on a big-endian architecture like SPARC or PowerPC.
Architecture Detection in Source Code
7-Zip and the underlying LZMA SDK determine the host system's byte
order at compile time using preprocessor directives. The source defines
target macros such as MY_CPU_LE for little-endian
architectures and MY_CPU_BE for big-endian architectures.
These flags inspect predefined compiler macros (such as
__BYTE_ORDER__, __BIG_ENDIAN__, or
architecture-specific identifiers like __powerpc__ and
__sparc__) to compile the appropriate data access
pathways.
Memory Access and Byte Swapping
The handling difference between the two architectures lies primarily in memory serialization and deserialization:
- Little-Endian Systems: Because the host memory representation matches the archive format, 7-Zip performs direct memory reads and writes for multi-byte values where aligned (or unaligned, on architectures supporting it). This eliminates conversion overhead and maximizes compression and extraction throughput.
- Big-Endian Systems: When compiled for big-endian
processors, 7-Zip cannot cast raw byte buffers directly to native
multi-byte integer types. Instead, it routes reads and writes through
inline conversion functions. These functions reconstruct values byte by
byte or utilize compiler intrinsics (such as byte-swapping instructions
like
bswap) to reverse the byte order before processing or writing to disk.
Algorithmic Stream Independence
The core compression engines in 7-Zip, such as LZMA and LZMA2, operate primarily on bitstreams and byte-level arithmetic rather than CPU word-sized integers. The range coder and probability models read and write serialized data as individual 8-bit bytes. As a result, the compression and decompression algorithms remain inherently architecture-neutral, requiring endian-specific handling only for dictionary pointers, match finders, and header framing.