How 7-Zip Processes DOCX and XLSX as ZIP Files
Modern Microsoft Office documents, such as DOCX and XLSX files, are packaged using the Open Packaging Conventions (OPC) standard, which inherently relies on the standard ZIP archive format. This article explains how the 7-Zip compression utility identifies, reads, and unpacks these Office documents by scanning standard ZIP file signatures and directory records, enabling users to explore and extract the underlying XML files and media assets.
The Office Open XML Architecture
Files with .docx and .xlsx extensions are
part of the Office Open XML (OOXML) specification. Rather than storing
information in a single binary blob, these documents are directory
structures containing:
- XML Schemas: Text documents like
document.xmlor spreadsheets likeworkbook.xmlcontaining the document data. - Relationship Files (
.rels): Pointers that define the associations between data components. - Content Types (
[Content_Types].xml): MIME-type mappings for parts within the package. - Media Assets: Raw image, audio, or video files
embedded in folders like
word/mediaorxl/media.
To make distribution efficient, this entire structure is compressed into a standard ZIP archive.
File Signature Detection
File extensions are primarily used by operating systems to associate files with specific applications. When opening a file, 7-Zip does not rely solely on the extension; it reads the file's binary header (the "magic bytes").
A valid ZIP file begins with the hexadecimal sequence
50 4B 03 04 (ASCII characters PK followed by
standard header indicators). When 7-Zip encounters a .docx
or .xlsx file, it immediately identifies this header.
Recognizing the signature, 7-Zip treats the file as an ordinary ZIP
container regardless of whether Microsoft Word or Excel is installed on
the system.
Parsing the Central Directory
Instead of decompressing the entire file into memory to read its contents, 7-Zip processes the structure using the standard ZIP specification:
- Locating the End of Central Directory (EOCD): 7-Zip seeks the end of the file to find the EOCD record, which points to the start of the archive's central directory.
- Reading the Central Directory: 7-Zip reads the
central directory table to build a virtual representation of the file
tree (
_rels/,docProps/,word/,xl/, etc.) along with the metadata (file sizes, compression methods, and CRC-32 checksums) for each internal entry. - Displaying the Virtual File System: The 7-Zip GUI or command-line interface renders this catalog directly to the user as folders and files.
Decompressing and Extracting Streams
Most contents within a DOCX or XLSX file are compressed using the Deflate algorithm (compression method 8 in the ZIP specification). When a user extracts an XML file or an embedded image, 7-Zip looks up the exact byte offset defined in the central directory, locates the local file header, and streams the data through its built-in Deflate decompression engine.
Because 7-Zip treats the file strictly as a ZIP archive, it is agnostic to the semantic validity of the XML content inside. This allows users to recover uncompressed images, repair corrupted XML structures, or extract raw text directly from damaged documents that Microsoft Office may fail to open.