How 7-Zip Handles Unicode Filenames
7-Zip manages Unicode filenames by leveraging modern character
encoding standards to ensure that international characters, symbols, and
non-Latin scripts are accurately preserved across different operating
systems. Depending on the archive format used—primarily comparing its
native .7z format against the legacy .zip
specification—7-Zip implements specific encoding methods such as UTF-16
and UTF-8 to prevent filename corruption, character substitution, and
cross-platform compatibility issues during compression and
extraction.
Native .7z Format Implementation
In its native .7z format, 7-Zip handles filenames using
UTF-16 Little Endian encoding by default. Because UTF-16 is integrated
directly into the core design of the .7z format:
- Universal Character Support: Characters from any language (including Asian scripts, Cyrillic, Arabic, and emojis) are stored accurately without relying on the host operating system's active code page.
- Locale Independence: An archive created on a system with a Japanese locale will extract with identical filenames on a system running an English, German, or Hebrew locale without producing "mojibake" (garbled text).
Handling Standard ZIP Archives
The original ZIP specification predated widespread Unicode adoption
and relied on the host system's default OEM/ANSI code page. To resolve
this, 7-Zip implements Unicode support in .zip archives
through specific mechanisms:
- UTF-8 Flagging: When compressing to
.zip, 7-Zip sets the UTF-8 flag (Bit 11 of the general-purpose bit flag) if the filename contains characters outside the standard ASCII range. This instructs compliant unpackers to interpret the filename bytes as UTF-8. - Info-ZIP Unicode Path Extra Field: For backward
compatibility with older unpackers, 7-Zip can read and write standard
Info-ZIP Unicode Path extra fields (
0x7075), ensuring alternative utilities can properly resolve the filenames.
Extraction and Legacy Archive Fallbacks
When extracting archives that lack explicit Unicode flags (common in older ZIP files created by legacy software):
- Auto-Detection: 7-Zip attempts to detect the character encoding. If no UTF-8 flag or Unicode metadata is present, it defaults to the local operating system's current OEM/ANSI code page.
- Command-Line Override: Advanced users can force
specific code page interpretations using the command-line interface. By
using the
-mcuswitch (for example,-mcu=onto force UTF-8, or-mcu=932to interpret legacy Shift-JIS filenames), users can resolve filename display issues from legacy sources.
Cross-Platform Path Handling
Operating systems handle path separators and forbidden characters
differently (such as Windows disallowing characters like :,
*, ?, or <). When 7-Zip
extracts an archive containing Unicode filenames:
- Filename Sanitization: On Windows, 7-Zip automatically converts or replaces illegal path characters to prevent extraction errors, while retaining valid Unicode code points.
- Normalization: 7-Zip generally preserves the Unicode normalization form (such as NFC or NFD) provided in the archive, ensuring correct visual display across Windows, Linux, and macOS platforms.