How 7-Zip Handles Unicode Filenames

7-Zip manages Unicode filenames by leveraging modern character encoding standards to ensure that international characters, symbols, and non-Latin scripts are accurately preserved across different operating systems. Depending on the archive format used—primarily comparing its native .7z format against the legacy .zip specification—7-Zip implements specific encoding methods such as UTF-16 and UTF-8 to prevent filename corruption, character substitution, and cross-platform compatibility issues during compression and extraction.

Native .7z Format Implementation

In its native .7z format, 7-Zip handles filenames using UTF-16 Little Endian encoding by default. Because UTF-16 is integrated directly into the core design of the .7z format:

  • Universal Character Support: Characters from any language (including Asian scripts, Cyrillic, Arabic, and emojis) are stored accurately without relying on the host operating system's active code page.
  • Locale Independence: An archive created on a system with a Japanese locale will extract with identical filenames on a system running an English, German, or Hebrew locale without producing "mojibake" (garbled text).

Handling Standard ZIP Archives

The original ZIP specification predated widespread Unicode adoption and relied on the host system's default OEM/ANSI code page. To resolve this, 7-Zip implements Unicode support in .zip archives through specific mechanisms:

  1. UTF-8 Flagging: When compressing to .zip, 7-Zip sets the UTF-8 flag (Bit 11 of the general-purpose bit flag) if the filename contains characters outside the standard ASCII range. This instructs compliant unpackers to interpret the filename bytes as UTF-8.
  2. Info-ZIP Unicode Path Extra Field: For backward compatibility with older unpackers, 7-Zip can read and write standard Info-ZIP Unicode Path extra fields (0x7075), ensuring alternative utilities can properly resolve the filenames.

Extraction and Legacy Archive Fallbacks

When extracting archives that lack explicit Unicode flags (common in older ZIP files created by legacy software):

  • Auto-Detection: 7-Zip attempts to detect the character encoding. If no UTF-8 flag or Unicode metadata is present, it defaults to the local operating system's current OEM/ANSI code page.
  • Command-Line Override: Advanced users can force specific code page interpretations using the command-line interface. By using the -mcu switch (for example, -mcu=on to force UTF-8, or -mcu=932 to interpret legacy Shift-JIS filenames), users can resolve filename display issues from legacy sources.

Cross-Platform Path Handling

Operating systems handle path separators and forbidden characters differently (such as Windows disallowing characters like :, *, ?, or <). When 7-Zip extracts an archive containing Unicode filenames:

  • Filename Sanitization: On Windows, 7-Zip automatically converts or replaces illegal path characters to prevent extraction errors, while retaining valid Unicode code points.
  • Normalization: 7-Zip generally preserves the Unicode normalization form (such as NFC or NFD) provided in the archive, ensuring correct visual display across Windows, Linux, and macOS platforms.