How unrar Handles Unicode Filenames on macOS

This article provides an overview of how the unrar command-line utility handles Unicode normalization when extracting archives on macOS. When extracting files, unrar does not internally alter or re-normalize Unicode text; instead, it outputs standard UTF-8 strings directly to POSIX file system APIs, relying on the macOS file system layer (APFS or HFS+) to determine whether file names are preserved as precomposed (NFC) or decomposed (NFD) characters.

Unicode Representation in RAR Archives

Modern RAR archives (RAR 5.0 and later) store file names natively as UTF-8 byte sequences. Older archives (RAR 4.x and earlier) store file names in UTF-16 alongside local codepages, which unrar converts to UTF-8 during extraction. In most operating environments, particularly Windows and Linux, Unicode file names are saved using Normalization Form C (NFC), where accented characters are represented as single precomposed code points (for example, é as U+00E9).

The Behavior of unrar During Extraction

The portable Unix source code of unrar (provided by RARLAB) treats file paths on Unix-like operating systems as opaque null-terminated UTF-8 byte streams. When executing on macOS:

  1. No Application-Level Normalization: unrar does not invoke macOS-specific APIs (such as CFStringNormalize) to alter Unicode normalization forms. It preserves the exact UTF-8 character sequences embedded in the archive headers.
  2. Direct POSIX Calls: Filenames are passed directly to standard system calls such as open(), creat(), and mkdir().

Because unrar does not convert the encoding between NFC and NFD itself, the resulting normalization form on disk depends entirely on the file system hosting the destination directory.

macOS File System Normalization: APFS vs. HFS+

How the file name is actually written to the storage drive depends on the macOS version and the disk format:

  • APFS (Apple File System): APFS is normalization-preserving and normalization-insensitive. When unrar creates a file using an NFC-encoded path (common in Windows-generated archives), APFS writes the exact NFC byte sequence to disk. Lookups remain normalization-insensitive, meaning the operating system treats both NFC and NFD representations of the name as the same file.
  • HFS+ (Mac OS Extended): HFS+ enforces a proprietary variant of Normalization Form D (NFD). If unrar extracts a file to an HFS+ volume, the kernel's Virtual File System (VFS) layer intercepts the file creation call and automatically decomposes precomposed characters into multiple code points (for example, translating é into e + combining acute accent U+0065 U+0301).

Implications for macOS Users and Developers

Because modern macOS systems use APFS, unrar will typically leave extracted file names in the normalization format used by the archive's creator—usually NFC. This can cause subtle issues in workflows that expect native macOS NFD formatting:

  • Shell and Scripting Mismatches: Tools or scripts comparing file names via exact byte matching may fail if comparing an unrar-extracted NFC file name with an NFD file name generated natively by a standard macOS GUI application.
  • Archive Collisions: If an archive contains two distinct entries that differ only by Unicode normalization (one in NFC and one in NFD), extracting on macOS will cause the second file to overwrite the first due to the operating system's normalization-insensitive lookup rules.