How UnRAR Added Support for Unicode Filenames

Early versions of the UnRAR extraction utility relied on local code pages, causing filenames with non-Latin characters to become corrupted when extracted across different operating systems or language locales. To solve this, developers introduced a multi-stage overhaul of the RAR archive format and the UnRAR codebase: first implementing a backward-compatible compressed UTF-16 storage mechanism in the RAR 3.0 specification, later switching to native UTF-8 strings in RAR 5.0, and refactoring internal path processing to use wide-character system APIs across all target platforms.

The RAR 2.x Limitation

In RAR 2.x and earlier formats, filenames were saved strictly as single-byte character strings using the host machine's active ANSI or OEM code page (such as Windows-1252 or CP437). UnRAR read these raw bytes directly and passed them to the operating system's standard file creation routines. If an archive created on a Japanese system was extracted on a machine configured for English or Cyrillic, the missing codepage mapping resulted in unreadable characters (mojibake) or invalid path errors.

The RAR 3.0 Hybrid UTF-16 System

To introduce Unicode support without breaking compatibility with older extractors, the RAR 3.0 specification altered the main file header structure:

  1. Header Flag Activation: A specific flag (LHD_UNICODE, value 0x0200) was introduced into the file header's flags field. If this flag was unset, UnRAR treated the filename as legacy single-byte text.
  2. Dual-String Storage: If the Unicode flag was set, the archive stored two versions of the name within the same byte block: a standard null-terminated 8-bit ASCII-compatible fallback name, followed immediately by an encoded Unicode data stream.
  3. Delta and High-Byte Compression: Storing raw 16-bit UTF-16 characters would have doubled the header overhead for filenames. UnRAR implemented a specialized decoding function (commonly implemented as ExtrNameToWide) that parsed a stateful byte stream:
    • A control byte dictated whether following characters inherited the current high byte or switched to a new base page.
    • Single-byte characters matching the ASCII fallback were referenced directly to reduce payload size.
    • Multi-byte characters were reassembled into full 16-bit integers (wchar_t / UTF-16LE values).

This allowed older versions of UnRAR to safely read the first ASCII string and ignore the trailing payload, while updated versions detected the flag, ran the decompression routine, and recovered the full Unicode representation.

The RAR 5.0 Overhaul to UTF-8

With the introduction of the RAR 5.0 format, the backward-compatibility layer with legacy 1990s formats was deprecated in favor of a cleaner architecture:

  • Direct UTF-8 Storage: RAR 5.0 abandoned both single-byte code page fallbacks and the proprietary compressed UTF-16 stream. All file paths, directory entries, and archive comments are now stored natively as variable-length UTF-8 byte sequences.
  • Header Structure Simplification: Filenames are stored with a length-prefixed UTF-8 byte array. UnRAR reads the string directly, eliminating the need for stateful decompression passes over header text.

Internal UnRAR Source Code Changes

Accommodating Unicode required deep architectural changes within the UnRAR C++ source code:

  • Wide Character Buffers: Internal path management shifted from simple char arrays to wchar_t arrays and custom wide-string wrappers.
  • Platform-Specific Conversion Layers: On Windows systems, UnRAR was rewritten to bypass ANSI Win32 APIs (such as CreateFileA and FindFirstFileA) in favor of Unicode-native APIs (CreateFileW, FindFirstFileW, and _wstat). On POSIX platforms (Linux and macOS), UnRAR converts UTF-16 inputs into UTF-8 strings before interacting with system calls, as modern POSIX filesystems treat filenames as byte sequences typically interpreted as UTF-8.
  • Normalization and Invalid Sequence Handling: UnRAR added validation checks to identify and sanitize illegal characters, mismatched surrogate pairs, and path traversal tokens (../) across different character encodings before files are written to disk.