How UnRAR Added Support for Unicode Filenames
Early versions of the UnRAR extraction utility relied on local code pages, causing filenames with non-Latin characters to become corrupted when extracted across different operating systems or language locales. To solve this, developers introduced a multi-stage overhaul of the RAR archive format and the UnRAR codebase: first implementing a backward-compatible compressed UTF-16 storage mechanism in the RAR 3.0 specification, later switching to native UTF-8 strings in RAR 5.0, and refactoring internal path processing to use wide-character system APIs across all target platforms.
The RAR 2.x Limitation
In RAR 2.x and earlier formats, filenames were saved strictly as single-byte character strings using the host machine's active ANSI or OEM code page (such as Windows-1252 or CP437). UnRAR read these raw bytes directly and passed them to the operating system's standard file creation routines. If an archive created on a Japanese system was extracted on a machine configured for English or Cyrillic, the missing codepage mapping resulted in unreadable characters (mojibake) or invalid path errors.
The RAR 3.0 Hybrid UTF-16 System
To introduce Unicode support without breaking compatibility with older extractors, the RAR 3.0 specification altered the main file header structure:
- Header Flag Activation: A specific flag
(
LHD_UNICODE, value0x0200) was introduced into the file header's flags field. If this flag was unset, UnRAR treated the filename as legacy single-byte text. - Dual-String Storage: If the Unicode flag was set, the archive stored two versions of the name within the same byte block: a standard null-terminated 8-bit ASCII-compatible fallback name, followed immediately by an encoded Unicode data stream.
- Delta and High-Byte Compression: Storing raw 16-bit
UTF-16 characters would have doubled the header overhead for filenames.
UnRAR implemented a specialized decoding function (commonly implemented
as
ExtrNameToWide) that parsed a stateful byte stream:- A control byte dictated whether following characters inherited the current high byte or switched to a new base page.
- Single-byte characters matching the ASCII fallback were referenced directly to reduce payload size.
- Multi-byte characters were reassembled into full 16-bit integers
(
wchar_t/ UTF-16LE values).
This allowed older versions of UnRAR to safely read the first ASCII string and ignore the trailing payload, while updated versions detected the flag, ran the decompression routine, and recovered the full Unicode representation.
The RAR 5.0 Overhaul to UTF-8
With the introduction of the RAR 5.0 format, the backward-compatibility layer with legacy 1990s formats was deprecated in favor of a cleaner architecture:
- Direct UTF-8 Storage: RAR 5.0 abandoned both single-byte code page fallbacks and the proprietary compressed UTF-16 stream. All file paths, directory entries, and archive comments are now stored natively as variable-length UTF-8 byte sequences.
- Header Structure Simplification: Filenames are stored with a length-prefixed UTF-8 byte array. UnRAR reads the string directly, eliminating the need for stateful decompression passes over header text.
Internal UnRAR Source Code Changes
Accommodating Unicode required deep architectural changes within the UnRAR C++ source code:
- Wide Character Buffers: Internal path management
shifted from simple
chararrays towchar_tarrays and custom wide-string wrappers. - Platform-Specific Conversion Layers: On Windows
systems, UnRAR was rewritten to bypass ANSI Win32 APIs (such as
CreateFileAandFindFirstFileA) in favor of Unicode-native APIs (CreateFileW,FindFirstFileW, and_wstat). On POSIX platforms (Linux and macOS), UnRAR converts UTF-16 inputs into UTF-8 strings before interacting with system calls, as modern POSIX filesystems treat filenames as byte sequences typically interpreted as UTF-8. - Normalization and Invalid Sequence Handling: UnRAR
added validation checks to identify and sanitize illegal characters,
mismatched surrogate pairs, and path traversal tokens (
../) across different character encodings before files are written to disk.