How Unrar Handles Invalid UTF-8 Filenames

When extracting archives containing filenames with invalid or malformed UTF-8 sequences, the command-line utility unrar utilizes internal character conversion routines that attempt to sanitize, replace, or adapt the problematic characters to fit the host operating system's filesystem rules. Depending on the archive format version (RAR4 versus RAR5) and the system locale, unrar either replaces unmappable bytes with generic substitution characters, transliterates them, or outputs the raw byte stream directly to the filesystem.

RAR Header Encodings: RAR4 vs. RAR5

To understand how unrar deals with invalid UTF-8, it is necessary to differentiate how filename metadata is stored:

  • RAR4 (Legacy): Filenames were primarily stored using a two-byte or single-byte character array mapped to local codepages (such as Windows CP-1252 or OEM codepages), alongside an optional UTF-16LE field for Unicode support. When parsing RAR4 archives, unrar attempts to read the UTF-16 sequence first. If it is corrupted or missing, it falls back to the single-byte raw representation.
  • RAR5 (Modern): Filenames are strictly defined to be encoded in UTF-8 directly inside the header data.

Character Conversion and Replacement

Internally, official RARLAB unrar processes paths by converting multibyte streams into wide-character strings (wchar_t) before writing to the target filesystem.

When encountering an invalid byte sequence that violates UTF-8 syntax—such as unexpected continuation bytes, overlong encodings, or invalid start bytes—the conversion routine fails to map the sequence to a valid Unicode code point. Instead of terminating the entire extraction process, unrar handles the error through the following mechanisms:

  1. Substitution with Fallback Characters: If the conversion routine detects an illegal byte sequence, it often replaces the offending byte with a standard placeholder (such as an underscore _ or a question mark ?, depending on the platform) to construct a valid string.
  2. Locale-Based Multibyte Conversion: On POSIX platforms (Linux and macOS), unrar relies on standard C library functions such as mbstowcs and wcstombs, guided by the active system locale (LC_CTYPE or LANG). If a filename contains bytes that cannot be represented in the current locale, these functions return an error, prompting unrar to use a sanitized fallback name or emit a warning indicating that the path could not be mapped cleanly.
  3. Hexadecimal Escaping: In certain builds and fork implementations (such as unrar-free or tools built on libarchive), bytes that do not conform to UTF-8 are converted to hexadecimal escape sequences (e.g., #XX) to preserve the uniqueness of the file while ensuring the final path remains valid UTF-8.

Filesystem-Level Considerations

The underlying operating system plays a major role in how unrar completes extraction:

  • Linux and POSIX Systems: The Linux kernel treats filenames as arbitrary, null-terminated byte sequences without enforcing strict UTF-8 validity (apart from disallowing ASCII NUL and /). If unrar bypasses wide-character conversion, raw invalid UTF-8 bytes may be written directly to the disk, resulting in filenames that display as broken characters or replacement symbols (such as ``) in standard terminal emulators and file managers.
  • Windows: Because the Windows kernel and NTFS enforce UTF-16LE for file naming, any invalid UTF-8 sequence must be resolved before calling the Win32 API (CreateFileW). Any character that cannot be converted into a valid UTF-16 code unit must either be replaced or the file creation call will fail with an invalid path error.

Resolving Extraction Failures

If unrar encounters an archive where filename corruption prevents extraction entirely, the issue can often be mitigated by adjusting the runtime environment:

  • Setting LC_ALL=C forces POSIX systems to treat paths as raw byte streams, avoiding UTF-8 validation errors during extraction.
  • Using specialized tools like 7z or bsdtar provides alternative fallback handling for malformed codepages, allowing manual specification of archive character encodings when unrar's default sanitization fails.