Configure unrar Character Set for Filenames

When extracting RAR archives created on foreign operating systems or legacy platforms, filenames can frequently become corrupted or unreadable due to character encoding mismatches. This article outlines how to configure unrar to recognize specific character sets using built-in switches, how to manipulate system locale environments to force correct decoding, and how to utilize purpose-built alternatives when standard unrar tools fail to interpret legacy code pages.

Using the Built-In -sc Switch

The official command-line unrar utility provided by RARLAB includes the -sc (Specify Character set) switch. This parameter defines the character encoding for handling text streams such as comments, file lists, and redirected output.

The syntax for the switch is:

unrar x -sc<charset> archive.rar

The <charset> parameter accepts the following values:

  • u: Unicode (UTF-8)
  • a: ANSI (standard Windows code page)
  • o: OEM (DOS code page)

For example, to extract an archive while enforcing UTF-8 character interpretation:

unrar x -scu archive.rar

Forcing Locale for Legacy Code Pages

Older RAR formats (such as RAR 2.0 and RAR 3.0) do not store filenames natively in Unicode; instead, they rely on the creator's local system code page. Because standard unrar relies on system libraries to map these characters, you can force unrar to use a specific code page by setting the system's LC_ALL or LANG environment variables before running the extraction command.

To temporarily change the environment for a single extraction command:

For Japanese (Shift-JIS / CP932):

LC_ALL=ja_JP.SJIS unrar x archive.rar

For Simplified Chinese (GBK / CP936):

LC_ALL=zh_CN.GBK unrar x archive.rar

For Cyrillic (CP1251):

LC_ALL=ru_RU.CP1251 unrar x archive.rar

Ensure the target locale is generated on your system. On Debian and Ubuntu systems, you can generate missing locales by running sudo dpkg-reconfigure locales or adding the required locale to /etc/locale.gen and executing sudo locale-gen.

Alternative Solution: Using unar for Exact Encodings

The standard unrar utility does not provide an explicit parameter to define arbitrary code pages (such as Big5, EUC-KR, or CP936) by name. If setting the locale fails, the command-line utility unar is the standard solution for handling non-Unicode archives.

Install unar via your package manager:

# Debian/Ubuntu
sudo apt install unar

# Arch Linux
sudo pacman -S unar

# macOS (Homebrew)
brew install unar

View the encodings supported by the tool:

unar -e help

Extract the archive by passing the exact encoding identifier directly with the -e flag:

unar -e gbk archive.rar
unar -e shift-jis archive.rar
unar -e cp949 archive.rar

Renaming Already Extracted Files

If files have already been extracted with corrupted names (mojibake), you can repair the filenames without re-extracting using the convmv utility.

Install convmv and convert filenames from the source encoding to UTF-8:

convmv -f gbk -t utf-8 -r --notest /path/to/extracted/folder

Remove --notest to run a dry run first, which displays what the filenames will look like before applying changes.