Configure unrar Character Set for Filenames
When extracting RAR archives created on foreign operating systems or
legacy platforms, filenames can frequently become corrupted or
unreadable due to character encoding mismatches. This article outlines
how to configure unrar to recognize specific character sets
using built-in switches, how to manipulate system locale environments to
force correct decoding, and how to utilize purpose-built alternatives
when standard unrar tools fail to interpret legacy code
pages.
Using the Built-In
-sc Switch
The official command-line unrar utility provided by
RARLAB includes the -sc (Specify Character set) switch.
This parameter defines the character encoding for handling text streams
such as comments, file lists, and redirected output.
The syntax for the switch is:
unrar x -sc<charset> archive.rarThe <charset> parameter accepts the following
values:
u: Unicode (UTF-8)a: ANSI (standard Windows code page)o: OEM (DOS code page)
For example, to extract an archive while enforcing UTF-8 character interpretation:
unrar x -scu archive.rarForcing Locale for Legacy Code Pages
Older RAR formats (such as RAR 2.0 and RAR 3.0) do not store
filenames natively in Unicode; instead, they rely on the creator's local
system code page. Because standard unrar relies on system
libraries to map these characters, you can force unrar to
use a specific code page by setting the system's LC_ALL or
LANG environment variables before running the extraction
command.
To temporarily change the environment for a single extraction command:
For Japanese (Shift-JIS / CP932):
LC_ALL=ja_JP.SJIS unrar x archive.rarFor Simplified Chinese (GBK / CP936):
LC_ALL=zh_CN.GBK unrar x archive.rarFor Cyrillic (CP1251):
LC_ALL=ru_RU.CP1251 unrar x archive.rarEnsure the target locale is generated on your system. On Debian and
Ubuntu systems, you can generate missing locales by running
sudo dpkg-reconfigure locales or adding the required locale
to /etc/locale.gen and executing
sudo locale-gen.
Alternative
Solution: Using unar for Exact Encodings
The standard unrar utility does not provide an explicit
parameter to define arbitrary code pages (such as Big5, EUC-KR, or
CP936) by name. If setting the locale fails, the command-line utility
unar is the standard solution for handling non-Unicode
archives.
Install unar via your package manager:
# Debian/Ubuntu
sudo apt install unar
# Arch Linux
sudo pacman -S unar
# macOS (Homebrew)
brew install unarView the encodings supported by the tool:
unar -e helpExtract the archive by passing the exact encoding identifier directly
with the -e flag:
unar -e gbk archive.rar
unar -e shift-jis archive.rar
unar -e cp949 archive.rarRenaming Already Extracted Files
If files have already been extracted with corrupted names (mojibake),
you can repair the filenames without re-extracting using the
convmv utility.
Install convmv and convert filenames from the source
encoding to UTF-8:
convmv -f gbk -t utf-8 -r --notest /path/to/extracted/folderRemove --notest to run a dry run first, which displays
what the filenames will look like before applying changes.