How Unrar Extracts Files Larger Than RAM
When extracting an archive containing a file larger than the system's
available physical memory, the unrar utility does not load
the entire uncompressed file into RAM. Instead, it relies on
stream-based processing, a bounded sliding dictionary, and buffered disk
writes. This architecture ensures that memory consumption remains tied
strictly to the compression parameters rather than the actual size of
the extracted contents.
Stream-Based Processing
The RAR decompression algorithm operates on sequential streams of
data rather than monolithic datasets. The unrar program
reads small chunks of compressed data from the archive on the storage
drive, processes that data in memory, and immediately writes the
resulting decompressed stream back to the target storage device. Because
the data passes through the program continuously, the total file
size—whether 10 gigabytes or several terabytes—has no direct bearing on
the amount of RAM required for the operation.
The Sliding Dictionary Window
Memory usage during extraction is primarily determined by the "dictionary size" selected when the archive was created. RAR uses a Lempel-Ziv-based compression algorithm, which relies on a sliding history window to identify repeated byte sequences:
- RAR4 Archives: Typically use a dictionary size of up to 4 megabytes.
- RAR5 Archives: Support larger dictionary sizes ranging from 128 kilobytes up to 1 gigabyte (or up to 4 gigabytes in newer revisions), with 32 megabytes being the standard default.
During extraction, unrar allocates only enough memory to
hold this sliding dictionary window, along with internal decompression
tables and small I/O buffers. For example, extracting a 100-gigabyte
file created with a 32-megabyte dictionary requires only roughly 40 to
60 megabytes of working memory.
Direct Disk Flushing
As the sliding window updates, data that falls out of the history window is flushed to the destination file on the disk via standard operating system write buffers. Once written to disk or the OS write cache, that memory is recycled for subsequent chunks of the stream.
If the system's RAM is smaller than the archive's required dictionary
size—such as attempting to extract an archive using a 1-gigabyte
dictionary on a legacy system with only 512 megabytes of RAM—the process
will rely on the operating system's swap space or pagefile. In cases
where virtual memory is also exhausted, unrar will
terminate with an out-of-memory error rather than attempting extraction.
However, under normal conditions with standard dictionary sizes,
unrar routinely extracts files many times larger than the
host system's RAM without performance degradation.