Extract Files by Checksum Using Unrar

Extracting files from a RAR archive based on a list of checksums requires a multi-step approach because the unrar utility does not natively support filtering by checksums. However, RAR archives store the CRC32 (or BLAKE2sp) checksum for every archived file in their metadata. By generating a technical listing of the archive with unrar, matching those internal checksums against your target list, and passing the corresponding file names back to unrar, you can extract only the exact files you need.

Step 1: View Archive Checksums

To see the checksums stored inside a RAR archive, use the technical list command lt:

unrar lt archive.rar

This output displays metadata blocks for each file, including fields for Name and Checksum (typically an 8-character hexadecimal CRC32 string).

Step 2: Prepare the Checksum List

Create a plain text file named checksums.txt containing the checksums you want to extract, with one checksum per line:

A1B2C3D4
E5F67890
12345678

Ensure your list uses uppercase characters to match standard unrar output.

Step 3: Extract Matching Files Using a Bash Script

Because unrar cannot filter by checksum directly, use the following Bash script to parse the archive metadata, match the checksums, and extract the matching files:

#!/usr/bin/env bash

ARCHIVE="archive.rar"
CHECKSUM_FILE="checksums.txt"

# Parse "Name" and "Checksum" fields from unrar technical output
awk -F': ' '
    /^Name: / { name = $2 }
    /^Checksum: / {
        crc = $2
        print crc, name
    }
' <(unrar lt "$ARCHIVE") | while read -r crc filename; do
    # Check if the CRC exists in the target checksums file
    if grep -qi "^$crc$" "$CHECKSUM_FILE"; then
        echo "Match found: $filename ($crc) -> Extracting..."
        unrar x -y "$ARCHIVE" "$filename"
    fi
done

Alternative: Automated Python Script

If you are working across platforms (Linux, macOS, or Windows) or dealing with complex file paths, Python provides a more robust parser:

import subprocess
import re

archive = "archive.rar"
checksums_file = "checksums.txt"

# Read target checksums
with open(checksums_file, "r") as f:
    target_checksums = {line.strip().upper() for line in f if line.strip()}

# Run unrar lt to get metadata
output = subprocess.check_output(["unrar", "lt", archive], text=True)

# Parse output for Name and Checksum pairs
matches = []
current_name = None

for line in output.splitlines():
    if line.startswith("Name: "):
        current_name = line.replace("Name: ", "").strip()
    elif line.startswith("Checksum: "):
        checksum = line.replace("Checksum: ", "").strip().upper()
        if checksum in target_checksums and current_name:
            matches.append(current_name)
        current_name = None

# Extract the matched files
if matches:
    print(f"Extracting {len(matches)} matching files...")
    subprocess.run(["unrar", "x", "-y", archive] + matches)
else:
    print("No matching files found in the archive.")

Important Considerations

  • Checksum Format: Standard RAR4 archives use CRC32 (8 hexadecimal characters). RAR5 archives may use BLAKE2sp (64 hexadecimal characters) depending on how the archive was created. Ensure the entries in your target checksum file match the algorithm used by the archive.
  • Directory Structure: Using unrar x preserves the folder structure upon extraction. If you want all extracted files placed directly into the current directory without subfolders, replace unrar x with unrar e.