Extract Files by Checksum Using Unrar
Extracting files from a RAR archive based on a list of checksums
requires a multi-step approach because the unrar utility
does not natively support filtering by checksums. However, RAR archives
store the CRC32 (or BLAKE2sp) checksum for every archived file in their
metadata. By generating a technical listing of the archive with
unrar, matching those internal checksums against your
target list, and passing the corresponding file names back to
unrar, you can extract only the exact files you need.
Step 1: View Archive Checksums
To see the checksums stored inside a RAR archive, use the technical
list command lt:
unrar lt archive.rarThis output displays metadata blocks for each file, including fields
for Name and Checksum (typically an
8-character hexadecimal CRC32 string).
Step 2: Prepare the Checksum List
Create a plain text file named checksums.txt containing
the checksums you want to extract, with one checksum per line:
A1B2C3D4
E5F67890
12345678
Ensure your list uses uppercase characters to match standard
unrar output.
Step 3: Extract Matching Files Using a Bash Script
Because unrar cannot filter by checksum directly, use
the following Bash script to parse the archive metadata, match the
checksums, and extract the matching files:
#!/usr/bin/env bash
ARCHIVE="archive.rar"
CHECKSUM_FILE="checksums.txt"
# Parse "Name" and "Checksum" fields from unrar technical output
awk -F': ' '
/^Name: / { name = $2 }
/^Checksum: / {
crc = $2
print crc, name
}
' <(unrar lt "$ARCHIVE") | while read -r crc filename; do
# Check if the CRC exists in the target checksums file
if grep -qi "^$crc$" "$CHECKSUM_FILE"; then
echo "Match found: $filename ($crc) -> Extracting..."
unrar x -y "$ARCHIVE" "$filename"
fi
doneAlternative: Automated Python Script
If you are working across platforms (Linux, macOS, or Windows) or dealing with complex file paths, Python provides a more robust parser:
import subprocess
import re
archive = "archive.rar"
checksums_file = "checksums.txt"
# Read target checksums
with open(checksums_file, "r") as f:
target_checksums = {line.strip().upper() for line in f if line.strip()}
# Run unrar lt to get metadata
output = subprocess.check_output(["unrar", "lt", archive], text=True)
# Parse output for Name and Checksum pairs
matches = []
current_name = None
for line in output.splitlines():
if line.startswith("Name: "):
current_name = line.replace("Name: ", "").strip()
elif line.startswith("Checksum: "):
checksum = line.replace("Checksum: ", "").strip().upper()
if checksum in target_checksums and current_name:
matches.append(current_name)
current_name = None
# Extract the matched files
if matches:
print(f"Extracting {len(matches)} matching files...")
subprocess.run(["unrar", "x", "-y", archive] + matches)
else:
print("No matching files found in the archive.")Important Considerations
- Checksum Format: Standard RAR4 archives use CRC32 (8 hexadecimal characters). RAR5 archives may use BLAKE2sp (64 hexadecimal characters) depending on how the archive was created. Ensure the entries in your target checksum file match the algorithm used by the archive.
- Directory Structure: Using
unrar xpreserves the folder structure upon extraction. If you want all extracted files placed directly into the current directory without subfolders, replaceunrar xwithunrar e.