How 7-Zip Handles Apple HFS and HFS+ Disk Images

7-Zip provides native, read-only support for Apple’s Hierarchical File System (HFS) and HFS+ (Mac OS Extended) disk images, allowing users on non-macOS platforms to inspect and unpack Mac-formatted archives. This article explains how the utility unwraps Apple disk image containers, traverses internal B-tree metadata structures, handles dual-fork architectures, and navigates file system limitations across different operating systems.

Disk Image Container Unwrapping

Before interacting with an HFS or HFS+ filesystem, 7-Zip inspects the underlying container format. Apple disk images typically use the Universal Disk Image Format (UDIF), commonly distributed with a .dmg file extension, or raw sector copies (.img, .iso). For UDIF images, 7-Zip reads the embedded trailer (the koly block) located in the final 512 bytes of the file. This block contains metadata pointing to partition maps and compressed data chunks. 7-Zip decompresses standard chunk algorithms—such as zlib, bzip2, and raw data runs—to present the underlying virtual block device containing the HFS/HFS+ partition.

Volume Header Parsing

Once the raw partition boundaries are identified, 7-Zip accesses the filesystem metadata:

  • Offset Verification: In HFS+, the Volume Header is located at byte offset 1024 from the start of the partition.
  • Header Validation: 7-Zip checks the signature bytes (H+ for standard HFS+, HX for case-sensitive HFSX) and verifies key parameters, including allocation block size, total block count, and the locations of special system files.
  • Special System Files: 7-Zip reads the extents records for the Catalog File, the Allocation File (bitmap), and the Extents Overflow File, which are necessary to map files that are fragmented or distributed across the volume.

B-Tree Catalog Traversal

HFS and HFS+ organize directory hierarchies and file metadata using balanced B-trees. 7-Zip’s parsing engine reads the Catalog File’s B-tree to reconstruct directory trees:

  1. Node Processing: The utility reads the node descriptors, beginning at the root node and traversing downward through index nodes to leaf nodes.
  2. Key Evaluation: Leaf nodes contain keyed records representing folder threads, directory entries, and file records. 7-Zip parses these keys, which include parent folder IDs and Unicode (UTF-16) file names.
  3. Data Mapping: File records provide the exact block locations and extents (start block and block count) for the file's contents, enabling 7-Zip to extract file payloads directly without scanning the entire disk sequentially.

Handling Data and Resource Forks

Classic Apple filesystems utilize dual-fork architecture, separating content into a "data fork" and a "resource fork":

  • Data Fork: Contains the primary payload of the file. 7-Zip extracts this fork as the standard file output.
  • Resource Fork: Contains legacy macOS metadata, icons, or interface elements. Because modern host operating systems like Windows do not natively use Apple resource forks, 7-Zip generally ignores them or may extract them as separate companion streams, depending on the image format and extraction settings, preventing data collisions on standard filesystems.

Extraction Constraints and Platform Considerations

7-Zip operates strictly as a parser and unpacker for Apple filesystems; it cannot create, update, or write to HFS or HFS+ volumes. When extracting files onto Windows NTFS or FAT systems, 7-Zip sanitizes characters that are valid in HFS+ but reserved in Windows (such as :, \, /, and *). Additionally, while HFS+ is case-preserving and HFSX is case-sensitive, extracting to case-insensitive filesystems may cause file overwrites if two files share the same name with different casing.