How does Memcached implement thread safety and lock granularity?

Memcached implements thread safety and fine-grained locking across multi-core processors by combining a master-worker thread architecture with granular mutexes that isolate critical sections across its hash table, memory slabs, and LRU eviction queues.

Thread Architecture and Event Dispatching

Memcached uses a multi-threaded, event-driven architecture powered by libevent. A single primary thread, known as the dispatcher thread, accepts incoming client connections and assigns them to worker threads using a round-robin distribution model over internal pipes. Each worker thread runs its own event loop, handling network I/O, parsing client commands, and executing read or write operations against shared in-memory data structures.

Granular Lock Design

To prevent race conditions without serializing access across CPU cores, Memcached breaks down system locking into distinct functional domains rather than relying on a single global lock:

  • Hash Table Item Locks: Memcached uses a array of fine-grained mutexes (item locks) mapped to key hash values. When accessing or updating a specific key, a worker thread acquires only the mutex covering that hash bucket range, allowing concurrent threads to read and write different key ranges simultaneously.
  • Slab Class Locks: Memory allocation relies on a slab allocator, where memory is divided into pre-allocated chunk sizes grouped by slab classes. Memory allocation and chunk re-assignment are protected by locks scoped to individual slab classes, preventing threads from blocking each other when allocating memory for different item sizes.
  • Segmented LRU Locks: Memcached organizes cache eviction queues into separate sub-LRUs (such as HOT, WARM, and COLD lists) per slab class. Rather than updating the main LRU list on every read access—which previously caused high lock contention—items are bumped between temperature queues using asynchronous workers and per-queue locks, drastically reducing mutex contention during read-heavy workloads.
  • Hash Table Expansion Lock: When the internal hash table grows, Memcached migrates items to a larger table using a dedicated maintenance thread. Worker threads briefly acquire a lock on individual buckets during migration, ensuring ongoing operations continue with minimal delay.

Through this combination of event-driven worker threads and split mutex domain partitioning, Memcached minimizes thread contention and achieves high throughput on modern multi-core systems.