moodycamel::ConcurrentQueue: An industrial-strength lock-free queue for C++
Moodycamel’s ConcurrentQueue: A Comprehensive Guide to a High-Performance Lock-Free Queue for C++
Introduction
In the world of multi-threaded C++ programming, a robust, fast, and flexible queue is a rare and valuable asset. Moodycamel’s ConcurrentQueue delivers exactly that: an industrial-strength, lock-free queue designed for multi-producer, multi-consumer (MPMC) workloads in C++11 and beyond. It’s implemented as a single-header library, enabling easy drop-in use without complex build steps or external dependencies. The design emphasizes speed, scalability, and practical features such as bulk operations, optional preallocation, and a lightweight blocking variant. While not every behavior is identical to a traditional linearizable queue, the library is engineered to minimize contention, maximize throughput under heavy workloads, and remain well tested and portable across standard C++ platforms.
Key Features at a Glance
- Blazing-fast performance with benchmarks that showcase notable throughput under contention.
- Single-header implementation: drop it into a project and go.
- Fully thread-safe lock-free queue usable from any number of threads.
- C++11-centric: elements are moved rather than copied wherever possible.
- Templated design eliminates the need to manage raw pointers; memory management is handled for you.
- No artificial limitations on element types or maximum count.
- Flexible memory model: pre-allocate up-front or allocate on the fly as needed.
- Fully portable: no assembly, relying on standard C++11 primitives.
- Supports high-speed bulk operations (enqueue/dequeue) to reduce per-item overhead.
- Includes a low-overhead blocking wrapper (BlockingConcurrentQueue) for wait-based patterns.
- Strong exception safety guarantees: the queue itself does not throw; failures surface as return values.
Why you’d use a lock-free queue (and where this one fits)
- Lock-free approaches reduce contention and often improve throughput in producer/consumer scenarios, especially under heavy load.
- Moodycamel’s queue is designed to be fast and practical, offering more features and fewer restrictions than many alternatives: it supports bulk operations, per-producer tokens, and a flexible memory model.
- It’s not a silver bullet. The queue is not linearizable, not NUMA-aware, and not sequentially consistent. These trade-offs are deliberate: the design emphasizes performance and reasonable correctness for many real-world streaming and processing patterns. If strict total ordering across independent producers or strict global memory ordering matters for your use case, consider other approaches or combine this queue with your own synchronization.
Important limitations to be aware of
- Ordering across independent producers is not guaranteed. If two producers enqueue at the same time, their relative order in the queue after dequeue may be undefined unless you manage ordering externally.
- The queue is not NUMA-aware and relies on internal memory reuse; scalability on NUMA architectures may be limited.
- It is not sequentially consistent; to achieve certain effects, you may need explicit memory ordering or careful design around pumping patterns. The recommended approach is to follow the provided samples and documented use cases.
High-level design: how it achieves speed
- Internals rely on contiguous blocks rather than linked lists, improving cache locality and throughput.
- The queue is composed of multiple sub-queues, one per producer. When a consumer dequeues, it scans sub-queues to find a non-empty one.
- The design makes the user mostly oblivious to the internal distribution: the bulk of the work is managed behind the scenes.
- A notable consequence of this architecture: if two producers enqueue concurrently, there is no defined ordering between their enqueued items when later dequeued—this is a natural outcome of the parallelism and the absence of global synchronization.
Basic usage: getting started
- The core implementation lives in a single header, concurrentqueue.h. There is a blocking variant in blockingconcurrentqueue.h that depends on the core header and a lightweight semaphore.
- Example usage (simplified):
- Include the header: #include "concurrentqueue.h"
- Create a queue: moodycamel::ConcurrentQueue q;
- Enqueue and dequeue:
- q.enqueue(25);
- int item;
- bool found = q.try_dequeue(item);
- // found and item == 25 if true
- Key methods you’ll encounter:
- Constructor: ConcurrentQueue(size_t initialSizeEstimate)
- enqueue(T&& item): Enqueue one item, expanding storage as needed
- try_enqueue(T&& item): Enqueue only if memory is pre-allocated
- try_dequeue(T& item): Dequeue one item if available
- Token-based variants:
- ProducerToken and ConsumerToken offer per-thread optimization: enqueue(ptok, item) and try_dequeue(ctok, item)
- Tokens can speed up operations and help with bulk operations
- The core API also includes bulk operations (enqueuebulk, trydequeue_bulk) for higher performance when processing multiple items at once.
A quick glance at the basic API (conceptual)
- Constructors: ConcurrentQueue(size_t initialSizeEstimate)
- enqueue(item)
- enqueue(prod_token, item)
- enqueuebulk(itemfirst, count) or enqueuebulk(prodtoken, item_first, count)
- try_enqueue(item)
- tryenqueue(prodtoken, item)
- tryenqueuebulk(item_first, count)
- tryenqueuebulk(prodtoken, itemfirst, count)
- try_dequeue(item&)
- trydequeue(constoken, item&)
- trydequeuebulk(item_first, max)
- trydequeuebulk(constoken, itemfirst, max)
- trydequeuefromproducer(prodtoken, item&) for targeted producer
- size_approx() for a rough count
Blocking version: wait-based semantics
- BlockingConcurrentQueue adds waitdequeue and waitdequeuebulk, with timed variants via waitdequeue_timed and a std::chrono-based timeout.
- A caution: destroying the queue while a consumer is waiting can be problematic; coordinate shutdown to ensure no thread is blocked on a wait when destruction occurs.
- Example snippet (conceptual):
- Blocking queue q;
- Producer thread pushes items with enqueue
- Consumer thread uses wait_dequeue(item) to obtain items, sometimes with a timeout
Advanced features that boost performance and flexibility Tokens and per-thread storage
- The queue can leverage extra per-producer and per-consumer storage if available, through tokens.
- You create a ProducerToken for a producer thread and use it to enqueue, and a ConsumerToken for a consumer thread to dequeue.
- Example:
- moodycamel::ProducerToken ptok(q);
- q.enqueue(ptok, 17);
- moodycamel::ConsumerToken ctok(q);
- int item;
- q.try_dequeue(ctok, item);
- Tokens are not strictly tied to a thread; they’re associated with a producer or consumer to improve performance, and there is one recommended per-thread token usage pattern.
Bulk operations
- Bulk enqueue and dequeue allow transferring multiple items with the same call, dramatically reducing per-item overhead under heavy workloads.
- Example:
- int items[] = {1, 2, 3, 4, 5};
- q.enqueue_bulk(items, 5);
- int results[5];
- sizet count = q.trydequeue_bulk(results, 5);
- for (size_t i = 0; i < count; ++i) assert(results[i] == items[i]);
Preallocation and memory sizing
- try_enqueue is non-allocating: if there isn’t enough room, it returns false.
- Pre-allocation requires careful calculation because blocks are the granularity of memory, and slots cannot be reused until blocks are fully filled and emptied.
- Block size and internal layout influence how many blocks you must reserve to guarantee space for a desired occupancy level.
- Simple formulas are provided for explicit versus implicit producers to determine total pre-allocated memory, or you can rely on the constructor overloads that compute it for you.
- Important caveat: even with correct pre-allocation, contention can cause try_enqueue to fail; you must code retry or fallback logic if you rely on non-blocking behavior.
Exception safety
- The queue is designed to be exception-safe. It does not throw; failures manifest as false returns (e.g., in allocation failures).
- Move semantics are preferred for enqueue when available; bulk enqueues may copy elements to maintain a consistent rollback if exceptions occur during construction.
- If destructors or copy/move operations can throw, annotate them as noexcept where possible to minimize exception-checking overhead.
- If a type’s destructor can throw, be mindful: dequeued elements are destructed after dequeue, but exceptions can complicate recovery.
Traits: customizing the allocator and constants
- The queue supports a traits template argument to override constants and allocation/deallocation behavior.
- Example:
- struct MyTraits : public moodycamel::ConcurrentQueueDefaultTraits { static const sizet BLOCKSIZE = 256; };
- moodycamel::ConcurrentQueue q;
- This allows tuning block size, allocator hooks, and other internal parameters to fit your application's memory characteristics.
Dequeueing without constructing items
- A common requirement is to dequeue into an object that isn’t easy to default-construct or to avoid extra construction work.
- The library provides patterns (wrappers) to copy the memory from dequeued elements or defer construction until assignment.
- Examples outline manual memcpy-based wrappers or wrappers that use placement-new to construct on dequeue with careful lifecycle management.
Samples, tests, and benchmarks
- The project ships with a variety of samples and tests to illustrate usage and verify behavior.
- There are unit tests, a fuzz tester, and model-checking with Relacy and CDSChecker, across Linux and Windows environments.
- The project includes benchmark code to compare against other queues (e.g., boost::lockfree::queue, tbb::concurrent_queue) and to showcase the performance characteristics of bulk operations.
- Documentation and blog posts provide deeper context on performance characteristics and design decisions.
Relationship to other tooling and licensing
- The queue integrates with popular toolchains: it requires a reasonably modern compiler (VS2012+, g++ 4.8+).
- It is available via vcpkg, with ongoing maintenance by contributors. The port is kept up to date by Microsoft and the community contributors.
- Licensing is generous: a simplified BSD license and dual-licensing under the Boost Software License. The licensing applies to the code itself, with caveats about third-party code used for benchmarking or testing (e.g., Boost queue, TBB, Relacy, and other tools).
- Note about patents: lock-free programming can intersect with patents; while the code itself is designed to avoid infringement, the author notes this is a potential risk in theory and encourages users to review licensing and patent implications for their jurisdiction and use case.
Diving into the code: a high-level map
- The codebase is organized to facilitate a clean separation of concerns:
- Helpers: utility functions such as rounding to power-of-two and fast operations.
- Traits: default traits live here, providing the constants and allocator hooks.
- Tokens: producer and consumer tokens that optimize per-thread usage.
- Public API: constructors, destructors, swap/assignment, and the bulk of enqueue/dequeue methods.
- Inline dequeue logic: designed to be straightforward yet efficient.
- Data structures: main internal blocks, the free-list, and the block pool used to recycle memory.
- Produer and consumer’s internal engines: two specialized SPMC sub-queues (one for explicit producers, one for implicit producers) with the capacity to recycle and optimize memory.
- Internal helpers: memory management, initial block pools, and the mapping between indices and blocks.
- Metadata: the queue’s core data members, followed by non-member swap helpers.
- The project’s narrative (as described in the original blog posts) emphasizes the “bulk-first” philosophy: if you batch operations, you will realize the greatest gains.
Practical considerations and best practices
- Read the samples and follow the recommended usage patterns. They reflect how the library is intended to be used in real applications.
- If you need deterministic ordering between multiple producers, consider using a single producer token or a stricter external synchronization strategy.
- When using preallocation, carefully compute an appropriate initial size to avoid unexpected try_enqueue failures under peak load.
- For tiny, simple producer/consumer scenarios (e.g., one producer and one or more consumers), compare with the SPSC variant mentioned in related projects to ensure you’re picking the right tool for the job.
- If you’re building with a toolchain that is not C++11-compliant or lacks certain atomics, ensure compatibility or apply the recommended minimum compiler versions.
A closing note on design choices
- Moodycamel’s ConcurrentQueue is designed as a practical, high-performance MP queue for real-world workloads. It embraces the realities of multi-threaded programming: it prioritizes throughput and safety, even if that means relaxing some strict ordering and memory-ordering guarantees.
- The trade-offs are explicit and documented, enabling you to decide if the library fits your use case. For many streaming, log processing, or event-driven scenarios, its bulk operations and token-based optimizations offer tangible benefits with manageable complexity.
Licensing and where to learn more
- If you’re curious for deeper context, the original blog posts and code references discuss the design decisions, the internal architecture, and benchmark results:
- A fast general-purpose lock-free queue for C++: blog post and benchmarks
- Detailed design overview: blog post
- Full source and samples: the project repository
- For integration details, you can consult the LICENSE and the project’s repository pages, and the vcpkg instructions for shipping and distribution in your project.
In summary
Moodycamel’s ConcurrentQueue stands as a practical, battle-tested MP/multi-consumer queue for C++11, designed to maximize throughput with minimal fuss. Its feature set—bulk operations, per-thread tokens, preallocation strategies, a blocking variant, and strong exception safety—addresses the practical needs of modern concurrent software. While it sacrifices some strict ordering guarantees and NUMA awareness in favor of speed and simplicity, it remains a robust choice for high-performance queuing in multi-threaded applications, backed by thorough testing and an active ecosystem of usage patterns and tooling. If you’re building a high-throughput producer-consumer pipeline, this queue is well worth a serious look, and its single-header footprint makes integration particularly painless.
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/cameron314/concurrentqueue
GitHub - cameron314/concurrentqueue: moodycamel::ConcurrentQueue: An industrial-strength lock-free queue for C++
Moodycamel’s ConcurrentQueue is an industrial-strength, lock-free queue designed for multi-producer, multi-consumer (MPMC) workloads in C++11 and beyond. It's a...
github - cameron314/concurrentqueue


