
A Deep Dive into Producer-Consumer Queues in C++
Six MPMC queue variants measured across producer/consumer mixes on x86 and ARM. Splitting one lock into two is 6.8x faster for fan-out and slower for fan-in. A deque beats a linked list 5x uncontended. And the object pool we added to skip malloc made things worse. Real numbers, real hardware.







