Linux membarrier() can move part of a synchronization cost from a frequently executed path to a less frequent coordination path. With MEMBARRIER_CMD_PRIVATE_EXPEDITED, one thread asks the kernel to establish a memory-ordering point across the running threads in the same process. The caller pays for the system call when coordination is needed instead of requiring every fast-path execution to carry an explicit hardware memory barrier.

This is narrower than a general thread rendezvous. membarrier() does not run an application callback on sibling threads, does not wait for application-level acknowledgements, and does not turn ordinary data races into valid synchronization. Its contract concerns ordering of userspace memory accesses around the barrier.

The private expedited command targets one process

MEMBARRIER_CMD_PRIVATE_EXPEDITED applies to running thread siblings of the calling thread. When the call returns successfully, those running siblings have passed through a state in which their userspace memory accesses are ordered with respect to the membarrier() operation. Threads that are not running are already considered to satisfy the required state.

The scope matters. The private command does not impose the corresponding barrier on unrelated processes. Linux also provides global forms, but their target set and registration rules are different.

A process must register before it can issue the private expedited operation:

syscall(SYS_membarrier,
        MEMBARRIER_CMD_REGISTER_PRIVATE_EXPEDITED,
        0, 0);

syscall(SYS_membarrier,
        MEMBARRIER_CMD_PRIVATE_EXPEDITED,
        0, 0);

There is no glibc wrapper for membarrier(), so applications invoke it through syscall().

Registration is not optional bookkeeping. Issuing a private expedited command without the required prior registration fails with EPERM. Registration lets the kernel prepare the process for the expedited mechanism before the latency-sensitive barrier request occurs.

Querying support is part of initialization

MEMBARRIER_CMD_QUERY returns a bit mask describing commands supported by the running kernel. Software can test the bits for both registration and execution commands before selecting this synchronization strategy.

int mask = syscall(SYS_membarrier, MEMBARRIER_CMD_QUERY, 0, 0);

if ((mask & MEMBARRIER_CMD_PRIVATE_EXPEDITED) &&
    (mask & MEMBARRIER_CMD_REGISTER_PRIVATE_EXPEDITED)) {
    /* This mode is available. */
}

The query boundary is useful because membarrier() is Linux-specific and command availability varies with kernel and architecture capabilities. Code that depends on an expedited mode should establish support during initialization rather than assuming that a kernel version alone implies every command is usable.

For calls whose flags value is zero, the documented result for a given command remains stable until reboot. That makes one-time capability and registration checks a natural fit for process startup.

Expedited shifts cost toward the coordination side

A conventional synchronization design can place hardware memory barriers in a path executed by many threads. If that path is extremely frequent, even a small per-operation cost can dominate the less frequent coordination event.

membarrier() supports the opposite allocation. The frequent path can use compiler barriers where the algorithm permits, while the infrequent path invokes the syscall to force the required cross-thread ordering state. Linux documentation cites Read-Copy-Update libraries and garbage collectors as examples of systems where this pattern can be useful.

The trade is not free. Expedited commands cause kernel work and extra overhead when invoked. Their value depends on an asymmetric workload: the fast side must be frequent enough, and the barrier side infrequent enough, for moving cost to the slow side to be beneficial.

That is an algorithmic condition rather than a universal performance claim. A design that invokes membarrier() frequently can lose the advantage it was intended to create.

A compiler barrier and a CPU memory barrier are different

The synchronization pattern relies on a distinction between compiler ordering and hardware memory ordering. A compiler barrier can prevent the compiler from moving memory operations across a point in generated code, but it does not by itself force another CPU core to observe memory accesses in the order required by a concurrent algorithm.

A full CPU memory barrier provides hardware ordering constraints, but placing one on every fast-path execution can be costly on architectures or workloads where that instruction has meaningful overhead.

membarrier() supplies an ordering relation between the caller’s barrier operation and memory accesses performed by targeted threads. It does not erase the need to reason about compiler transformations, atomic operations, or the language memory model. Correct algorithms still need the appropriate compiler-level and language-level primitives around shared state.

This boundary is especially important in C and C++. A kernel ordering primitive cannot retroactively make a program with undefined behavior from data races valid under the language rules.

Registration and execution are separate lifecycle phases

The two-command shape creates a useful lifecycle boundary. Registration is a setup operation; execution is the synchronization event that may occur repeatedly afterward.

A runtime can query support, register once, and only then enable a fast-path design that assumes private expedited barriers are available. If registration fails, the runtime can keep a fallback synchronization strategy instead of entering a mode whose ordering contract cannot be fulfilled.

This separation also avoids treating EPERM during execution as a transient scheduling failure. For the private expedited command, EPERM indicates that the process did not complete the required registration.

The command set includes stronger or specialized variants as well. MEMBARRIER_CMD_PRIVATE_EXPEDITED_SYNC_CORE adds a core-serialization guarantee, while the RSEQ variant can restart running restartable-sequence critical sections. Those commands have their own registration and support requirements and should not be treated as aliases for the basic private expedited barrier.

The barrier is not a stop-the-world primitive

A successful private expedited call establishes the documented ordering state, but sibling threads continue under normal scheduling. The syscall does not expose a point at which application code may safely inspect arbitrary non-atomic shared objects without the synchronization required by the program’s concurrency model.

It also does not provide a transactional snapshot of process memory. Threads can continue modifying shared state before and after the ordering point. Algorithms using membarrier() therefore need a protocol that gives the ordering event meaning, such as publishing state before the call and interpreting observations after the call according to the algorithm’s atomic and compiler-ordering rules.

This is the core boundary of the interface: membarrier() can provide an efficient kernel-assisted ordering event across a defined thread set, but the surrounding concurrency protocol still determines which state transitions that event makes safe.