Shared Memory, Semaphores, and eventfd
Zero-copy IPC and lightweight synchronization primitives
POSIX shared memory
POSIX shared memory creates a named object backed by tmpfs that multiple processes can mmap into their address spaces:
#include <sys/mman.h>
#include <fcntl.h>
/* Process A: create and write */
int fd = shm_open("/myshm",
O_CREAT | O_RDWR,
0600);
ftruncate(fd, 4096); /* set size */
void *ptr = mmap(NULL, 4096,
PROT_READ | PROT_WRITE,
MAP_SHARED, fd, 0);
close(fd); /* fd no longer needed after mmap */
*(int *)ptr = 42;
munmap(ptr, 4096);
/* Remove the name (pages freed when last mmap is gone) */
shm_unlink("/myshm");
/* Process B: open and read */
int fd = shm_open("/myshm", O_RDONLY, 0);
void *ptr = mmap(NULL, 4096, PROT_READ, MAP_SHARED, fd, 0);
close(fd);
int val = *(int *)ptr; /* val == 42 */
munmap(ptr, 4096);
shm_open is implemented as open on /dev/shm/ (a tmpfs mount). The shared pages are in the page cache, mapped into both processes' page tables.
System V shared memory
The older but more widely available SHM API:
#include <sys/shm.h>
#include <sys/ipc.h>
/* Create/get a segment */
key_t key = ftok("/tmp/myfile", 'A'); /* generate key from path+id */
int shmid = shmget(key,
4096, /* size */
IPC_CREAT | 0600); /* flags */
/* Attach (map into address space) */
void *ptr = shmat(shmid, NULL, 0); /* NULL = kernel chooses addr */
/* Use it */
*(int *)ptr = 99;
/* Detach */
shmdt(ptr);
/* Remove (kernel frees when all processes detach) */
shmctl(shmid, IPC_RMID, NULL);
# Show System V shared memory
ipcs -m
# ------ Shared Memory Segments --------
# key shmid owner perms bytes nattch status
# 0x00000000 0 root 600 4096 0
# Remove a segment
ipcrm -m <shmid>
Synchronization: POSIX semaphores
Shared memory requires external synchronization. POSIX semaphores provide a counter that can be atomically incremented/decremented:
Named semaphores (between unrelated processes)
#include <semaphore.h>
/* Process A */
sem_t *sem = sem_open("/mysem",
O_CREAT, 0600, 1); /* initial value = 1 */
sem_wait(sem); /* P(): decrement, block if 0 */
/* critical section */
sem_post(sem); /* V(): increment, wake one waiter */
sem_close(sem); /* close fd */
sem_unlink("/mysem"); /* remove name */
/* Process B */
sem_t *sem = sem_open("/mysem", 0); /* open existing */
sem_wait(sem);
/* ... */
sem_post(sem);
sem_close(sem);
Unnamed semaphores (shared memory or threads)
/* In shared memory, accessible to multiple processes */
struct shared {
sem_t sem;
int data;
};
struct shared *sh = mmap(NULL, sizeof(*sh),
PROT_READ|PROT_WRITE,
MAP_SHARED|MAP_ANONYMOUS, -1, 0);
sem_init(&sh->sem, 1, 1); /* pshared=1: shared between processes */
/* Process A/B: */
sem_wait(&sh->sem);
sh->data++;
sem_post(&sh->sem);
/* Cleanup */
sem_destroy(&sh->sem);
munmap(sh, sizeof(*sh));
Kernel implementation
POSIX semaphores use futexes internally:
/* glibc sem_wait: */
sem_wait(sem) {
if (__atomic_sub_fetch(&sem->value, 1, __ATOMIC_ACQ_REL) >= 0)
return; /* fast path: counter was > 0 */
/* slow path: wait */
futex_wait(&sem->value, ...);
}
sem_post(sem) {
if (__atomic_add_fetch(&sem->value, 1, __ATOMIC_RELEASE) <= 0)
futex_wake(&sem->value, 1); /* wake one waiter */
}
eventfd: lightweight event notification
eventfd creates a file descriptor backed by a 64-bit counter — the lightest-weight
notification mechanism, and the standard way to wake an epoll-driven event loop from another
thread or context:
#include <sys/eventfd.h>
int efd = eventfd(0, EFD_CLOEXEC | EFD_NONBLOCK); /* initial value 0 */
uint64_t val = 1;
write(efd, &val, sizeof(val)); /* signal: counter += 1 */
uint64_t count;
read(efd, &count, sizeof(count)); /* consume: returns counter, resets to 0 */
It's used extensively in QEMU/KVM virtio notifications, io_uring completion notification,
libuv/libevent event loop backends, and cgroup event notification. See
eventfd and signalfd for the full treatment — EFD_SEMAPHORE mode, the
epoll integration pattern, and the real kernel implementation (struct eventfd_ctx,
eventfd_write()/eventfd_read()).
memfd: anonymous file-backed shared memory
memfd_create creates an anonymous file in memory (no filesystem path), useful for sharing without name collisions:
#include <sys/mman.h>
/* Create anonymous file */
int fd = memfd_create("mydata", MFD_CLOEXEC | MFD_ALLOW_SEALING);
ftruncate(fd, size);
void *ptr = mmap(NULL, size, PROT_READ|PROT_WRITE, MAP_SHARED, fd, 0);
/* Seal: prevent future modifications (useful for read-only sharing) */
fcntl(fd, F_ADD_SEALS, F_SEAL_WRITE | F_SEAL_GROW | F_SEAL_SHRINK);
/* Pass fd to another process via Unix socket (SCM_RIGHTS) */
send_fd_over_socket(socket_fd, fd);
/* Other process can mmap the fd — no name needed */
memfd_create is used by:
- Graphics/Wayland: sharing framebuffers between client and compositor
- dbus-broker: when an oversized log entry doesn't fit in a single datagram to the systemd
journal socket, it's sealed into a memfd and sent as the datagram's payload instead — the
journal's own convention for large fields, not part of the D-Bus wire protocol itself
(
misc_memfd()in dbus-broker'ssrc/util/misc.c, used fromsrc/util/log.c)
Comparing shared memory approaches
| Approach | Setup | Name in fs | fd | Sealing | Best for |
|---|---|---|---|---|---|
POSIX shm (shm_open) |
Path in /dev/shm/ |
Yes | After open | No | Named, persistent |
System V (shmget) |
IPC key | Via ipcs |
No | No | Legacy compatibility |
mmap(MAP_SHARED | MAP_ANONYMOUS) |
Anonymous | No | No | Parent-child only | |
memfd_create |
Anonymous file | No | Yes | Yes | Dynamic, fd-passing |
Further reading
Kernel source
- ipc/shm.c —
shmget(),shmat()/shmdt(), andshmctl()syscalls;newseg()creates the segment's hidden backing file viashmem_kernel_file_setup()(orhugetlb_file_setup()forSHM_HUGETLB) - mm/shmem.c —
shmem_file_setup()/shmem_kernel_file_setup(): the tmpfs backing shared by both POSIXshm_open()objects and SysVshmget()segments - mm/memfd.c —
memfd_create()syscall and theF_SEAL_*sealing checks - fs/eventfd.c —
eventfd_write()/eventfd_read()and thestruct eventfd_ctxcounter/wait-queue implementation
Man pages
shmget(2)— allocate a System V shared memory segmentshmat(2)/shmdt(2)— attach/detach a System V segmentshmctl(2)— control operations, includingIPC_RMIDshm_open(3)— POSIX shared memory objects; documents the/dev/shmtmpfs implementation on Linuxshm_overview(7)— POSIX shared memory API overviewsem_overview(7)— POSIX semaphores: named vs. unnamed,/dev/shmpersistence for named semaphoresmemfd_create(2)— anonymous sealable memory fileseventfd(2)— counter-backed notification fd, includingEFD_SEMAPHORE
Related pages
- Process Address Space — VMAs and the
MAP_SHAREDflag that both POSIX shm and anonymous shared mappings rely on - File-Backed mmap and Page Faults —
MAP_PRIVATEvsMAP_SHAREDfault handling in detail - Futex Internals — how
sem_wait()/sem_post()fall back to the kernel when the uncontended fast-path atomic isn't enough - SysV IPC: Semaphores and Message Queues —
semget/semop,SEM_UNDO, and the SysV IPC leak problem - eventfd and signalfd —
struct eventfd_ctx, epoll integration, and signalfd covered in depth - Unix Domain Sockets —
SCM_RIGHTS, the generic mechanism for passing any open file descriptor (including amemfd_create()fd) between processes
LWN articles
- Sealed files (2014) — Jonathan Corbet on the
memfd_create()andF_SEAL_*sealing mechanism described in this page's memfd section
External
- tmpfs — The Linux Kernel documentation — confirms both POSIX
shm_open()//dev/shmand SysVshmget()segments are backed by tmpfs (SysV via an internal, non-visible mount)