Skip to content

Shared Memory, Semaphores, and eventfd

Zero-copy IPC and lightweight synchronization primitives

POSIX shared memory

POSIX shared memory creates a named object backed by tmpfs that multiple processes can mmap into their address spaces:

#include <sys/mman.h>
#include <fcntl.h>

/* Process A: create and write */
int fd = shm_open("/myshm",
                  O_CREAT | O_RDWR,
                  0600);
ftruncate(fd, 4096);  /* set size */

void *ptr = mmap(NULL, 4096,
                 PROT_READ | PROT_WRITE,
                 MAP_SHARED, fd, 0);
close(fd);  /* fd no longer needed after mmap */

*(int *)ptr = 42;
munmap(ptr, 4096);

/* Remove the name (pages freed when last mmap is gone) */
shm_unlink("/myshm");

/* Process B: open and read */
int fd = shm_open("/myshm", O_RDONLY, 0);
void *ptr = mmap(NULL, 4096, PROT_READ, MAP_SHARED, fd, 0);
close(fd);
int val = *(int *)ptr;   /* val == 42 */
munmap(ptr, 4096);

shm_open is implemented as open on /dev/shm/ (a tmpfs mount). The shared pages are in the page cache, mapped into both processes' page tables.

ls /dev/shm/   # see current shared memory objects
ipcs -m        # see System V shared memory segments

System V shared memory

The older but more widely available SHM API:

#include <sys/shm.h>
#include <sys/ipc.h>

/* Create/get a segment */
key_t key = ftok("/tmp/myfile", 'A');  /* generate key from path+id */
int shmid = shmget(key,
                   4096,               /* size */
                   IPC_CREAT | 0600);  /* flags */

/* Attach (map into address space) */
void *ptr = shmat(shmid, NULL, 0);     /* NULL = kernel chooses addr */

/* Use it */
*(int *)ptr = 99;

/* Detach */
shmdt(ptr);

/* Remove (kernel frees when all processes detach) */
shmctl(shmid, IPC_RMID, NULL);
# Show System V shared memory
ipcs -m
# ------ Shared Memory Segments --------
# key        shmid    owner   perms  bytes  nattch  status
# 0x00000000 0        root    600    4096   0

# Remove a segment
ipcrm -m <shmid>

Synchronization: POSIX semaphores

Shared memory requires external synchronization. POSIX semaphores provide a counter that can be atomically incremented/decremented:

Named semaphores (between unrelated processes)

#include <semaphore.h>

/* Process A */
sem_t *sem = sem_open("/mysem",
                      O_CREAT, 0600, 1);  /* initial value = 1 */

sem_wait(sem);       /* P(): decrement, block if 0 */
/* critical section */
sem_post(sem);       /* V(): increment, wake one waiter */

sem_close(sem);      /* close fd */
sem_unlink("/mysem"); /* remove name */

/* Process B */
sem_t *sem = sem_open("/mysem", 0);  /* open existing */
sem_wait(sem);
/* ... */
sem_post(sem);
sem_close(sem);

Unnamed semaphores (shared memory or threads)

/* In shared memory, accessible to multiple processes */
struct shared {
    sem_t sem;
    int   data;
};

struct shared *sh = mmap(NULL, sizeof(*sh),
                         PROT_READ|PROT_WRITE,
                         MAP_SHARED|MAP_ANONYMOUS, -1, 0);

sem_init(&sh->sem, 1, 1);  /* pshared=1: shared between processes */

/* Process A/B: */
sem_wait(&sh->sem);
sh->data++;
sem_post(&sh->sem);

/* Cleanup */
sem_destroy(&sh->sem);
munmap(sh, sizeof(*sh));

Kernel implementation

POSIX semaphores use futexes internally:

/* glibc sem_wait: */
sem_wait(sem) {
    if (__atomic_sub_fetch(&sem->value, 1, __ATOMIC_ACQ_REL) >= 0)
        return;  /* fast path: counter was > 0 */
    /* slow path: wait */
    futex_wait(&sem->value, ...);
}

sem_post(sem) {
    if (__atomic_add_fetch(&sem->value, 1, __ATOMIC_RELEASE) <= 0)
        futex_wake(&sem->value, 1);  /* wake one waiter */
}

eventfd: lightweight event notification

eventfd creates a file descriptor backed by a 64-bit counter — the lightest-weight notification mechanism, and the standard way to wake an epoll-driven event loop from another thread or context:

#include <sys/eventfd.h>

int efd = eventfd(0, EFD_CLOEXEC | EFD_NONBLOCK);  /* initial value 0 */

uint64_t val = 1;
write(efd, &val, sizeof(val));      /* signal: counter += 1 */

uint64_t count;
read(efd, &count, sizeof(count));   /* consume: returns counter, resets to 0 */

It's used extensively in QEMU/KVM virtio notifications, io_uring completion notification, libuv/libevent event loop backends, and cgroup event notification. See eventfd and signalfd for the full treatment — EFD_SEMAPHORE mode, the epoll integration pattern, and the real kernel implementation (struct eventfd_ctx, eventfd_write()/eventfd_read()).

memfd: anonymous file-backed shared memory

memfd_create creates an anonymous file in memory (no filesystem path), useful for sharing without name collisions:

#include <sys/mman.h>

/* Create anonymous file */
int fd = memfd_create("mydata", MFD_CLOEXEC | MFD_ALLOW_SEALING);
ftruncate(fd, size);

void *ptr = mmap(NULL, size, PROT_READ|PROT_WRITE, MAP_SHARED, fd, 0);

/* Seal: prevent future modifications (useful for read-only sharing) */
fcntl(fd, F_ADD_SEALS, F_SEAL_WRITE | F_SEAL_GROW | F_SEAL_SHRINK);

/* Pass fd to another process via Unix socket (SCM_RIGHTS) */
send_fd_over_socket(socket_fd, fd);
/* Other process can mmap the fd — no name needed */

memfd_create is used by:

  • Graphics/Wayland: sharing framebuffers between client and compositor
  • dbus-broker: when an oversized log entry doesn't fit in a single datagram to the systemd journal socket, it's sealed into a memfd and sent as the datagram's payload instead — the journal's own convention for large fields, not part of the D-Bus wire protocol itself (misc_memfd() in dbus-broker's src/util/misc.c, used from src/util/log.c)

Comparing shared memory approaches

Approach Setup Name in fs fd Sealing Best for
POSIX shm (shm_open) Path in /dev/shm/ Yes After open No Named, persistent
System V (shmget) IPC key Via ipcs No No Legacy compatibility
mmap(MAP_SHARED | MAP_ANONYMOUS) Anonymous No No Parent-child only
memfd_create Anonymous file No Yes Yes Dynamic, fd-passing

Further reading

Kernel source

  • ipc/shm.c — shmget(), shmat()/shmdt(), and shmctl() syscalls; newseg() creates the segment's hidden backing file via shmem_kernel_file_setup() (or hugetlb_file_setup() for SHM_HUGETLB)
  • mm/shmem.c — shmem_file_setup()/shmem_kernel_file_setup(): the tmpfs backing shared by both POSIX shm_open() objects and SysV shmget() segments
  • mm/memfd.c — memfd_create() syscall and the F_SEAL_* sealing checks
  • fs/eventfd.c — eventfd_write()/eventfd_read() and the struct eventfd_ctx counter/wait-queue implementation

Man pages

  • shmget(2) — allocate a System V shared memory segment
  • shmat(2)/shmdt(2) — attach/detach a System V segment
  • shmctl(2) — control operations, including IPC_RMID
  • shm_open(3) — POSIX shared memory objects; documents the /dev/shm tmpfs implementation on Linux
  • shm_overview(7) — POSIX shared memory API overview
  • sem_overview(7) — POSIX semaphores: named vs. unnamed, /dev/shm persistence for named semaphores
  • memfd_create(2) — anonymous sealable memory files
  • eventfd(2) — counter-backed notification fd, including EFD_SEMAPHORE

LWN articles

  • Sealed files (2014) — Jonathan Corbet on the memfd_create() and F_SEAL_* sealing mechanism described in this page's memfd section

External