System Call

System Call

Definition: The programmatic interface by which a user-mode application requests privileged services from the operating system kernel.

How It Works

  • The CPU operates in privilege rings (or “modes”): user mode (Ring 3 on x86) runs application code with restricted access to hardware and memory; kernel mode (Ring 0) runs OS code with full hardware access.
  • To make a system call, an application places a syscall number and its arguments into specific CPU registers, then executes a trap instruction, syscall on modern x86-64, svc on ARM, historically int 0x80 on older x86 Linux.
  • That instruction causes a controlled transition into kernel mode: the CPU switches privilege level, jumps to a fixed, kernel-defined entry point, and the kernel’s syscall dispatcher reads the syscall number from the register and looks it up in a syscall table, an array of function pointers, to find and run the corresponding kernel handler.
  • The kernel handler executes with full privileges, does the actual work (reading a file, allocating memory, sending a network packet), places a return value/error code into a register, and executes a return-from-trap instruction that switches the CPU back to user mode and resumes the application right after the trap instruction.
  • Standard library functions like read(), write(), open(), and fork() are thin wrappers around this trap mechanism; the actual syscall convention (which register holds what) is architecture- and OS-specific, which is exactly what the C library abstracts away from application code.

Under the Hood

The trap instruction itself is the security boundary. Because entering kernel mode is only possible through this single, fixed, kernel-controlled entry point, and the CPU itself refuses to execute privileged instructions from user mode without it, user code can’t simply jump into arbitrary kernel code or fabricate its own kernel-mode execution. The kernel’s syscall entry handler validates every argument, pointers get checked that they actually point into the calling process’s valid address space, sizes get bounds-checked, before trusting anything the user process passed in, since a malicious or buggy program can put arbitrary garbage in those registers.

On Linux x86-64, the syscall table is sys_call_table, and each syscall has a fixed number, read is 0, write is 1, open is 2, and so on, stable across kernel versions for ABI compatibility. strace works by using the ptrace() syscall to intercept every syscall entry and exit for a traced process, reading the syscall number and arguments straight out of the registers at the trap boundary, which is why it can show every syscall a program makes without needing the program’s source code or cooperation.

Not every kernel entry is a deliberate syscall. Page faults, timer interrupts, and hardware interrupts also force a mode switch into kernel code, but through different CPU exception/interrupt mechanisms, not the syscall trap instruction, even though the underlying “drop into Ring 0, run kernel code, return to Ring 3” pattern is the same. Some frequently-called, read-only kernel data (like the current time) is exposed through the vDSO (virtual dynamic shared object), a small chunk of kernel code mapped into every process’s address space, letting calls like gettimeofday() run entirely in user mode with no trap at all, because the actual mode switch is the expensive part of a syscall, not the kernel logic itself.

Debugging Workflow

The default tool for understanding what a program is actually asking the kernel to do is strace, which attaches via ptrace() and prints every syscall, its arguments, and its return value:

$ strace -f -tt ./app
14:22:01.001234 openat(AT_FDCWD, "/etc/config.json", O_RDONLY) = 3
14:22:01.001456 read(3, "{\"timeout\":30}", 4096) = 15
14:22:01.001501 close(3)               = 0
14:22:01.031820 connect(4, {sa_family=AF_INET, ...}, 16) = -1 EINPROGRESS

The -tt timestamps make it easy to spot a syscall that took an unexpectedly long time between its line and the next, exactly where a program is stuck. strace -c ./program instead aggregates: total time, call count, and average latency per syscall, the fastest way to answer “is this program slow because of the kernel or because of its own code.”

$ strace -f -e trace=network -p <pid>     # only network-related syscalls
$ perf trace -p <pid>                      # lower-overhead alternative to strace
$ ltrace ./app                             # library calls instead of syscalls, for comparison

strace’s ptrace()-based interception adds real overhead, often 2-10x slowdown, since every syscall now round-trips through the tracer; perf trace uses kernel tracepoints instead and is significantly cheaper, worth reaching for for production or performance-sensitive tracing rather than a full strace attach.

Why It Matters

  • Protects hardware, disk, memory, and network resources from unverified user-level code. Every disk write, every network packet, every process creation must pass through kernel-validated code, which is what makes the OS the enforcement point for security and resource isolation between processes.
  • The syscall interface is the actual stable contract between applications and the kernel. Linux’s promise to “never break userspace” is fundamentally a promise about syscall numbers and semantics staying compatible, letting binaries compiled years ago still run on a current kernel.
  • Understanding the user-mode/kernel-mode boundary explains why sandboxing technologies (seccomp-bpf, gVisor) work by filtering or intercepting exactly this syscall interface, restricting which syscalls a process is allowed to make is a direct way to shrink its attack surface against the kernel.

Common Pitfalls

  • Frequent, small system calls introduce real overhead from the repeated mode switches, which is why user-space buffering (fwrite/stdio buffering, rather than calling write() per byte) exists specifically to batch work into fewer syscalls.
  • Assuming a syscall like write() is synchronous all the way to disk. It typically only copies data into the kernel’s page cache and returns; actual durability requires fsync(). See File Systems.
  • Not checking syscall return values for errors. Nearly every syscall can fail (ENOMEM, EACCES, EINTR from a signal interrupting a blocking call), and ignoring this is a common source of subtle bugs that only appear under resource pressure or signal delivery.
  • Treating syscalls as free in hot loops. Tools like strace -c or perf trace reveal that a surprising fraction of “slow” application code turns out to be syscall-bound, not CPU-bound, once you actually measure it.

History

  • Early Unix used a software interrupt, int 0x80 on x86, to trap into the kernel. It worked but was relatively slow because software interrupts weren’t designed or optimized specifically for this frequent use case.
  • Intel and AMD later added dedicated fast syscall instructions, sysenter/sysexit and syscall/sysret respectively, purpose-built to minimize the mode-switch overhead compared to a general interrupt, and modern 64-bit Linux uses syscall exclusively.
  • Sandboxing built directly on the syscall interface, seccomp-bpf (2012, Linux), lets a process install a filter restricting exactly which syscalls, and with which arguments, it’s allowed to make, a foundational building block for container runtimes and browser sandboxes today.

FAQ

Does every library function make a system call? No. Many, like strlen() or memcpy(), do pure computation in userspace with no kernel involvement at all. Only functions that need a kernel-managed resource, files, memory, processes, network, actually trap into the kernel.

Why is fork() considered an expensive syscall relative to others? Because it must set up an entire new address space for the child, even with copy-on-write deferring the actual page copying, the page table setup and PCB creation is real work compared to a simple syscall like getpid().

Can user-mode code intercept another process’s system calls? Not directly without kernel help. ptrace() (what strace and debuggers use) is itself a syscall that the kernel provides specifically to let one process observe another’s execution, including syscalls, under controlled, kernel-mediated conditions.

Comparison

System CallLibrary Function CallSignal
Crosses privilege boundaryYes (user to kernel mode)No (stays in user mode)Delivered by kernel, handled in user mode
MechanismTrap instruction (syscall/svc)Ordinary function callKernel-delivered asynchronous interrupt to a process
OverheadHigher (mode switch)Low (normal call)Delivery cost similar to a syscall; handling is user-mode
Exampleread(), write(), fork()strlen(), printf() (may wrap syscalls internally)SIGSEGV, SIGTERM
InitiatorThe process itself, deliberatelyThe process itself, deliberatelyThe kernel or another process, asynchronously
Can be intercepted/tracedYes (ptrace/strace)Not directly (no kernel boundary crossed)Yes (signal handlers)

Common Syscalls Quick Reference

SyscallPurpose
read() / write()Transfer bytes to/from a file descriptor
open() / close()Obtain/release a file descriptor
fork() / execve()Duplicate a process / replace its memory image with a new program
mmap() / munmap()Map/unmap memory, including files, into a process’s address space
socket() / connect() / bind()Create and configure network endpoints
wait() / waitpid()Block until a child process changes state (commonly, exits)
ioctl()Device-specific control operations that don’t fit the read/write model

Example

Calling read() in C compiles down to a trap into the kernel:

ssize_t n = read(fd, buf, count);
1. libc's read() places syscall number (0 on x86-64 Linux), fd, buf, count into registers
2. CPU executes `syscall`, switching from Ring 3 (user) to Ring 0 (kernel)
3. Kernel's sys_read() handler validates fd and buf, copies data from the file into buf
4. Return value (bytes read, or -1 with errno set) placed in a register
5. CPU executes sysret, switching back to Ring 3, program resumes after the syscall

strace -c ./program summarizes every syscall a program makes and how much time is spent in each, a standard first step when diagnosing whether a slow program is spending its time in the kernel or in its own user-mode code.

Dig deeper