System Call
System Call
Definition: The programmatic interface by which a user-mode application requests privileged services from the operating system kernel.
How It Works
- The CPU operates in privilege rings (or “modes”): user mode (Ring 3 on x86) runs application code with restricted access to hardware and memory; kernel mode (Ring 0) runs OS code with full hardware access.
- To make a system call, an application places a syscall number and its arguments into specific CPU registers, then executes a trap instruction,
syscallon modern x86-64,svcon ARM, historicallyint 0x80on older x86 Linux. - That instruction causes a controlled transition into kernel mode: the CPU switches privilege level, jumps to a fixed, kernel-defined entry point, and the kernel’s syscall dispatcher reads the syscall number from the register and looks it up in a syscall table, an array of function pointers, to find and run the corresponding kernel handler.
- The kernel handler executes with full privileges, does the actual work (reading a file, allocating memory, sending a network packet), places a return value/error code into a register, and executes a return-from-trap instruction that switches the CPU back to user mode and resumes the application right after the trap instruction.
- Standard library functions like
read(),write(),open(), andfork()are thin wrappers around this trap mechanism; the actual syscall convention (which register holds what) is architecture- and OS-specific, which is exactly what the C library abstracts away from application code.
Under the Hood
The trap instruction itself is the security boundary. Because entering kernel mode is only possible through this single, fixed, kernel-controlled entry point, and the CPU itself refuses to execute privileged instructions from user mode without it, user code can’t simply jump into arbitrary kernel code or fabricate its own kernel-mode execution. The kernel’s syscall entry handler validates every argument, pointers get checked that they actually point into the calling process’s valid address space, sizes get bounds-checked, before trusting anything the user process passed in, since a malicious or buggy program can put arbitrary garbage in those registers.
On Linux x86-64, the syscall table is sys_call_table, and each syscall has a fixed number, read is 0, write is 1, open is 2, and so on, stable across kernel versions for ABI compatibility. strace works by using the ptrace() syscall to intercept every syscall entry and exit for a traced process, reading the syscall number and arguments straight out of the registers at the trap boundary, which is why it can show every syscall a program makes without needing the program’s source code or cooperation.
Not every kernel entry is a deliberate syscall. Page faults, timer interrupts, and hardware interrupts also force a mode switch into kernel code, but through different CPU exception/interrupt mechanisms, not the syscall trap instruction, even though the underlying “drop into Ring 0, run kernel code, return to Ring 3” pattern is the same. Some frequently-called, read-only kernel data (like the current time) is exposed through the vDSO (virtual dynamic shared object), a small chunk of kernel code mapped into every process’s address space, letting calls like gettimeofday() run entirely in user mode with no trap at all, because the actual mode switch is the expensive part of a syscall, not the kernel logic itself.
Debugging Workflow
The default tool for understanding what a program is actually asking the kernel to do is strace, which attaches via ptrace() and prints every syscall, its arguments, and its return value:
$ strace -f -tt ./app
14:22:01.001234 openat(AT_FDCWD, "/etc/config.json", O_RDONLY) = 3
14:22:01.001456 read(3, "{\"timeout\":30}", 4096) = 15
14:22:01.001501 close(3) = 0
14:22:01.031820 connect(4, {sa_family=AF_INET, ...}, 16) = -1 EINPROGRESS
The -tt timestamps make it easy to spot a syscall that took an unexpectedly long time between its line and the next, exactly where a program is stuck. strace -c ./program instead aggregates: total time, call count, and average latency per syscall, the fastest way to answer “is this program slow because of the kernel or because of its own code.”
$ strace -f -e trace=network -p <pid> # only network-related syscalls
$ perf trace -p <pid> # lower-overhead alternative to strace
$ ltrace ./app # library calls instead of syscalls, for comparison
strace’s ptrace()-based interception adds real overhead, often 2-10x slowdown, since every syscall now round-trips through the tracer; perf trace uses kernel tracepoints instead and is significantly cheaper, worth reaching for for production or performance-sensitive tracing rather than a full strace attach.
Why It Matters
- Protects hardware, disk, memory, and network resources from unverified user-level code. Every disk write, every network packet, every process creation must pass through kernel-validated code, which is what makes the OS the enforcement point for security and resource isolation between processes.
- The syscall interface is the actual stable contract between applications and the kernel. Linux’s promise to “never break userspace” is fundamentally a promise about syscall numbers and semantics staying compatible, letting binaries compiled years ago still run on a current kernel.
- Understanding the user-mode/kernel-mode boundary explains why sandboxing technologies (seccomp-bpf, gVisor) work by filtering or intercepting exactly this syscall interface, restricting which syscalls a process is allowed to make is a direct way to shrink its attack surface against the kernel.
Common Pitfalls
- Frequent, small system calls introduce real overhead from the repeated mode switches, which is why user-space buffering (
fwrite/stdiobuffering, rather than callingwrite()per byte) exists specifically to batch work into fewer syscalls. - Assuming a syscall like
write()is synchronous all the way to disk. It typically only copies data into the kernel’s page cache and returns; actual durability requiresfsync(). See File Systems. - Not checking syscall return values for errors. Nearly every syscall can fail (
ENOMEM,EACCES,EINTRfrom a signal interrupting a blocking call), and ignoring this is a common source of subtle bugs that only appear under resource pressure or signal delivery. - Treating syscalls as free in hot loops. Tools like
strace -corperf tracereveal that a surprising fraction of “slow” application code turns out to be syscall-bound, not CPU-bound, once you actually measure it.
History
- Early Unix used a software interrupt,
int 0x80on x86, to trap into the kernel. It worked but was relatively slow because software interrupts weren’t designed or optimized specifically for this frequent use case. - Intel and AMD later added dedicated fast syscall instructions,
sysenter/sysexitandsyscall/sysretrespectively, purpose-built to minimize the mode-switch overhead compared to a general interrupt, and modern 64-bit Linux usessyscallexclusively. - Sandboxing built directly on the syscall interface, seccomp-bpf (2012, Linux), lets a process install a filter restricting exactly which syscalls, and with which arguments, it’s allowed to make, a foundational building block for container runtimes and browser sandboxes today.
FAQ
Does every library function make a system call? No. Many, like strlen() or memcpy(), do pure computation in userspace with no kernel involvement at all. Only functions that need a kernel-managed resource, files, memory, processes, network, actually trap into the kernel.
Why is fork() considered an expensive syscall relative to others? Because it must set up an entire new address space for the child, even with copy-on-write deferring the actual page copying, the page table setup and PCB creation is real work compared to a simple syscall like getpid().
Can user-mode code intercept another process’s system calls? Not directly without kernel help. ptrace() (what strace and debuggers use) is itself a syscall that the kernel provides specifically to let one process observe another’s execution, including syscalls, under controlled, kernel-mediated conditions.
Comparison
| System Call | Library Function Call | Signal | |
|---|---|---|---|
| Crosses privilege boundary | Yes (user to kernel mode) | No (stays in user mode) | Delivered by kernel, handled in user mode |
| Mechanism | Trap instruction (syscall/svc) | Ordinary function call | Kernel-delivered asynchronous interrupt to a process |
| Overhead | Higher (mode switch) | Low (normal call) | Delivery cost similar to a syscall; handling is user-mode |
| Example | read(), write(), fork() | strlen(), printf() (may wrap syscalls internally) | SIGSEGV, SIGTERM |
| Initiator | The process itself, deliberately | The process itself, deliberately | The kernel or another process, asynchronously |
| Can be intercepted/traced | Yes (ptrace/strace) | Not directly (no kernel boundary crossed) | Yes (signal handlers) |
Common Syscalls Quick Reference
| Syscall | Purpose |
|---|---|
read() / write() | Transfer bytes to/from a file descriptor |
open() / close() | Obtain/release a file descriptor |
fork() / execve() | Duplicate a process / replace its memory image with a new program |
mmap() / munmap() | Map/unmap memory, including files, into a process’s address space |
socket() / connect() / bind() | Create and configure network endpoints |
wait() / waitpid() | Block until a child process changes state (commonly, exits) |
ioctl() | Device-specific control operations that don’t fit the read/write model |
Example
Calling read() in C compiles down to a trap into the kernel:
ssize_t n = read(fd, buf, count);
1. libc's read() places syscall number (0 on x86-64 Linux), fd, buf, count into registers
2. CPU executes `syscall`, switching from Ring 3 (user) to Ring 0 (kernel)
3. Kernel's sys_read() handler validates fd and buf, copies data from the file into buf
4. Return value (bytes read, or -1 with errno set) placed in a register
5. CPU executes sysret, switching back to Ring 3, program resumes after the syscall
strace -c ./program summarizes every syscall a program makes and how much time is spent in each, a standard first step when diagnosing whether a slow program is spending its time in the kernel or in its own user-mode code.
Related Terms
Referenced by