Process and Thread
Process and Thread
Definition: A process is an isolated executing instance of a program with its own memory space; a thread is a lightweight execution path within a process that shares memory with sibling threads.
How It Works
- Processes maintain distinct address spaces, code, data, heap, and stack segments, guarded by the MMU and enforced by the OS, so one process cannot read or corrupt another’s memory without going through an explicit IPC mechanism.
- Threads within the same process share the heap, code segment, and open file descriptor table, but each thread keeps its own set of CPU registers, program counter, and stack, so it can execute independently while still reading/writing the same shared heap data as its sibling threads.
- Every process has a Process Control Block (PCB) in the kernel: PID, memory map pointers, open file table, scheduling state, and saved register values when not running. Threads have an analogous, lighter Thread Control Block (TCB), which mostly just needs its own register/stack state since it shares everything else with the parent process.
- Context switching: the OS saves the currently running task’s execution state (registers, program counter) into its PCB/TCB and restores the next task’s saved state into the CPU. Switching between threads of the same process skips reloading the page table (CR3 register on x86) since the address space is identical; switching between different processes must reload it, which flushes the TLB and is significantly more expensive.
- Process creation on Unix-like systems happens via
fork()(duplicates the calling process, parent and child then diverge) often paired withexec()(replaces the calling process’s memory image with a new program). Thread creation happens viapthread_create()or a language runtime’s thread API, which allocates a new stack and TCB within the existing process.
Under the Hood
fork() historically copied the entire parent address space into the child, which was expensive. Modern implementations use copy-on-write (COW): the child gets its own page table, but each entry initially points at the same physical pages as the parent, marked read-only. Only when either process actually writes to a page does the kernel take a page fault, allocate a real private copy for the writer, and update that one page table entry, deferring the cost of copying until it’s actually needed and often avoiding most of it entirely for a fork()+exec() pair, since the child usually calls exec() (discarding the copied address space) before writing much of anything.
On Linux, there isn’t actually a separate kernel primitive for “process” versus “thread.” Both are created by the same underlying clone() syscall; what differs is which resources are shared versus duplicated, controlled by clone flags like CLONE_VM (share the address space), CLONE_FILES (share the file descriptor table), and CLONE_SIGHAND (share signal handlers). fork() calls clone() with none of these sharing flags set (full duplication, with COW for memory); pthread_create() calls clone() with essentially all of them set. This is why Linux tools sometimes call threads “lightweight processes” (LWPs), each thread genuinely has its own kernel scheduling entity (its own entry in /proc/<pid>/task/), it just shares almost everything with its siblings.
Every thread and process is represented by exactly one task_struct in the kernel, regardless of which one it is; there’s no separate struct thread type. What distinguishes a “process” from one of its “threads” is purely bookkeeping: all task_structs created via clone(CLONE_THREAD) share the same thread-group ID (tgid), which is what user-space tools report as the PID, while each thread additionally has its own unique pid value internally. getpid() in a multi-threaded program actually returns the shared tgid, not each thread’s individual kernel-level ID, which is why every thread in a process appears to have “the same PID” from user space even though the kernel schedules and tracks them as distinct entities.
History
- Early Unix systems (1970s-80s) had only processes; there was no OS-level thread concept, concurrency within a program meant either multiple processes or user-space cooperative scheduling with no kernel awareness at all.
- Kernel-level threading support arrived through the 1990s in different forms across Unix variants, eventually standardized by POSIX Threads (Pthreads) in 1995, giving portable thread creation, mutexes, and condition variables across compliant systems.
- Linux’s
clone()syscall (added early, refined through the 1990s-2000s) took a different architectural path than dedicated thread primitives in other kernels, generalizing process and thread creation into one syscall parameterized by which resources to share, which is why Linux threads and processes remain conceptually unified at the kernel level today.
Debugging Workflow
Diagnosing which of many threads in a process is stuck, consuming CPU, or leaking usually starts with per-thread visibility rather than treating the process as one opaque unit:
$ ps -eLf | grep <pid> # every thread of a process, each with its own LWP id
$ top -H -p <pid> # live per-thread CPU usage
$ cat /proc/<pid>/status | grep Threads # thread count
$ cat /proc/<pid>/task/<tid>/stack # kernel-side blocking point for one thread
A thread count that climbs steadily over a service’s lifetime without leveling off is a thread leak, commonly a thread pool that creates workers but never joins/reaps finished ones. Cross-referencing top -H’s per-thread CPU column against /proc/<pid>/task/<tid>/stack for the busiest thread narrows down whether a specific thread is spinning (high CPU, changing stack) or genuinely stuck (zero CPU, unchanging blocked stack) far faster than reasoning about the process as a whole.
Why It Matters
- Underpins multitasking OS design, multi-core CPU utilization, and essentially every concurrency model in modern software, from browser tab isolation to web server request handling.
- The process/thread tradeoff is a real architectural decision: processes give strong isolation (a crash in one can’t corrupt another) at the cost of heavier creation and communication overhead; threads give cheap creation and fast shared-memory communication at the cost of one bad thread being able to corrupt the whole process.
- Multi-process isolation is a security boundary. Browsers moved to per-tab (or per-site) process isolation specifically so a memory-corruption bug in one tab’s renderer can’t read another tab’s memory, something impossible to guarantee between threads in the same process.
Common Pitfalls
- Sharing mutable memory across threads without locks causes data races and undefined behavior. See Concurrency and Race Condition for the full mechanics.
- Assuming thread creation is “free.” It’s much cheaper than a process, but each thread still needs its own stack (megabytes by default on many platforms) and TCB, so spawning tens of thousands of threads for tens of thousands of concurrent tasks doesn’t scale the way I-O Multiplexing or async I/O does.
- Forgetting that process context switching is significantly heavier than thread context switching, because it requires a page table (CR3) reload and the resulting TLB flush, on top of the register save/restore every context switch needs regardless of process or thread.
- Zombie processes: a child process that has exited but whose exit status hasn’t been read yet by the parent (via
wait()/waitpid()) stays in the process table as a zombie, consuming a PID slot until reaped. - Orphan processes: a child whose parent exits before it does gets re-parented (to
init/PID 1 on Unix, or a designated reaper), which is normal and handled automatically, but code that assumes a specific parent will always exist can break here. - Thread pool exhaustion: sizing a thread pool too small for a workload that blocks (waiting on I/O, waiting on another thread) causes requests to queue behind busy threads even though the CPU itself is mostly idle, a throughput problem that looks like a CPU bottleneck but isn’t.
- Assuming
fork()in a multi-threaded process is safe in general. Only the calling thread survives into the child; any lock held by a different thread at the moment offork()stays locked forever in the child, since the thread that would unlock it doesn’t exist there, a classic source of child-process hangs right afterfork().
Green Threads and Goroutines
Not every “thread” is a kernel-scheduled entity. Green threads (used historically by early Java, and by languages like Erlang) and Go’s goroutines are scheduled entirely by a language runtime in userspace, with the runtime multiplexing many logical threads onto a much smaller number of actual OS threads. This avoids the OS-level cost of thread creation and context switching for workloads that spawn huge numbers of short-lived concurrent tasks (Go programs routinely run hundreds of thousands of goroutines), at the cost of needing the runtime itself to handle blocking syscalls carefully, typically by handing them off to a dedicated OS thread so one blocked goroutine doesn’t stall the others multiplexed on the same underlying thread.
Comparison
| Process | Thread | |
|---|---|---|
| Address space | Own, isolated | Shared with sibling threads |
| Creation cost | Higher (new address space, COW setup) | Lower (new stack + TCB only) |
| Context switch cost | Higher (page table/TLB reload) | Lower (registers/stack only) |
| Crash isolation | Strong, one process can’t corrupt another | Weak, one thread’s bug can corrupt the whole process |
| Communication | Explicit IPC (pipes, sockets, shared memory) | Direct shared memory access |
| Linux kernel primitive | clone() with no sharing flags | clone() with CLONE_VM/CLONE_FILES/etc. |
| Kernel representation | Own task_struct, own tgid | Own task_struct, shares parent’s tgid |
| Typical stack size | N/A (has code/data/heap/stack) | 1-8MB default, configurable |
NUMA Considerations
On multi-socket servers, memory access isn’t uniform: each CPU socket has “local” RAM it can reach quickly and “remote” RAM attached to another socket that takes longer to access over the interconnect, a design called NUMA (Non-Uniform Memory Access). A process or thread scheduled on one socket but allocating memory that ends up on another socket’s RAM pays a real latency penalty on every access. NUMA-aware schedulers try to keep a thread and the memory it touches on the same node, and tools like numactl let an application pin threads and memory allocation to a specific NUMA node explicitly when the default placement isn’t good enough.
Example
A Chrome-style browser uses separate processes per tab for crash and security isolation, while a typical web server uses threads (or an async event loop) to handle many concurrent requests cheaply within one process:
Browser: [Tab A process] [Tab B process] [GPU process] [Network process]
↑ isolated address spaces, IPC via Chrome's Mojo/IPC layer
Web server: [Process] -> Thread 1 (request A)
-> Thread 2 (request B) # share heap, DB connection pool, etc.
-> Thread 3 (request C)
ps -eLf on Linux lists every thread of every process (the L flag), and htop can be toggled to show threads as separate rows, both reflecting the fact that the kernel schedules threads, not just whole processes.
FAQ
Why do people say threads are “lightweight processes”? Because on Linux they’re created by the same clone() syscall as processes, with sharing flags set so they reuse the parent’s address space, file table, and signal handlers instead of duplicating them, making creation and switching cheaper than a full process.
Can two threads in the same process have different memory protection? Not for the shared heap and code segment, those are identical across threads by definition. Each thread’s own stack, however, is a separate region, so a stack overflow in one thread doesn’t directly overwrite another thread’s stack.
Does killing a process kill all its threads? Yes. Terminating a process tears down its entire address space and every thread running within it; there’s no way for one thread to outlive the process that owns its shared memory.
Related Terms
Referenced by