Buffer Overflow and Memory Safety
Buffer Overflow and Memory Safety
Definition: A memory corruption vulnerability that occurs when a program writes data past the boundary of an allocated buffer, overwriting adjacent memory on the stack or heap.
How It Works
A buffer overflow happens whenever code writes more bytes into a fixed-size region than that region can hold, and nothing checks the write against the buffer’s actual size.
- Stack-based overflow: a local array declared on the stack (
char buf[64]) is filled with attacker-controlled input using an unbounded copy. The overflow spills into adjacent stack data: other locals, the saved frame pointer, and eventually the return address. - Heap-based overflow: a dynamically allocated block is overrun, corrupting the allocator’s bookkeeping metadata (chunk size fields, free-list pointers) or an adjacent object such as a function pointer or C++ vtable pointer.
- Global/BSS overflow: less common, but a fixed-size global buffer can be overrun into adjacent statically allocated data, including function pointers held in global structs.
- Root causes: unbounded string copies, off-by-one loop bounds, format string bugs (
printf(user_input)instead ofprintf("%s", user_input)), and integer overflows in a size calculation that produce an undersized allocation before a correctly-sized write into it.
If the attacker can control what ends up in the return address or a function pointer, they can redirect the CPU’s instruction pointer to code of their choosing the moment the function returns or the pointer is invoked.
Under the Hood
Stack frame layout on x86/x86-64 puts local buffers, then the saved base pointer, then the return address, then the caller’s arguments, while the stack grows toward lower addresses and a sequential write inside a buffer proceeds toward higher addresses. So strcpy(buf, input) with input longer than buf overwrites buf itself, then the saved base pointer, then the return address, in that order. An attacker crafts the input so the bytes landing at the return-address offset spell out the address of injected shellcode, or the address of the first gadget in a chain.
Concretely, on a 64-bit system: char buf[64] might sit at stack address 0x7ffee1000000, the saved base pointer at 0x7ffee1000040 (immediately after the 64-byte buffer), and the 8-byte return address at 0x7ffee1000048. An input of 80 bytes overflows buf by 16 bytes, exactly enough to overwrite the saved base pointer and the low bytes of the return address. A working exploit pads the first 64 bytes with filler, writes 8 bytes of arbitrary filler over the saved base pointer, then writes the target address as the final 8 bytes, in little-endian byte order since x86-64 stores multi-byte values least-significant-byte first (a target address 0x4141414142424242 is written to memory as the byte sequence 42 42 42 42 41 41 41 41).
Mitigations attack different points in that chain:
- Stack canaries: the compiler inserts a random value between local buffers and the saved registers. On function exit, the canary is checked before the return address is used; a smashed canary triggers an immediate abort. Canaries don’t help if the overflow leaks the canary value first through a separate info-disclosure bug, or if it overwrites a function pointer that gets called before the function returns.
- ASLR: randomizes the base addresses of the stack, heap, and loaded libraries on each process start, so hardcoded shellcode addresses from a previous run are wrong. Attackers defeat it with an information leak that reveals one real runtime address, then compute the rest by fixed offset.
- DEP/NX bit: marks stack and heap pages non-executable, so injected shellcode can’t simply be jumped to and run.
- ROP (return-oriented programming): the standard bypass for DEP/NX. The attacker chains short instruction sequences (“gadgets”) already present in the binary or libc, each ending in
ret, to perform arbitrary computation using only code that was always marked executable. - Heap corruption primitives: classic exploits overwrite the size and pointer fields of a heap chunk header so that when the chunk is later freed, the allocator’s unlink/coalesce logic performs an attacker-controlled write to an attacker-controlled address, an arbitrary-write primitive usable for further exploitation.
Stack Layout and the Overwrite Path
A normal 64-bit stack frame and the same frame after an unbounded strcpy(buf, attacker_input) look like this, address order running low to high:
The overflow doesn’t need to touch every byte between the buffer and the return address deliberately, it just needs to write enough bytes for the tail end of the input to land exactly on the return address’s offset. Everything before that offset in the payload is filler.
Why It Matters
- Buffer overflows remain a leading root cause of critical, remotely exploitable vulnerabilities in C and C++ codebases: browsers, kernels, network daemons, and embedded firmware.
- Unlike a logic bug, a memory corruption bug in the right spot gives an attacker arbitrary code execution with no valid credentials, which is why it consistently ranks among the highest-severity CVSS scores.
- Fixing the underlying language-level unsafety (moving to Rust, adding bounds checks, fuzzing C/C++ interfaces) is now a stated priority for major OS vendors, not just an application-level concern.
Common Pitfalls
- Using unbounded C string functions (
strcpy,gets,sprintf,strcat) instead of length-checked equivalents (strncpy,snprintf,strlcpy) - Assuming a stack canary makes a function immune to overflow-driven hijacking, when adjacent function pointers or heap objects can still be corrupted without ever touching the canary
- Computing a buffer size with signed or narrow integer arithmetic that silently wraps before the allocation
- Trusting a length field supplied by the client (a network packet, a file header) without validating it against the actual buffer size
- Disabling stack protection (
-fno-stack-protector) or NX for perceived performance gains without measuring the actual cost - Treating fuzzing or static analysis as optional for code that parses untrusted input
- Ignoring compiler warnings about implicit truncation or sign conversion in a size calculation, since that’s exactly where undersized-allocation bugs hide
- Assuming a crash from a fuzzer-found input is “just a denial of service” without triaging whether the same write primitive is also controllable enough to redirect execution
Comparison
| Buffer Overflow | Use-After-Free | Integer Overflow | Memory-Safe Language | |
|---|---|---|---|---|
| Trigger | Write past buffer bounds | Access to freed memory | Arithmetic wraps past type limit | N/A, enforced by design |
| Typical outcome | Return address / pointer overwrite | Type confusion, dangling pointer reuse | Undersized allocation, then overflow | Panic or exception instead of corruption |
| Common in | C, C++ | C, C++ | Any language with fixed-width integers | Rust, Go, Java, Python |
| Primary mitigation | Canaries, ASLR, DEP/NX, bounds checks | Smart pointers, garbage collection, borrow checking | Checked or saturating arithmetic | Compiler and runtime enforcement |
| Exploit reliability trend | Declining as mitigations layer up | Increasingly the preferred bug class in modern browsers/kernels | Often a stepping stone to another bug class, not the payload itself | N/A, bug class largely eliminated |
| Representative CVE | CVE-2021-3156 (sudo) | CVE-2019-0708 (BlueKeep, RDP) | CVE-2015-1538 (Stagefright, Android) | N/A |
| Found by | Manual review, fuzzing, ASan | Fuzzing, static use-after-free detectors | Static analysis, code review of arithmetic | N/A, compiler/runtime rejects the pattern |
| Typical fix | Bounds check, length-checked copy function | Smart pointer, nulling freed pointers, garbage collection | Checked/saturating arithmetic, wider integer type | N/A, language default |
Example
CVE-2014-0160 (Heartbleed) was a buffer over-read in OpenSSL’s heartbeat extension: the server trusted a client-supplied length field without checking it against the actual payload size, leaking up to 64KB of adjacent heap memory, including private keys, per request. CVE-2021-3156 (“Baron Samedit”) was a heap-based buffer overflow in sudo’s command-line parsing of escaped characters, allowing local privilege escalation to root. Both were memory-safety bugs in mature, widely audited C codebases, which is a large part of why the industry has pushed toward memory-safe languages for new security-critical code.
Real-World Case Study
CVE-2021-3156 (“Baron Samedit”), disclosed by Qualys in 2021, is a heap-based buffer overflow in sudo that had gone unnoticed for close to a decade. The bug lived in how sudo parsed command-line arguments containing escaped characters when invoked in a specific mode (via sudoedit-style argument parsing). A crafted argument caused sudo’s internal length calculation for an unescaped buffer to be miscounted, so the code that copied the argument into a heap buffer wrote past its allocated end. Because the vulnerable parsing ran before sudo’s normal permission checks completed, any local, unprivileged user could trigger it, turning a heap corruption bug into root-level privilege escalation on effectively every major Linux distribution shipping the affected sudo versions. It illustrates two recurring buffer overflow lessons: the bug sat in decades-old, widely audited C code, and the most damaging overflows are ones reachable before any authorization gate, not ones requiring existing privileges to exploit.
Real-World Mitigation Stack
A modern Linux binary compiled with hardening flags layers several defenses at once, so a single bug rarely leads directly to code execution:
- Compiled with
-fstack-protector-strongfor canaries - Linked as a Position Independent Executable (PIE) so ASLR covers the binary itself, not just libraries
- Built with
-D_FORTIFY_SOURCE=2so the compiler substitutes bounds-checked variants of risky libc calls where the buffer size is known at compile time - Marked
NXin the ELF header so the loader maps stack and heap as non-executable - Optionally run under a fuzzer (AFL++, libFuzzer) in CI to catch overflow-triggering inputs before release
History
- 1988: the Morris Worm exploited a buffer overflow in the Unix
fingerddaemon, one of the first widely documented uses of the technique in the wild. - 1996: Aleph One’s essay “Smashing the Stack for Fun and Profit” formalized stack-smashing technique for a mainstream security audience and became the canonical reference.
- Late 1990s–2000s: stack canaries (StackGuard, ProPolice) and non-executable stack/heap protections (PaX, then DEP) were adopted as compiler and OS defaults.
- Mid-2000s: ASLR shipped broadly across major operating systems, forcing attackers to combine overflow bugs with separate information disclosure bugs.
- ~2007 onward: return-oriented programming was formalized as the standard bypass for non-executable memory, followed by more general code-reuse techniques (jump-oriented programming, counterfeit object-oriented programming).
- 2010s–present: hardware-assisted mitigations (Intel CET shadow stacks, ARM Pointer Authentication) and Control Flow Integrity narrow the attack surface further, while major vendors (Google, Microsoft) publicly commit to memory-safe languages like Rust for new systems code specifically to eliminate this bug class at the source.
Detection and Tools
- Static analysis: tools like Coverity, Clang Static Analyzer, and
cppcheckflag unbounded copies and suspicious pointer arithmetic before code ships. - Dynamic instrumentation: AddressSanitizer (ASan) instruments memory accesses at compile time and aborts immediately on an out-of-bounds read or write, far more precisely than waiting for a crash.
- Fuzzing: AFL++ and libFuzzer generate malformed inputs automatically and feed them to a target repeatedly, looking for crashes; most modern buffer overflow CVEs in open-source C/C++ projects are now found this way rather than by manual review.
- Runtime mitigations: even without finding the bug, canaries, ASLR, DEP/NX, and CFI (Control Flow Integrity) reduce how far a discovered bug can be pushed toward reliable exploitation.
FAQ
Does a memory-safe language make an application immune to buffer overflows? For safe-mode code, effectively yes: Rust’s borrow checker and bounds-checked indexing turn what would be a silent overflow into a compile error or a runtime panic. unsafe blocks in Rust, or JNI/native calls from Java and Go, reopen the same C-level risk.
Why does ASLR need to be combined with an info leak to be bypassed? Because ASLR only hides addresses, it doesn’t remove the underlying bug. A separate vulnerability that reveals one pointer value (a stack address, a libc address) lets the attacker compute every other randomized base by fixed offset, restoring a working exploit.
Is a heap overflow generally easier or harder to exploit than a stack overflow? Harder in modern allocators. glibc and other hardened allocators added integrity checks (like Safe-Linking) specifically to make the classic unlink-based arbitrary-write technique unreliable, forcing attackers toward more allocator-specific exploitation techniques.
Does DEP/NX alone stop code injection attacks? It stops naive shellcode injection, but not ROP, which reuses existing executable code instead of injecting new code, so DEP/NX is a mitigation that raises the bar, not a complete fix on its own.
Do Windows and Linux use the same mitigation names? The concepts line up but the names differ. Windows’ DEP corresponds to Linux’s NX bit, Windows’ ASLR is functionally the same idea on both, and Windows adds SafeSEH and Control Flow Guard (CFG) to protect exception handler chains and indirect call targets, roughly analogous to Linux’s stack canaries and CFI schemes but implemented at the OS/compiler level differently.
Related Terms
Referenced by