CPU Core and Registers
CPU Core and Registers
Definition: A CPU core is an independent processing unit containing an ALU, a Control Unit, and a small set of high-speed internal Registers capable of fetching and executing instructions.
How It Works
- Registers: the smallest, fastest storage elements directly inside the CPU, holding operands and results for the instruction currently executing (e.g. Program Counter PC, Stack Pointer SP, general-purpose registers)
- ALU (Arithmetic Logic Unit): performs integer arithmetic (add, subtract, multiply) and bitwise logic operations (AND, OR, XOR, shifts) on register values
- Control Unit (CU): fetches instructions from memory, decodes their opcodes, and generates the control signals that direct data through the ALU, registers, and buses
- Floating-Point Unit (FPU): a separate execution unit for floating-point arithmetic, since IEEE 754 operations need different hardware than integer ALUs
- Register File: the physical block of storage holding all general-purpose registers, built from extremely fast SRAM cells directly wired to the ALU’s inputs and outputs, so an operand read takes a fraction of a clock cycle
- Registers are named in the ISA (architectural registers, e.g.
%raxon x86-64) but modern CPUs implement far more physical registers than the ISA exposes, using register renaming to map architectural names onto a larger physical pool and avoid false dependencies between unrelated instructions reusing the same name - Special-purpose registers include the Program Counter (address of the next instruction), Stack Pointer (top of the current call stack), Status/Flags register (carry, zero, overflow, sign flags set by the last ALU operation), and often a Link Register (return address, used instead of a stack push/pop on some RISC ISAs like ARM)
- Vector/SIMD registers (XMM/YMM/ZMM on x86-64, NEON
vregisters on ARM) are wide registers, 128 to 512 bits, holding multiple packed values processed by a single instruction; see SIMD and Vector Processing - A modern core is really several cooperating units: a front end (fetch, decode, branch prediction), a back end of multiple execution ports (integer ALUs, FPU, load/store units) that can run in parallel, and a retirement unit that commits results in original program order even though execution itself may have happened out of order
- Out-of-order execution lets independent instructions execute as soon as their operands are ready rather than strictly in program order, using a reorder buffer (ROB) to track in-flight instructions and commit their results in the original order so the programmer-visible state stays correct
- The register file typically has multiple read/write ports, letting several instructions read or write different registers in the same clock cycle, which is a direct hardware prerequisite for issuing more than one instruction per cycle (superscalar execution)
Under the Hood
Given: the instruction ADD %rax, %rbx on x86-64, adding %rbx into %rax.
Step 1: the control unit fetches the instruction bytes from memory at the address in the Program Counter, then decodes the opcode and identifies %rax and %rbx as source/destination operands.
Step 2: the register file reads the current values of %rax and %rbx and presents them to the ALU’s input ports.
Step 3: the ALU computes the sum and sets the Zero, Carry, and Overflow flags based on the result.
Step 4: the result is written back into %rax in the register file, and the Program Counter advances to the next instruction.
Answer: one instruction touches the control unit, register file, ALU, and flags register in sequence, all without a single access to main memory.
Given: a CPU with 16 architectural integer registers but 180 physical registers internally.
Step: two unrelated instructions both writing to %rax (a write-after-write hazard on the architectural name) get mapped to two different physical registers by the renamer.
Answer: the second write no longer has to wait for the first to retire, letting both execute out of order, which is exactly what register renaming buys beyond what the ISA’s named registers alone would allow.
Given: a function body needs 20 live values at once, but x86-64 exposes only 16 general-purpose architectural registers. Step 1: the compiler’s register allocator picks the 16 most frequently used values to keep in registers. Step 2: the remaining values are spilled, stored to stack memory and reloaded whenever needed.
; compiler-generated spill code for one extra live value
mov QWORD PTR [rbp-8], r15 ; spill r15 to stack
... ; r15 is now free for reuse
mov r15, QWORD PTR [rbp-8] ; reload r15 when needed again
Answer: each spill/reload pair costs a memory access (a handful of cycles from L1, far more on an L1 miss) that a pure register access would never pay, which is exactly why compilers treat register pressure as a real optimization target, not just a bookkeeping detail.
Why It Matters
- Registers provide sub-nanosecond operand access, several times faster than even L1 cache, which is why compilers work hard to keep frequently used values in registers rather than memory
- The ALU and control unit’s design directly determines instructions-per-cycle (IPC) and clock frequency limits, since every instruction ultimately funnels through them
- Register count and width (32-bit vs 64-bit) shape what a compiler’s register allocator can do; more architectural registers mean fewer spills to the stack for a given piece of code
Common Pitfalls
- Limited architectural register count forces register spilling, saving active values to the stack/RAM and reloading them later, which costs memory-latency cycles a register access would never pay
- Assuming register renaming means architectural register count doesn’t matter: it helps with hazards between independent instructions, but a compiler still can’t address more values at once than the ISA exposes, so spills still happen under register pressure
- Confusing the ALU with the FPU: integer overflow and floating-point exceptions are detected and handled by entirely separate hardware paths with different flag semantics
- Forgetting that flags are a shared, mutable side effect: an instruction that doesn’t need to set flags but does anyway can create a false dependency that stalls out-of-order execution on some microarchitectures
Comparison
| Storage | Typical latency | Typical size per core | Managed by |
|---|---|---|---|
| Registers | < 1 cycle | Tens to low hundreds of bytes | Compiler (architectural), hardware (renaming) |
| L1 cache | ~1 ns | ~32-64 KB | Hardware, transparent to software |
| Main memory (RAM) | ~50-100 ns | GBs | OS virtual memory, hardware |
| ISA | General-purpose integer registers | Register width | Naming example |
|---|---|---|---|
| x86-64 | 16 | 64-bit | %rax, %rbx, %rsp |
| ARM64 (AArch64) | 31 | 64-bit | x0-x30 |
| RISC-V (RV64) | 32 (x0 hardwired to zero) | 64-bit | x0-x31, ABI names like a0, t0, ra |
| Unit | Job |
|---|---|
| Control Unit | Fetch, decode, sequence instructions |
| ALU | Integer arithmetic and logic |
| FPU | Floating-point arithmetic |
| Register file | Fast operand storage, read/written every instruction |
| Reorder buffer | Tracks in-flight out-of-order instructions, commits in program order |
Example
A 64-bit x86-64 core exposes architectural registers like %rax, %rbx, %rsp (stack pointer), and %rip (instruction pointer), but internally an Intel or AMD core renames these onto a physical register file several times larger to support out-of-order execution. Modern Intel/AMD cores (e.g. AMD Zen 4) implement several hundred physical integer registers behind just 16 architectural names.
ARM64 exposes 31 general-purpose registers (x0-x30) plus a dedicated stack pointer and program counter, a noticeably larger architectural register set than x86-64’s 16, which is one reason ARM64 code needs fewer stack spills for the same workload and one reason Apple Silicon (M-series) chips, built on ARM64, can sustain high instructions-per-cycle with a comparatively simple compiler-facing register model.
RISC-V takes transparency further: x0 is hardwired to the constant zero at the ISA level, so instructions that need a zero operand (or need to discard a result) simply reference x0 instead of needing a dedicated “compare and branch on zero” instruction form.
Given: a debugger needs to show a programmer the value of a local variable that the compiler decided to keep entirely in a register, never spilling it to the stack. Step 1: the compiler emits debug metadata (DWARF on Linux, PDB on Windows) mapping that source variable to a specific register at each program point, since the variable has no fixed memory address to point to. Step 2: when the programmer inspects that variable in the debugger, the debugger reads the live register value directly rather than dereferencing a stack address. Answer: this is exactly why optimized builds can make debuggers report a variable as “optimized out” at certain points, the register holding it got reused for something else once the compiler proved the old value was no longer needed.
Related Terms
Referenced by