CPU Instruction Cycle
CPU Instruction Cycle
Definition: The continuous operational sequence a CPU core executes to process one machine instruction: Fetch, Decode, Execute, Memory Access, Writeback.
How It Works
- 1. Fetch: the control unit reads the instruction byte(s) from the memory address held in the Program Counter (PC), using the instruction/address/data buses
- 2. Decode: the control unit parses the opcode, identifies the instruction type, and extracts source and destination register/immediate operands from the instruction encoding
- 3. Execute: the ALU performs the arithmetic/logic operation, or evaluates a branch condition and computes the target address for a jump
- 4. Memory Access: if the instruction reads or writes memory (a load or store), the memory stage performs that access; instructions with no memory operand pass through this stage doing nothing
- 5. Writeback: the final result is written into the destination register, and the Program Counter is updated, either incremented to the next sequential instruction or set to a branch target
- This is the classic five-stage RISC model (as in the MIPS and early ARM pipelines); real CISC cores like x86-64 further split decode into multiple sub-stages, since a single x86 instruction can be 1 to 15 bytes and needs microcode translation into simpler internal micro-ops first
- The cycle repeats indefinitely at the CPU’s clock rate; a non-pipelined implementation completes one full cycle before starting the next instruction’s fetch, while a pipelined implementation overlaps stages of consecutive instructions (see CPU Pipelining)
- Interrupts and exceptions are checked at a defined point in the cycle, typically after an instruction completes writeback, so the CPU can service them between instructions without corrupting the in-progress one
Under the Hood
Given: the instruction ADD RAX, RBX (add RBX into RAX) sitting at address 0x1000, with PC currently pointing there.
Step 1, Fetch: the control unit reads the instruction bytes at 0x1000 from memory into the instruction register.
Step 2, Decode: the control unit recognizes the ADD opcode and identifies RAX as destination/source-1 and RBX as source-2.
Step 3, Execute: the ALU adds the current values of RAX and RBX.
Step 4, Memory: skipped, since this instruction has no memory operand.
Step 5, Writeback: the sum is written into RAX, and PC advances to 0x1004 (the next instruction).
Answer: five stages, one memory access (the initial fetch), and one register write, to execute a single add.
Given: the instruction MOV RAX, [RBX] (load the value at the address in RBX into RAX).
Step: Fetch reads the instruction itself; Decode identifies it as a load; Execute computes the effective address (just RBX’s value here, but could include an offset or scaled index); Memory reads 8 bytes from that address; Writeback stores the loaded value into RAX.
Answer: unlike the pure ADD example, this instruction actually uses the Memory stage, which is why load/store instructions typically have higher latency than register-only ALU instructions.
Given: a conditional branch JZ target (jump if zero flag set) follows a comparison instruction.
Step 1, Fetch: the branch instruction’s bytes are read from memory as normal.
Step 2, Decode: the control unit recognizes this as a conditional branch and identifies the target address encoding.
Step 3, Execute: the branch unit checks the zero flag (set by the earlier comparison) and computes whether to take the branch.
Step 4, Memory: skipped, branches have no memory operand.
Step 5, Writeback: instead of writing a register, this stage updates the Program Counter, either to the branch target (if taken) or to the next sequential instruction (if not taken).
Answer: a branch’s “writeback” is a PC update rather than a register write, which is exactly the value a pipelined CPU doesn’t know for certain until Execute completes, forcing a stall or a guess (branch prediction) in the stages before it.
cycle: 1 2 3 4 5
ADD RAX,RBX: F -> D -> E -> M -> W
(Memory stage does nothing; Execute produces the sum)
Interrupt Handling in the Cycle
- Between instructions, typically right after Writeback, the control unit checks a pending-interrupt line
- If an interrupt (timer, keyboard, disk I/O completion) is pending and not masked, the CPU saves the current Program Counter and processor status, then jumps to an Interrupt Service Routine (ISR) address looked up from an interrupt vector table
- After the ISR finishes (typically via an
IRET/ERET-style instruction), the saved PC and status are restored, and the instruction cycle resumes exactly where it left off - This check-after-writeback timing is what guarantees interrupts never corrupt a partially executed instruction: the interrupted instruction has always fully completed its own cycle first
Why It Matters
- Represents the fundamental clock-driven loop underlying all CPU execution, from a microcontroller to a server-class processor; every higher-level abstraction (instructions, programs, operating systems) ultimately reduces to this cycle running billions of times per second
- Understanding which stage does what explains real performance behavior: memory-bound code spends disproportionate time in the Memory stage waiting on cache/RAM, while ALU-bound code is limited by Execute-stage throughput
- The cycle’s Fetch-Decode dependency on the Program Counter is exactly what a correctly or incorrectly predicted branch speeds up or slows down; see CPU Pipelining for how pipelining turns a mispredicted PC update into a costly flush
Common Pitfalls
- Unoptimized memory access during the Fetch or Memory stages causes clock cycle stalls; an L1 cache miss during Fetch can stall the entire cycle for the instruction that needed it
- Assuming every instruction uses every stage equally: register-only ALU instructions skip meaningful work in the Memory stage, while loads/stores depend heavily on it, so cycle counts per instruction type vary in practice even in a non-pipelined design
- Treating the five-stage model as universally accurate: real CPUs vary stage count and granularity heavily, and CISC decode in particular is far more complex than a single textbook “Decode” box suggests
- Forgetting that Writeback’s Program Counter update is itself a dependency: the next Fetch cannot begin correctly until this cycle’s PC update completes, which is exactly the sequential dependency pipelining is designed to hide
- Assuming clock frequency alone predicts performance: two CPUs at the same clock speed can have very different real throughput if one stalls far more often in Fetch/Memory waiting on cache misses
- Ignoring microcode: some CISC instructions (x86’s string operations, or complex addressing modes) decode into many internal micro-ops, meaning a single “instruction” in source code can represent many actual passes through Execute internally
Comparison
| Stage | Uses | Skipped when |
|---|---|---|
| Fetch | Instruction/address bus, PC | Never |
| Decode | Control unit, opcode tables/microcode | Never |
| Execute | ALU or branch unit | Never (even a no-op computes something trivial) |
| Memory | Data bus, cache/RAM | Instruction has no memory operand |
| Writeback | Register file | Instruction produces no result (e.g. an unconditional jump) |
| Design | Instructions in flight | Throughput |
|---|---|---|
| Non-pipelined (classic cycle) | 1 | 1 instruction per 5 cycles (5-stage model) |
| Pipelined | Up to number of stages | Approaches 1 instruction per cycle once filled |
| Superscalar pipelined | Multiple per stage | More than 1 instruction per cycle |
| Instruction type | Fetch | Decode | Execute | Memory | Writeback |
|---|---|---|---|---|---|
Register ALU op (ADD) | Yes | Yes | Computes result | Idle | Writes register |
Load (MOV reg, [mem]) | Yes | Yes | Computes address | Reads memory | Writes register |
Store (MOV [mem], reg) | Yes | Yes | Computes address | Writes memory | Idle |
| Unconditional jump | Yes | Yes | Computes target | Idle | Updates PC only |
Example
Executing ADD RAX, RBX on an x86-64 core performs Fetch, Decode, ALU addition (Execute), no Memory-stage work, and Writeback to RAX. A simple 8-bit microcontroller like the classic AVR (used in early Arduino boards) implements this exact five-stage sequence close to literally, with datasheet-documented cycle counts per instruction, which is why AVR assembly programmers can predict exact timing by counting instructions, something essentially impossible on a modern out-of-order x86-64 core.
Given: an embedded systems engineer needs a hardware timing loop that takes exactly 100 microseconds on an AVR running at 16 MHz, with no hardware timer available. Step 1: at 16 MHz, one clock cycle is 62.5 nanoseconds, and the datasheet states most single-word AVR instructions complete their instruction cycle in exactly 1 clock cycle (some, like multiply, take 2). Step 2: the engineer counts instructions in a small delay loop, multiplies by their known cycle counts, and multiplies by 62.5 ns per cycle to get total elapsed time. Answer: 100 microseconds needs exactly 1,600 cycles’ worth of instructions, a calculation only possible because the instruction cycle’s timing is fixed and documented, unlike a modern out-of-order superscalar core where the same instruction sequence’s timing varies with cache state, branch prediction, and neighboring instructions.
Related Terms
Referenced by