CPU Instruction Cycle

CPU Instruction Cycle

Definition: The continuous operational sequence a CPU core executes to process one machine instruction: Fetch, Decode, Execute, Memory Access, Writeback.

How It Works

  • 1. Fetch: the control unit reads the instruction byte(s) from the memory address held in the Program Counter (PC), using the instruction/address/data buses
  • 2. Decode: the control unit parses the opcode, identifies the instruction type, and extracts source and destination register/immediate operands from the instruction encoding
  • 3. Execute: the ALU performs the arithmetic/logic operation, or evaluates a branch condition and computes the target address for a jump
  • 4. Memory Access: if the instruction reads or writes memory (a load or store), the memory stage performs that access; instructions with no memory operand pass through this stage doing nothing
  • 5. Writeback: the final result is written into the destination register, and the Program Counter is updated, either incremented to the next sequential instruction or set to a branch target
  • This is the classic five-stage RISC model (as in the MIPS and early ARM pipelines); real CISC cores like x86-64 further split decode into multiple sub-stages, since a single x86 instruction can be 1 to 15 bytes and needs microcode translation into simpler internal micro-ops first
  • The cycle repeats indefinitely at the CPU’s clock rate; a non-pipelined implementation completes one full cycle before starting the next instruction’s fetch, while a pipelined implementation overlaps stages of consecutive instructions (see CPU Pipelining)
  • Interrupts and exceptions are checked at a defined point in the cycle, typically after an instruction completes writeback, so the CPU can service them between instructions without corrupting the in-progress one

Under the Hood

Given: the instruction ADD RAX, RBX (add RBX into RAX) sitting at address 0x1000, with PC currently pointing there. Step 1, Fetch: the control unit reads the instruction bytes at 0x1000 from memory into the instruction register. Step 2, Decode: the control unit recognizes the ADD opcode and identifies RAX as destination/source-1 and RBX as source-2. Step 3, Execute: the ALU adds the current values of RAX and RBX. Step 4, Memory: skipped, since this instruction has no memory operand. Step 5, Writeback: the sum is written into RAX, and PC advances to 0x1004 (the next instruction). Answer: five stages, one memory access (the initial fetch), and one register write, to execute a single add.

Given: the instruction MOV RAX, [RBX] (load the value at the address in RBX into RAX). Step: Fetch reads the instruction itself; Decode identifies it as a load; Execute computes the effective address (just RBX’s value here, but could include an offset or scaled index); Memory reads 8 bytes from that address; Writeback stores the loaded value into RAX. Answer: unlike the pure ADD example, this instruction actually uses the Memory stage, which is why load/store instructions typically have higher latency than register-only ALU instructions.

Given: a conditional branch JZ target (jump if zero flag set) follows a comparison instruction. Step 1, Fetch: the branch instruction’s bytes are read from memory as normal. Step 2, Decode: the control unit recognizes this as a conditional branch and identifies the target address encoding. Step 3, Execute: the branch unit checks the zero flag (set by the earlier comparison) and computes whether to take the branch. Step 4, Memory: skipped, branches have no memory operand. Step 5, Writeback: instead of writing a register, this stage updates the Program Counter, either to the branch target (if taken) or to the next sequential instruction (if not taken). Answer: a branch’s “writeback” is a PC update rather than a register write, which is exactly the value a pipelined CPU doesn’t know for certain until Execute completes, forcing a stall or a guess (branch prediction) in the stages before it.

cycle:        1     2     3     4     5
ADD RAX,RBX:  F  -> D  -> E  -> M  -> W
              (Memory stage does nothing; Execute produces the sum)

Interrupt Handling in the Cycle

  • Between instructions, typically right after Writeback, the control unit checks a pending-interrupt line
  • If an interrupt (timer, keyboard, disk I/O completion) is pending and not masked, the CPU saves the current Program Counter and processor status, then jumps to an Interrupt Service Routine (ISR) address looked up from an interrupt vector table
  • After the ISR finishes (typically via an IRET/ERET-style instruction), the saved PC and status are restored, and the instruction cycle resumes exactly where it left off
  • This check-after-writeback timing is what guarantees interrupts never corrupt a partially executed instruction: the interrupted instruction has always fully completed its own cycle first

Why It Matters

  • Represents the fundamental clock-driven loop underlying all CPU execution, from a microcontroller to a server-class processor; every higher-level abstraction (instructions, programs, operating systems) ultimately reduces to this cycle running billions of times per second
  • Understanding which stage does what explains real performance behavior: memory-bound code spends disproportionate time in the Memory stage waiting on cache/RAM, while ALU-bound code is limited by Execute-stage throughput
  • The cycle’s Fetch-Decode dependency on the Program Counter is exactly what a correctly or incorrectly predicted branch speeds up or slows down; see CPU Pipelining for how pipelining turns a mispredicted PC update into a costly flush

Common Pitfalls

  • Unoptimized memory access during the Fetch or Memory stages causes clock cycle stalls; an L1 cache miss during Fetch can stall the entire cycle for the instruction that needed it
  • Assuming every instruction uses every stage equally: register-only ALU instructions skip meaningful work in the Memory stage, while loads/stores depend heavily on it, so cycle counts per instruction type vary in practice even in a non-pipelined design
  • Treating the five-stage model as universally accurate: real CPUs vary stage count and granularity heavily, and CISC decode in particular is far more complex than a single textbook “Decode” box suggests
  • Forgetting that Writeback’s Program Counter update is itself a dependency: the next Fetch cannot begin correctly until this cycle’s PC update completes, which is exactly the sequential dependency pipelining is designed to hide
  • Assuming clock frequency alone predicts performance: two CPUs at the same clock speed can have very different real throughput if one stalls far more often in Fetch/Memory waiting on cache misses
  • Ignoring microcode: some CISC instructions (x86’s string operations, or complex addressing modes) decode into many internal micro-ops, meaning a single “instruction” in source code can represent many actual passes through Execute internally

Comparison

StageUsesSkipped when
FetchInstruction/address bus, PCNever
DecodeControl unit, opcode tables/microcodeNever
ExecuteALU or branch unitNever (even a no-op computes something trivial)
MemoryData bus, cache/RAMInstruction has no memory operand
WritebackRegister fileInstruction produces no result (e.g. an unconditional jump)
DesignInstructions in flightThroughput
Non-pipelined (classic cycle)11 instruction per 5 cycles (5-stage model)
PipelinedUp to number of stagesApproaches 1 instruction per cycle once filled
Superscalar pipelinedMultiple per stageMore than 1 instruction per cycle
Instruction typeFetchDecodeExecuteMemoryWriteback
Register ALU op (ADD)YesYesComputes resultIdleWrites register
Load (MOV reg, [mem])YesYesComputes addressReads memoryWrites register
Store (MOV [mem], reg)YesYesComputes addressWrites memoryIdle
Unconditional jumpYesYesComputes targetIdleUpdates PC only

Example

Executing ADD RAX, RBX on an x86-64 core performs Fetch, Decode, ALU addition (Execute), no Memory-stage work, and Writeback to RAX. A simple 8-bit microcontroller like the classic AVR (used in early Arduino boards) implements this exact five-stage sequence close to literally, with datasheet-documented cycle counts per instruction, which is why AVR assembly programmers can predict exact timing by counting instructions, something essentially impossible on a modern out-of-order x86-64 core.

Given: an embedded systems engineer needs a hardware timing loop that takes exactly 100 microseconds on an AVR running at 16 MHz, with no hardware timer available. Step 1: at 16 MHz, one clock cycle is 62.5 nanoseconds, and the datasheet states most single-word AVR instructions complete their instruction cycle in exactly 1 clock cycle (some, like multiply, take 2). Step 2: the engineer counts instructions in a small delay loop, multiplies by their known cycle counts, and multiplies by 62.5 ns per cycle to get total elapsed time. Answer: 100 microseconds needs exactly 1,600 cycles’ worth of instructions, a calculation only possible because the instruction cycle’s timing is fixed and documented, unlike a modern out-of-order superscalar core where the same instruction sequence’s timing varies with cache state, branch prediction, and neighboring instructions.

Dig deeper