Instruction Set Architecture
Instruction Set Architecture
Definition: The abstract interface contract between software and CPU hardware, defining the available instructions, addressing modes, register set, and data types a processor implements (e.g. x86-64, ARM64, RISC-V).
How It Works
- The ISA is a specification, not a physical implementation; many different physical chips (an Intel Core i9 and an AMD Ryzen) can implement the same ISA (x86-64) while using completely different internal microarchitectures
- RISC (Reduced Instruction Set Computer): simple, typically fixed-length instructions, each doing one small thing, designed to execute in close to one clock cycle each and to be easy to pipeline (ARM, RISC-V, MIPS)
- CISC (Complex Instruction Set Computer): variable-length instructions, some capable of complex multi-step operations (a single x86 instruction can do a memory load, an arithmetic operation, and a store), reflecting an era when instruction density and compiler simplicity mattered more than pipelining (x86-64)
- The microcode layer inside a CISC CPU translates complex ISA instructions into simpler internal micro-ops that the actual execution hardware runs, meaning a modern x86-64 chip is internally more RISC-like than the ISA it exposes suggests
- Addressing modes define how an instruction specifies an operand’s location: immediate (a literal constant), register direct, register indirect (a memory address held in a register), and various indexed/scaled/displaced forms (
[base + index*scale + displacement]) common in CISC ISAs - The ISA also fixes the ABI-adjacent basics: register count and width, integer/float data types and sizes, endianness, and how conditions/flags work, all of which a compiler’s backend must target exactly for generated code to run correctly
- ISA extensions add optional instruction subsets on top of a base ISA without breaking compatibility, for example x86-64’s SSE/AVX for SIMD or ARM’s NEON, detected at runtime via CPU feature flags so software can use them only when present
- The load-store architecture principle, central to RISC design, means only dedicated load/store instructions touch memory; every arithmetic/logic instruction operates purely on registers, which keeps the pipeline’s Execute stage simple and predictable, see CPU Pipelining
- Instruction encoding is the actual bit layout of an instruction: which bits are the opcode, which are register fields, which are immediates; RISC-V’s fixed 32-bit encoding is famous for keeping this deliberately simple and extensible, with reserved bit patterns for future extensions
- An ISA also defines its privilege levels (user mode vs kernel/supervisor mode on x86-64’s rings, or RISC-V’s machine/supervisor/user modes), which is the hardware foundation operating systems build memory protection and system calls on top of
RISC vs CISC: Why the Split Happened
- In the 1970s-80s, memory was expensive and slow relative to the CPU, so CISC ISAs packed more work into each instruction to reduce total instructions fetched from memory, and compilers of that era weren’t sophisticated enough to exploit simple instructions well
- As memory and cache got faster and compilers got much better at optimization, the calculus flipped: uniform, simple, pipeline-friendly instructions became more valuable than instruction-count density, favoring RISC
- x86 survived by pivoting its outward-facing ISA into microcode-decoded RISC-like internals while keeping the CISC instruction encoding for backward compatibility with decades of existing software, a compromise rather than a pure design choice
Under the Hood
Given: compiling c = a + b[i] for a RISC target (ARM64) versus a CISC target (x86-64).
Step 1, RISC: the compiler must emit a separate load instruction to read b[i] into a register, a separate add, and a separate store, since ARM64 ALU instructions only operate on registers, never directly on memory.
Step 2, CISC: x86-64 permits a single ADD instruction with a memory operand, so the compiler can fold the load directly into the arithmetic instruction’s encoding.
Answer: RISC code is typically more instructions but each is uniform and pipeline-friendly; CISC code is fewer, denser instructions, but the CPU’s decoder does more work per instruction and the instruction itself takes internally more cycles/micro-ops to execute.
Given: a RISC-V core needs to load the 32-bit constant 0x12345678 into a register, but RISC-V’s fixed instruction width (32 bits) can’t fit a full 32-bit immediate plus an opcode in one instruction.
Step: the compiler emits two instructions, lui (load upper immediate, sets the high 20 bits) followed by addi (add immediate, sets the low 12 bits).
Answer: what would be one CISC instruction (MOV reg, 0x12345678 on x86-64) becomes two RISC instructions, the direct tradeoff for RISC-V’s uniform, easy-to-decode 32-bit instruction format.
Given: a decoder needs to figure out where one instruction ends and the next begins. Step 1, RISC-V: every instruction is exactly 32 bits (ignoring the optional compressed 16-bit extension), so the decoder always knows the next instruction starts 4 bytes later, with no lookahead needed. Step 2, x86-64: an instruction can be 1 to 15 bytes, with prefixes, opcode bytes, a ModRM byte, optional SIB byte, and optional displacement/immediate bytes, so the decoder must partially parse an instruction just to know its length before it can even start decoding the next one. Answer: this is why x86-64 decoders are a genuinely harder, more power-hungry piece of silicon than a RISC-V decoder, and why superscalar x86-64 chips dedicate significant die area and engineering effort just to decoding multiple variable-length instructions per cycle.
x86-64 encoding sketch for one ADD instruction:
[prefix?] [REX?] [opcode] [ModRM] [SIB?] [displacement?] [immediate?]
0-4 1 1-2 0-1 0-1 0,1,2,4 0,1,2,4 bytes
RISC-V encoding for any instruction:
[opcode: 7 bits][rest of the 32-bit word, fixed layout by instruction type]
Why It Matters
- Determines binary compatibility: an x86-64 binary cannot run natively on ARM64 hardware, and vice versa, without emulation/translation, which is the whole reason Apple’s Rosetta 2 and Microsoft’s Windows-on-ARM emulation layers exist
- Shapes compiler backend design directly: instruction selection has to target the specific instructions, addressing modes, and register set the ISA defines, described in Compiler Pipeline Architecture
- Drives power efficiency and die area tradeoffs at the hardware level: RISC’s simpler decode logic is a large part of why ARM-based mobile and Apple Silicon chips achieve strong performance-per-watt versus traditional x86-64 designs
Common Pitfalls
- Assuming x86 assembly or binaries can run natively on ARM hardware without binary translation or emulation; they cannot, the encodings and instruction semantics are entirely different
- Treating “RISC vs CISC” as a clean, absolute binary today: modern x86-64 decodes into RISC-like micro-ops internally, and modern ARM has accumulated some more complex instructions over decades, so the practical distinction is smaller at the microarchitecture level than the ISA-level story suggests
- Writing code that assumes a specific ISA’s endianness or integer width without checking; x86-64 is little-endian, some ARM configurations are switchable, and porting code that hardcodes byte order breaks silently rather than loudly
- Forgetting to runtime-check for ISA extensions (AVX2, AVX-512) before using them; code compiled assuming an extension is present will crash with an illegal-instruction fault on older or different hardware that lacks it
- Confusing ISA with microarchitecture in a performance discussion; two chips implementing the identical ISA (two different x86-64 CPUs) can have wildly different real-world performance because of microarchitectural differences, cache sizes, pipeline depth, branch predictor quality, that the ISA itself says nothing about
Comparison
| RISC | CISC | |
|---|---|---|
| Instruction length | Fixed (e.g. 32-bit) | Variable (1-15 bytes on x86-64) |
| Instructions per task | More | Fewer |
| Decode complexity | Simple | Complex, needs microcode |
| Memory operands in ALU ops | Not allowed, load/store separate | Allowed directly |
| Examples | ARM, RISC-V, MIPS | x86-64 |
| Typical strength | Power efficiency, simple pipelining | Code density, legacy compatibility |
| ISA | Style | License | Common in |
|---|---|---|---|
| x86-64 | CISC (RISC-like microcode internally) | Proprietary (Intel/AMD) | Desktops, servers, gaming |
| ARM64 (AArch64) | RISC | Licensed (Arm Holdings) | Mobile, Apple Silicon, growing server share |
| RISC-V | RISC | Open, royalty-free | Embedded, academia, emerging server/laptop chips |
Example
Apple Silicon M-series processors implement the ARM64 (AArch64) RISC architecture, which is a major reason they achieve high performance-per-watt versus comparable Intel/AMD x86-64 (CISC) laptop chips. RISC-V is a newer, open, royalty-free RISC ISA gaining adoption in embedded systems and increasingly in server and even laptop-class chips specifically because anyone can implement it without ISA licensing fees, unlike x86-64 (licensed by Intel/AMD) or ARM (licensed by Arm Holdings).
Apple’s Rosetta 2 is a concrete example of the compatibility problem an ISA boundary creates: when Apple moved Macs from x86-64 to ARM64, existing x86-64 applications couldn’t run natively, so Rosetta 2 translates x86-64 binaries into ARM64 instructions, mostly ahead-of-time at install time, with a JIT fallback for dynamically generated code, trading some performance for compatibility during the transition period.
Given: a compiler team needs to decide whether to add a new fused multiply-add instruction to their custom ISA. Step 1: adding it as a single instruction reduces instruction count and can improve floating-point precision (one rounding step instead of two), but it complicates every implementation that has to support the ISA, since now the ALU needs a genuine three-operand fused unit rather than reusing separate multiply and add hardware. Step 2: leaving it out keeps the ISA and hardware simpler, at the cost of compilers having to emit two instructions and accept the extra rounding step, for every multiply-accumulate operation across every program that ever runs on that ISA. Answer: this is the essential RISC-vs-CISC tension in miniature: every instruction added to an ISA is a permanent commitment every future implementation must honor, which is why ISA design changes so much more slowly and carefully than microarchitecture design.
Related Terms
Referenced by