JIT vs AOT Compilation
JIT vs AOT Compilation
Definition: Ahead-Of-Time (AOT) compiles source code directly into native machine binaries before execution; Just-In-Time (JIT) compiles bytecode into native machine code dynamically at runtime, while the program is already running.
How It Works
- AOT (C, C++, Rust, Go, Swift): the full program compiles to a static or dynamically-linked native binary before it ever runs, giving fast, predictable startup with no runtime compilation overhead
- JIT (JVM/HotSpot, V8 JavaScript, PyPy, .NET CLR): the runtime first interprets or baseline-compiles bytecode, then profiles execution to find “hot” methods and loops and recompiles just those into optimized native code, driven by live runtime data rather than a static training run
- Tiered Compilation: HotSpot uses C1 (client compiler, fast to compile, lightly optimized) for warm code and escalates the hottest methods to C2 (server compiler, slow to compile, heavily optimized); V8 layers Ignition (bytecode interpreter) -> Sparkplug (fast non-optimizing baseline JIT) -> Maglev/TurboFan (optimizing JIT)
- Speculative optimization and deoptimization: JITs compile assuming observed types or shapes hold, for example V8 assumes an object’s “hidden class” stays stable; if that assumption is later violated, the engine deoptimizes, discards the optimized code, and falls back to the interpreter or a slower tier
- Hybrid approaches blur the line: GraalVM’s
native-imageAOT-compiles JVM bytecode into a standalone binary; Android ART compiles DEX bytecode to native code at install/first-run time rather than every launch; .NET’s ReadyToRun pre-JITs common paths into the assembly while still allowing runtime re-JITing - Both ultimately run the same instruction-selection, register-allocation, and scheduling backend logic described in Compiler Pipeline Architecture; the difference is entirely about when that backend runs and what information it has access to
- On-Stack Replacement (OSR): lets a JIT swap a running interpreted method for compiled native code mid-execution, without waiting for the method to be called again, which matters for long-running loops that are already hot on their first call
- Inline caches: JITs attach a small cache at each call site recording the concrete type(s) seen there so far; a monomorphic site (always the same type) compiles to a direct call, a polymorphic site (a few types) compiles to a small type-checked dispatch table, and a megamorphic site (many types) falls back to a slow generic lookup
- AOT compilers can still use profile data if it’s supplied ahead of time: Profile-Guided Optimization (PGO) runs the program once on representative input, records which branches and functions are hot, and feeds that back into a second, offline compile, approximating some of what a JIT learns live
Warmup and Tiering in Practice
- A JIT’s total cost has two components: interpreter/baseline overhead paid before a method is recompiled, and compilation overhead paid to actually recompile it; tiering exists to balance these, since compiling every method with a heavy optimizer immediately would make startup unacceptably slow
- HotSpot’s default trigger is invocation counters: a method recompiles to C1 after roughly 1,500 invocations and to C2 after roughly 10,000 (tunable via
-XX:CompileThreshold), though loop back-edge counts can trigger OSR compilation even sooner for hot loops that only run once - V8 similarly promotes a function once it crosses an internal call-count and time-in-function threshold; functions that never get called enough simply stay in Sparkplug or even Ignition, which is correct, since compiling cold code with a heavy optimizer would waste time for no benefit
- Background compilation threads let tiering happen without pausing the running program: the interpreter or baseline tier keeps executing on the main thread while a separate compiler thread builds the optimized version, and the switch to optimized code happens atomically once it’s ready
Real Numbers
- A typical HotSpot JVM microbenchmark shows 10-50x throughput improvement between pure interpretation and fully C2-optimized code for numerically heavy loops, though this gap depends heavily on how much the JIT can prove about types and bounds
- Cold-start latency is the flip side: a JVM Lambda function can take 300-900ms just to initialize the JVM and JIT-warm the handler path, versus roughly 10-50ms for an AOT-compiled Go or Rust binary doing equivalent work, which is the core reason serverless platforms push developers toward AOT-friendly languages or GraalVM native-image for latency-sensitive functions
Under the Hood
Given: a Go program and an equivalent Node.js script both need to sum a large array.
Step 1, AOT: go build compiles the whole program to native code once; every run starts at full speed with no warm-up.
Step 2, JIT: Node starts by running the summing loop through V8’s Ignition interpreter; after a few thousand iterations, V8’s profiler marks the loop hot and TurboFan compiles it to native code mid-execution.
Answer: the Go binary runs at consistent speed from instruction one; the Node script starts slower and speeds up partway through, potentially reaching similar or better throughput if TurboFan’s runtime-informed optimization beats what a static compiler could prove safe ahead of time.
Given: a JavaScript function add(a, b) is called thousands of times with numbers, then once with a string.
Step 1: TurboFan speculatively compiles add assuming both arguments are numbers, based on observed call history.
Step 2: the call with a string violates that assumption.
Answer: V8 deoptimizes add, discarding the optimized native code and falling back to the interpreter or a slower tier for subsequent calls until the function’s behavior stabilizes again.
Given: a serverless function handling one HTTP request per cold-started container, using the JVM. Step 1: the container starts, the class loader loads bytecode, and HotSpot begins interpreting, since C1/C2 haven’t triggered yet. Step 2: the request completes and the container is torn down before the method ever ran enough times to justify JIT compilation. Answer: the JIT’s runtime specialization never pays off in this workload shape; the entire request runs on interpreted bytecode, which is why serverless JVM deployments often switch to an AOT-compiled runtime (GraalVM native-image) instead.
Why It Matters
- Governs runtime performance, startup latency, memory footprint, and, for serverless or CLI workloads, how quickly a process reaches full speed
- AOT wins where startup time and predictability matter most: CLI tools, serverless cold starts, embedded systems
- JIT wins for long-running processes where warm-up time is amortized and the compiler can specialize to the program’s actual observed runtime behavior in ways a static compiler never could, such as devirtualizing a call based on the one concrete type actually seen at that call site
- The choice shapes deployment strategy too: AOT binaries are self-contained and simple to ship; JIT runtimes need the JIT engine itself present on the target machine
- Memory footprint differs structurally: a JIT keeps both the interpreter/bytecode and the compiled native code for hot methods resident, plus profiling metadata per call site, while an AOT binary only carries the final machine code
- For interactive workloads (a web server handling requests for hours), the JIT’s specialization advantage compounds over the process lifetime, since the same optimized code path serves millions of requests once compiled
Common Pitfalls
- JIT warmup latency: initial execution suffers latency spikes and lower throughput while the engine still runs interpreted or baseline code, directly hurting short-lived processes like serverless functions
- Deoptimization storms: polymorphic call sites that keep violating the JIT’s type speculation, for example a JS function called with wildly different argument shapes, can repeatedly trigger deopt/reoptimize cycles, making code slower than if it had never been optimized
- Assuming AOT is always faster: AOT lacks runtime profile information, so it can’t perform speculative optimizations a JIT can, and heavy AOT optimization can bloat binary size and build times
- Benchmarking a JIT-compiled program for only a few seconds and concluding it’s slow, when the measurement window never let the JIT reach its optimized tier
- Forgetting that AOT binaries are tied to the CPU features chosen at compile time (e.g., whether AVX2 is assumed available), while a JIT can, in principle, generate code tuned to the exact CPU it’s actually running on
- Treating PGO as equivalent to a JIT: PGO’s profile is a snapshot from one training run and stays fixed once compiled, while a JIT keeps re-profiling and can adapt if the program’s actual behavior later diverges from what was first observed
- Ignoring inline cache degradation: code that starts monomorphic and drifts polymorphic or megamorphic over time (common in dynamically typed languages as codebases grow) silently loses JIT optimization without any visible error
Comparison
| AOT | JIT | |
|---|---|---|
| When compiled | Before execution, once | During execution, repeatedly as needed |
| Startup latency | Low, immediate full speed | Higher, warms up over time |
| Peak throughput | Fixed at compile time | Can exceed AOT via runtime specialization |
| Uses runtime profile data | No | Yes |
| Binary/deployment | Self-contained native binary | Needs the runtime/JIT engine present |
| Typical languages | C, C++, Rust, Go, Swift | Java, JavaScript, C#, Python (PyPy) |
| Tier | Compiles how fast | Optimization level | Trigger |
|---|---|---|---|
| Interpreter (Ignition, JVM bytecode interpreter) | Instant, no compile step | None | Always available, first execution |
| Baseline JIT (Sparkplug, C1) | Fast | Light | Method called a small number of times |
| Optimizing JIT (TurboFan, C2) | Slow | Heavy, speculative | Method identified as genuinely hot |
Example
The V8 engine compiles JavaScript through multiple tiers: Ignition interprets bytecode immediately for fast startup, Sparkplug compiles that bytecode nearly verbatim into machine code once a function runs a few times, and TurboFan later recompiles genuinely hot functions with full optimization, deoptimizing back to Ignition if a type assumption, like an object’s shape, turns out to be wrong.
Java’s HotSpot JVM follows a similar tiered path (interpreter -> C1 -> C2), while GraalVM’s native-image takes the opposite route for the same bytecode: it AOT-compiles the whole application ahead of time into a native executable, trading away JIT-style runtime specialization for near-instant startup, which is why it’s popular for serverless Java deployments.
.NET offers both in one runtime: the CLR JIT-compiles Intermediate Language (IL) to native code method-by-method at runtime by default, while crossgen/ReadyToRun and the newer Native AOT toolchain compile the whole assembly ahead of time, letting a team pick per-deployment whether they want JIT’s adaptability or AOT’s startup speed without switching languages.
Python’s reference interpreter, CPython, is a pure bytecode interpreter with no JIT at all, which is why CPU-bound Python code is often much slower than equivalent Java or JavaScript; PyPy exists specifically to add a tracing JIT on top of the same language, closing much of that gap for long-running programs.
Given: a company runs the same microservice two ways to compare, once on OpenJDK HotSpot, once compiled with GraalVM native-image. Step 1: the HotSpot version takes noticeably longer to reach peak request throughput after each deploy, but eventually reaches slightly higher peak throughput once C2 has fully optimized the hot paths. Step 2: the native-image version starts at essentially its full throughput immediately, since there’s no warm-up tier to climb, but that throughput ceiling is somewhat lower, since it lacks runtime profile-driven speculation. Answer: the tradeoff is concrete and measurable, not just theoretical: HotSpot wins on sustained peak throughput for long-lived instances, native-image wins on deploy-to-full-speed latency, which is exactly why teams pick per-workload rather than treating one as universally better.
Related Terms
Referenced by