Serverless Computing and Cold Starts
Serverless Computing and Cold Starts
Definition: An execution model in which the cloud provider dynamically provisions, runs, and tears down the underlying compute resources on a customer’s behalf, charging only for the exact compute time consumed rather than for pre-purchased, always-on capacity. The name is something of a misnomer — servers obviously still exist underneath — but from the developer’s perspective the server’s lifecycle, capacity planning, and patching disappear behind the platform’s API entirely. Amazon commercially popularized the model with AWS Lambda’s launch in 2014, cementing Function as a Service (FaaS) as the dominant meaning of “serverless,” though the term itself saw earlier, more niche use around 2012 for backend-as-a-service platforms.
How It Works
- Developers upload pure function code (a handler) rather than a deployable server image; the provider packages it into an execution environment and invokes it in response to events.
- Event-driven invocation: triggers include HTTP requests (via an API Gateway), message queue arrivals, storage bucket writes, database change streams, or cron-style schedules — the function only runs when one of these fires.
- Automatic, per-request scaling: the provider spins up as many parallel execution environments as there are concurrent invocations, from zero up to a provider-defined concurrency limit, with no manual scaling configuration.
- Fine-grained, usage-based billing: charged in small increments of execution time (AWS Lambda commonly bills per 1ms) multiplied by allocated memory, rather than per provisioned server-hour, so an idle function costs nothing.
- Cold starts: when a function is invoked and no warm execution environment exists — after a period of inactivity, or when scaling beyond currently-warm capacity — the provider must allocate a new sandbox, fetch the code package, initialize the language runtime, and run any module-level/global initialization code before the handler itself executes, adding latency that only that specific request pays.
- Warm starts: subsequent invocations landing on an already-initialized execution environment skip all of that setup and run the handler directly, which is why traffic shape (steady vs. bursty) heavily affects how often users actually experience cold-start latency.
- Managed concurrency, not managed servers: the provider decides how many execution environments to keep warm, when to recycle them, and how to route retries, so scaling and failover logic that a team would otherwise write themselves (health checks, instance replacement) is simply absent from the code.
Why It Matters
- Eliminates the operational overhead of patching operating systems, managing idle server costs, and configuring auto-scaling rules.
- Pay-per-use pricing makes intermittent or spiky workloads (a webhook handler firing a few times a day) dramatically cheaper than a provisioned server sitting idle most of the time.
- Forces, and enables, a naturally stateless, horizontally-scalable architecture by construction, since the platform can create or destroy execution environments at will and offers no guarantee any two invocations land on the same one.
- Shifts operational burden further up the Cloud Service Models stack than even PaaS — there’s no server to patch, size, or keep alive, only function code and the events that trigger it.
- Lowers the barrier to shipping small, event-driven pieces of logic (a webhook handler, an image resizer, a scheduled cleanup job) without standing up and maintaining a whole service just to host them.
Under the Hood: What Actually Causes a Cold Start
A cold start is the sum of several sequential costs the provider can’t skip on a genuinely fresh invocation: locating and provisioning compute capacity for the sandbox (a lightweight VM like Firecracker, or a container, depending on the provider), downloading and unpacking the deployment package or container image, booting the language runtime itself — the JVM and .NET CLR are notably slower to initialize than Node.js or Python, often adding hundreds of milliseconds to well over a second — and finally running the function’s module-level initialization code (imports, SDK client construction, database connection setup) before the actual handler logic ever gets to run. Providers mitigate this at several of those layers rather than eliminating it outright: Firecracker microVMs (used by AWS Lambda and Fargate) cut provisioning time to roughly 125ms by stripping the VM boot path down to the bare minimum a sandboxed function needs; “provisioned concurrency” (Lambda) or “minimum instances” (Cloud Run, Cloud Functions) let customers pay to keep a pool of environments permanently warm, trading away the scale-to-zero cost savings for a latency guarantee; and snapshot-based approaches, such as Lambda SnapStart for Java, resume a function from a pre-initialized memory snapshot taken after the expensive runtime and class-loading work is already done, instead of repeating it on every cold invocation. None of these eliminate cold starts entirely, they only shrink how often one happens or how expensive it is when it does, which is why cold-start-sensitive, latency-critical paths still frequently avoid pure serverless in favor of an always-on container or VM.
Comparison: Serverless vs Containers vs VMs for Compute
| Serverless (FaaS) | Containers | Virtual Machines | |
|---|---|---|---|
| Unit managed | function/handler | container image | full OS instance |
| Scale-to-zero | yes, native | only with extra tooling (e.g. Knative, KEDA) | rarely, usually always-on |
| Cold start | milliseconds to seconds, on demand | seconds if scaled to zero | tens of seconds to minutes |
| Billing granularity | per-invocation/ms | per-second while running | per-second/hour while running |
| Max execution duration | capped (Lambda: 15 minutes) | unbounded | unbounded |
| Best fit | event-driven, spiky/intermittent workloads | steady microservices, portable across clouds | long-running, stateful, or licensing-restricted workloads |
| Operational ownership | provider owns almost all of it | customer owns image, orchestration | customer owns OS, patching, scaling |
Common Pitfalls
- Storing local state (like files or session data) in memory between function executions, as serverless containers are ephemeral and can be destroyed at any time.
- Ignoring cold start latency in synchronous user-facing APIs, leading to intermittent slow responses for users.
- Building functions with heavy dependencies or large deployment packages, which directly worsens cold start time since more code has to be downloaded and initialized before the handler can run.
- Hitting the platform’s maximum execution duration (e.g., 15 minutes on Lambda) on long-running batch or processing jobs that were never a good fit for a request/response execution model to begin with.
- Underestimating the cost of “serverless glue” — chaining many small functions together via queues and event buses can rack up per-invocation costs and end-to-end latency that a single longer-running service wouldn’t incur.
- Not accounting for concurrency limits: a burst of traffic can hit an account- or function-level concurrency cap, causing throttled invocations exactly when the system is under the most load.
- Treating a function’s timeout as a safety net rather than a design constraint — retried, timed-out invocations against a non-idempotent handler (e.g. one that charges a payment) can silently duplicate side effects.
Code Example
// AWS Lambda handler (Node.js) — comments mark what runs on cold start vs every invocation
const { DynamoDBClient } = require("@aws-sdk/client-dynamodb");
// Module-level code runs ONCE per cold start, not per invocation —
// initializing the SDK client here (not inside the handler) lets warm
// invocations reuse the same connection instead of re-creating it.
const client = new DynamoDBClient({ region: "us-east-1" });
exports.handler = async (event) => {
// This line runs on every invocation, cold or warm.
const userId = event.pathParameters.id;
const result = await client.send(/* ...GetItemCommand... */);
return {
statusCode: 200,
body: JSON.stringify(result),
};
};
Best Practices
- Initialize SDK clients, DB connections, and other reusable objects at module scope (outside the handler) so warm invocations skip re-creating them.
- Keep deployment packages small and dependencies minimal — every extra megabyte is extra cold-start download and unpack time.
- Use provisioned concurrency or minimum instances for latency-sensitive, user-facing paths where occasional multi-second cold starts are unacceptable.
- Design functions to be stateless and idempotent, since the platform can create, reuse, or destroy execution environments at any time and may retry invocations on failure.
- Set realistic timeouts and monitor concurrency limits so a traffic spike fails predictably (queued or throttled) rather than silently.
- Log and alert on cold-start frequency and duration separately from warm-invocation latency, since averaging them together hides exactly the tail-latency spikes users notice most.
FAQ
Are cold starts avoidable entirely? Not completely, but the highest-impact case — a request-triggered function on a low-traffic path — can be mitigated with provisioned concurrency or by choosing a faster-initializing runtime.
Why do Java/.NET functions have worse cold starts than Node.js/Python? Runtime and framework initialization (JVM class loading, JIT warmup, .NET CLR startup) is inherently heavier than an interpreted language’s startup, though snapshot-based techniques like Lambda SnapStart specifically target this gap.
Is serverless always cheaper than running a server? No — it’s cheaper for intermittent or spiky workloads with real idle time; a service under sustained, predictable, high load is very often cheaper on reserved VM or container capacity, see Cloud Service Models for the broader cost tradeoff.
Does every invocation of a warm function skip initialization entirely? Yes for module-level code, but the handler body itself still runs in full on every invocation — only the one-time setup (imports, client construction) is skipped, which is exactly why that setup belongs outside the handler function.
History
- The term “serverless” saw early, more niche use around 2012 for backend-as-a-service platforms like Parse and early Firebase, before FaaS as it’s known today existed.
- AWS Lambda’s launch in 2014 is generally credited with popularizing the modern FaaS meaning of “serverless” and making it a mainstream compute option.
- Google Cloud Functions and Azure Functions both followed in 2016, and Cloudflare Workers (2017) pushed the model further to the edge using V8-isolate-based sandboxes for near-zero cold starts instead of container/VM-based ones.
- AWS’s 2018 release and open-sourcing of Firecracker gave the industry a shared, purpose-built microVM technology for sandboxing serverless workloads, since adopted well beyond Lambda itself.
- Cold-start mitigation kept advancing through the early 2020s — provisioned concurrency (2019) and Lambda SnapStart (2022) both target the same problem from different angles, provisioning ahead of time versus resuming from a pre-warmed snapshot.
Related Terms
- Cloud Service Models
- Content Delivery Network (CDN) and Edge Computing
- Virtual Machines (VMs)
- Auto-Scaling
- Event-Driven Architecture
- API Gateway
- Kubernetes (K8s)
Example
AWS Lambda and Google Cloud Functions allow developers to run backend code without managing servers, automatically scaling in response to traffic. A team might put a rarely-hit admin endpoint on Lambda for near-zero idle cost while accepting the occasional cold start, but pin their high-traffic checkout API to provisioned concurrency (or a regular container) specifically to keep tail latency low for paying customers.
Referenced by