Virtual Machines (VMs)
Virtual Machines (VMs)
Definition: A software-based emulation of a complete physical computer — CPU, memory, storage, and network interfaces — that runs its own independent Guest OS on top of physical hardware (the Host), with a Hypervisor arbitrating access to the underlying resources and enforcing isolation between guests. The concept traces back to IBM’s mainframe virtualization work in the 1960s-70s (CP-40 and VM/370), largely disappeared as cheap, single-purpose x86 servers took over in the 1980s-90s, and was commercially revived by VMware in 1999, which brought hardware virtualization to commodity x86 hardware for the first time. That revival is what made modern cloud computing possible — practically every IaaS offering today (AWS EC2, GCP Compute Engine, Azure VMs) is, underneath, a fleet of physical hosts sliced into rentable VMs.
How It Works
- Hypervisor (Type 1, “bare-metal”): runs directly on physical hardware with no host OS underneath — examples include VMware ESXi, Xen, and KVM (used internally by AWS, GCP, and most cloud providers); this is what production cloud infrastructure runs on.
- Hypervisor (Type 2, “hosted”): runs as an application on top of a conventional OS (VirtualBox, VMware Workstation, Parallels) — common for local development, less common in production.
- Resource allocation and isolation: the hypervisor allocates physical CPU (via time-slicing or dedicated cores), memory, and storage to each VM and enforces isolation between them; each VM runs a completely independent Guest OS kernel, unaware it’s sharing underlying physical hardware with other VMs.
- Contrast with containers: VMs virtualize the hardware, so each VM needs its own full OS kernel and typically takes tens of seconds to boot; containers virtualize the OS by sharing the host kernel via namespaces/cgroups, so they start in milliseconds and have a far smaller footprint — see Kubernetes (K8s) for how containers are orchestrated at scale.
- Lightweight microVMs: hypervisors like Firecracker (built by AWS, used for Lambda and Fargate) narrow this gap by booting a minimal VM in around 125ms, combining VM-level isolation with near-container startup speed — see Serverless Computing and Cold Starts.
- Instance families: cloud providers size VMs into instance families/types (e.g., AWS
t3.medium,m5.large) that bundle a fixed vCPU/RAM/network ratio tuned for general-purpose, compute-optimized, or memory-optimized workloads, and the same virtual disk image can be snapshotted, cloned, or migrated between physical hosts as a portable unit. - Virtual disks: persistent VM storage lives in a disk image (VMDK, QCOW2) or a cloud block volume (AWS EBS, GCP Persistent Disk) that the hypervisor presents to the guest as a raw block device, decoupling the VM’s data lifecycle from the physical host it happens to be running on.
Why It Matters
- Enables infrastructure consolidation (running 10 VMs on 1 physical server instead of buying 10 physical servers) and forms the foundation of modern Infrastructure as a Service (IaaS) cloud computing — see Cloud Service Models.
- Strong isolation (separate kernels, not just separate processes) makes VMs the default choice for multi-tenant environments where security boundaries matter more than density or startup latency.
- The pay-per-second/hour rental model turns capacity planning from a multi-week hardware procurement cycle into an API call, letting teams provision and discard compute on demand instead of over-buying physical servers for peak load.
- Snapshot- and image-based deployment makes VMs portable and recoverable by construction — a golden image or snapshot can relaunch an identical instance on a different host or region for disaster recovery, see High Availability (HA) and Disaster Recovery (DR).
- Broad OS and licensing compatibility: because a VM presents a full virtual machine rather than a shared kernel, it can run essentially any OS the hypervisor supports, which is why regulated, legacy, or Windows-licensed workloads default to VMs over containers.
Under the Hood: How a Hypervisor Virtualizes CPU and Memory
CPU virtualization historically relied on trap-and-emulate: privileged instructions issued by the guest kernel would trap into the hypervisor, which emulated their effect and returned control to the guest. The problem was that x86 wasn’t classically virtualizable — it has a handful of “sensitive but unprivileged” instructions that silently behave differently instead of trapping, so early solutions had to work around hardware rather than with it, either by rewriting guest code on the fly (VMware’s binary translation) or by modifying the guest OS to call the hypervisor directly instead of issuing those instructions at all (Xen’s paravirtualization). Intel VT-x and AMD-V, introduced in 2005-2006, solved this properly by adding a new hardware CPU mode: the hypervisor runs in “root mode” and the guest runs in a lower-privilege “non-root mode” where sensitive instructions now trap by design, letting most guest code execute at near-native speed with hardware handling the boundary. Memory virtualization followed a parallel path — early hypervisors maintained “shadow page tables,” a hypervisor-managed shadow copy of each guest’s page table that had to be kept in sync on every guest page table update, which was expensive. Extended/Nested Page Tables (EPT on Intel, NPT on AMD) moved that translation into hardware instead, adding a second layer of address translation (guest-virtual to guest-physical to host-physical) that the CPU’s memory management unit walks directly. Memory overcommit — running VMs whose combined allocated memory exceeds physical RAM — is handled through a balloon driver installed inside the guest, which the hypervisor inflates to reclaim pages back to the host under memory pressure, and through transparent page sharing (or KSM on Linux/KVM), which deduplicates identical memory pages across VMs, common when many guests run the same OS.
Comparison: VMs vs Containers vs Bare Metal
| Virtual Machines | Containers | Bare Metal | |
|---|---|---|---|
| Isolation boundary | separate kernel per VM, hypervisor-enforced | shared host kernel, namespace/cgroup-isolated | none — full physical machine, single tenant |
| Boot time | seconds to tens of seconds (microVMs ~125ms) | milliseconds | minutes for a full OS boot, or none if already running |
| Density per host | tens of VMs | hundreds to thousands of containers | one workload |
| Security isolation | strong, hardware-enforced | weaker — a kernel exploit can cross container boundaries | strongest possible, nothing is shared |
| Typical use | multi-tenant IaaS, legacy apps, strict isolation needs | microservices, CI/CD, high-density scheduling | latency-critical, licensing-restricted, or max-performance workloads |
| Cost model | pay per second/hour of provisioned instance | usually billed alongside the host it runs on | highest fixed cost, no multi-tenant sharing |
Common Pitfalls
- “VM Sprawl”: creating numerous virtual machines without proper lifecycle management, leading to wasted cloud spend on idle instances and unpatched, forgotten security vulnerabilities.
- Over-provisioning instance size “just in case” instead of right-sizing based on actual CPU/memory utilization, which is one of the single largest sources of avoidable cloud bill waste.
- Treating VMs as pets instead of cattle: manually configuring long-lived instances by hand makes them fragile and non-reproducible; see Infrastructure as Code (IaC) for the alternative.
- Letting long-lived instances drift out of patch compliance because nobody owns rebuilding their base image, turning “just a VM” into a standing security liability.
- Defaulting to on-demand pricing for steady-state workloads that would be meaningfully cheaper on reserved or committed-use pricing, or ignoring spot/preemptible capacity for interruption-tolerant batch jobs.
- Ignoring the “noisy neighbor” effect on burstable instance types (e.g., AWS
t3family): these accrue CPU credits while idle and get throttled hard once credits run out under sustained load, surprising teams who only load-tested in short bursts. - Skipping regular OS-level backups in the assumption that the cloud provider’s infrastructure durability also covers accidental deletion or corruption inside the guest OS — it generally doesn’t without an explicit backup/snapshot policy.
Code Example
# cloud-init: runs once on first boot of a new VM instance to configure it
#cloud-config
hostname: web-01
package_update: true
packages:
- nginx
- fail2ban
users:
- name: deploy
groups: sudo
shell: /bin/bash
ssh_authorized_keys:
- ssh-ed25519 AAAAC3Nza... deploy@ci
write_files:
- path: /etc/nginx/sites-available/default
content: |
server {
listen 80;
root /var/www/html;
}
runcmd:
- systemctl enable nginx
- systemctl start nginx
Best Practices
- Right-size instances based on actual CPU/memory utilization metrics collected over time, and revisit that sizing periodically rather than once at launch.
- Provision through Infrastructure as Code (Terraform, CloudFormation) instead of manually clicking through a console, so instances are reproducible cattle, not hand-tended pets.
- Bake golden images (with tools like Packer) that already include patches and hardening, instead of configuring each instance by hand after boot.
- Mix pricing models deliberately: reserved/committed-use for steady baseline load, spot/preemptible for interruption-tolerant batch work, on-demand only for genuine burst.
- Automate patching and schedule regular image rebuilds rather than patching long-lived instances in place indefinitely.
- Snapshot before risky changes and test restore procedures ahead of time, since a disaster-recovery plan nobody has actually exercised is not a reliable one.
FAQ
Are VMs still relevant now that containers are common? Yes — VMs remain the right choice wherever strong, kernel-level isolation matters (multi-tenant platforms, regulated workloads), for legacy applications tied to a specific OS, and for licensing or compliance requirements that assume a full dedicated machine.
What’s the difference between a VM snapshot and an image (or AMI)? A snapshot is a point-in-time capture of a running instance’s disk state, typically used for backup or rollback; an image (or AMI on AWS) is a reusable template — often built from a snapshot — used to launch new instances from scratch.
Why do “burstable” instance types slow down under sustained load? They run on a CPU credit system: the instance accumulates credits while idle or lightly loaded and spends them under load, and once credits are exhausted the CPU is throttled to its baseline performance until credits replenish.
How is a VM different from a Docker container at a technical level? A VM includes a full guest OS kernel and is isolated by the hypervisor at the hardware level; a container shares the host’s single kernel and is isolated only by kernel-level namespaces and cgroups, which is faster and lighter but a narrower security boundary.
History
- IBM’s CP-40 and later VM/370 pioneered virtualization on mainframes in the 1960s-70s, letting a single machine run multiple isolated operating system instances decades before x86 virtualization existed.
- x86 was considered practically unvirtualizable until VMware, founded in 1998, shipped VMware Workstation in 1999 using binary translation — the first mainstream virtualization on commodity hardware.
- Xen (2003, University of Cambridge) popularized paravirtualization and became the hypervisor behind early AWS EC2 (launched 2006), the moment renting a VM by the hour became a mainstream commercial product.
- Intel VT-x and AMD-V (2005-2006) added hardware-assisted virtualization support, and cloud providers later built custom hypervisors optimized for multi-tenant cloud workloads specifically, such as AWS’s Nitro system.
- Containers rose through the 2013-2015 era (Docker, then Kubernetes) as a lighter alternative for many workloads, but rather than replacing VMs outright, most production clouds today run containers on top of VMs for an extra layer of isolation.
Related Terms
- Cloud Service Models
- Kubernetes (K8s)
- Auto-Scaling
- Serverless Computing and Cold Starts
- Infrastructure as Code (IaC)
- High Availability (HA) and Disaster Recovery (DR)
- AWS (Amazon Web Services)
Example
Amazon EC2 and Google Compute Engine provide on-demand Virtual Machines in the cloud, billed per second/hour based on instance type; under the hood, AWS Lambda’s isolation between customer function invocations is itself implemented using lightweight Firecracker microVMs rather than plain OS containers. A team migrating a legacy monolith might snapshot its production VM, launch an identical clone in a second region as a disaster-recovery standby, and only later consider re-architecting the application onto containers once its VM-based operations are stable.
Referenced by