Virtual Private Cloud (VPC) and Subnets
Virtual Private Cloud (VPC) and Subnets
Definition: A Virtual Private Cloud is a logically isolated section of a public cloud provider’s network where an account can launch resources into a virtual network it defines — its own IP address range, routing rules, and gateways — rather than sharing a flat address space with every other customer. Subnets are the smaller IP ranges a VPC gets carved into, typically aligned to availability zones, and each subnet’s routing determines whether the resources inside it are reachable from the public internet. AWS introduced VPC in 2009, and every major cloud provider (GCP’s VPC, Azure’s Virtual Network) now ships some version of the same idea, evolving out of an earlier, less isolated model where cloud tenants shared address space more directly.
How It Works
- You define a CIDR block (e.g.
10.0.0.0/16) for the VPC, giving it roughly 65,000 usable addresses, then carve that block into smaller Subnets, each with its own CIDR range and availability zone - Public subnets have a route table entry pointing to an Internet Gateway (IGW), giving instances inside them a path to and from the public internet, typically via a public or elastic IP
- Private subnets have no direct internet route; instances inside must route outbound traffic through a NAT Gateway sitting in a public subnet, which lets them reach the internet (for patches, API calls) without being reachable from it
- Route tables are the actual mechanism that makes a subnet “public” or “private” — the label is just shorthand for what a subnet’s route table points its default route at
- Security Groups and Network ACLs act as virtual firewalls at two layers — Security Groups are stateful and attach to individual instances/interfaces, NACLs are stateless and attach to whole subnets, and both are typically used together for defense in depth
- VPC Peering and Transit Gateways let separate VPCs (different teams, different environments, or different accounts entirely) communicate privately without routing traffic over the public internet
- VPC Endpoints let resources in a private subnet reach cloud provider services (object storage, managed databases) directly over the provider’s internal network, avoiding a NAT Gateway hop and its associated cost for provider-native traffic
Why It Matters
- Forms the foundational security boundary for cloud infrastructure, preventing internal databases and backend services from being directly reachable from the public internet by default
- Lets an organization replicate a familiar on-premises network topology (DMZ-style public tier, private application tier, isolated data tier) inside a cloud provider, instead of accepting whatever flat network model the provider offers by default
- Multiple VPCs (per environment, per team, or per compliance boundary) give hard isolation between, say, production and staging, so a misconfiguration or compromise in one doesn’t automatically expose the other
- Almost every other cloud security control — Identity and Access Management (IAM) policies, Kubernetes (K8s) cluster networking, Load Balancer placement — is layered on top of a VPC’s network boundaries, making VPC design one of the first decisions that constrains everything built afterward
- Compliance regimes (PCI-DSS, HIPAA, SOC 2) frequently mandate demonstrable network segmentation between sensitive and non-sensitive systems, and VPC/subnet boundaries are usually the concrete artifact auditors point to as evidence that segmentation exists
Under the Hood: CIDR Math and Subnet Sizing
Every VPC and subnet boundary is ultimately just binary arithmetic on an IP address expressed as a CIDR block — the /16 in 10.0.0.0/16 means the first 16 bits of the 32-bit address are fixed as the network portion, leaving 16 bits (2^16 = 65,536 addresses) free for hosts and subnets within it. Carving that into subnets means borrowing bits from the host portion: a /24 subnet fixes 24 bits and leaves 8 free (256 addresses, though cloud providers reserve the first four and the last one for network, gateway, DNS, and broadcast addresses, leaving 251 usable). Getting this sizing wrong in either direction causes real problems — a VPC CIDR too small forces a disruptive re-architecture once an organization outgrows its address space (VPCs are difficult to resize after resources are already using their addresses), while subnets sized too small silently run out of IPs as auto-scaling groups or Kubernetes clusters try to launch more instances than the subnet has addresses left for, a failure mode that often only shows up under peak load, the worst possible time to discover it.
Comparison: VPC vs Flat Public Networking vs On-Premises Networking
| VPC (cloud, isolated) | Flat Public Cloud Networking | On-Premises Networking | |
|---|---|---|---|
| Isolation | Logical isolation per account/VPC | Minimal — shared address space by default | Physical isolation, fully self-controlled |
| Setup effort | Moderate — CIDR planning, route tables, gateways | Low, but little control once running | High — physical cabling, hardware firewalls |
| Elasticity | High — new subnets/peering created on demand | High | Low — bounded by physical hardware |
| Control granularity | Fine-grained (Security Groups, NACLs, route tables) | Coarse, provider-dependent | Fine-grained but manually operated |
| Cost model | Pay for gateways, peering, data transfer | Often bundled, less transparent | High upfront capital cost, low marginal cost |
Common Pitfalls
- Placing database instances in public subnets with permissive security groups, exposing them directly to automated internet scanning and brute-force attacks
- Choosing a VPC CIDR block too small for future growth (or one that overlaps with another VPC’s range), which makes later peering or account consolidation painful or impossible without re-addressing everything
- Forgetting that NAT Gateways bill per-hour and per-GB processed, letting a private subnet with heavy outbound traffic produce a surprisingly large bill
- Relying on Security Groups alone and ignoring NACLs (or vice versa), missing the defense-in-depth benefit of having both a stateful and a stateless filtering layer
- Over-peering VPCs into a dense mesh instead of using a Transit Gateway, creating an unmanageable web of point-to-point connections as the number of VPCs grows
- Leaving default VPC security groups wide open (“allow all from anywhere”) from initial setup and never tightening them once real workloads move in
- Spanning a single subnet across multiple availability zones (not possible on AWS, but a common conceptual mistake) instead of understanding that each subnet is tied to exactly one AZ, requiring at least one subnet per AZ for genuine high availability
Code Example
resource "aws_vpc" "main" {
cidr_block = "10.0.0.0/16"
}
resource "aws_subnet" "public" {
vpc_id = aws_vpc.main.id
cidr_block = "10.0.1.0/24"
availability_zone = "us-east-1a"
map_public_ip_on_launch = true
}
resource "aws_subnet" "private" {
vpc_id = aws_vpc.main.id
cidr_block = "10.0.2.0/24"
availability_zone = "us-east-1a"
}
resource "aws_internet_gateway" "igw" {
vpc_id = aws_vpc.main.id
}
resource "aws_route" "public_default" {
route_table_id = aws_vpc.main.default_route_table_id
destination_cidr_block = "0.0.0.0/0"
gateway_id = aws_internet_gateway.igw.id
}
Best Practices
- Plan CIDR ranges deliberately before launching anything, leaving room to grow and avoiding overlap with any VPC you might ever need to peer with
- Default to private subnets for anything that doesn’t need to be directly internet-facing, and put only load balancers and bastion hosts in public subnets
- Use Security Groups as the primary control (stateful, attached to resources) and NACLs as a coarse-grained backstop at the subnet boundary
- Prefer a Transit Gateway over a mesh of point-to-point VPC peering connections once an organization has more than a handful of VPCs to interconnect
- Tag and document subnets clearly (environment, tier, availability zone) so route table and security group intent stays legible months later
FAQ
What’s the difference between a Security Group and a NACL? A Security Group is stateful and attaches to individual instances or network interfaces — allow the inbound request and the response is automatically allowed back out — while a NACL is stateless and attaches to a whole subnet, requiring explicit inbound and outbound rules for both directions of traffic.
Do private subnet resources have any internet access at all? Yes, outbound-only, via a NAT Gateway in a public subnet — a private instance can call out to an external API or download a patch, but nothing on the internet can initiate a connection to it directly.
Can two VPCs in different regions talk to each other privately? Yes, through VPC peering or a Transit Gateway, both of which route traffic across the provider’s private backbone instead of the public internet, though peering connections don’t transitively route through a third VPC without a Transit Gateway.
Why does a subnet have fewer usable IPs than its CIDR block suggests?
Cloud providers reserve a handful of addresses in every subnet — typically the network address, the VPC router, DNS, a reserved-for-future-use address, and the broadcast address — so a /24 subnet’s theoretical 256 addresses works out to roughly 251 actually usable ones.
History
- Cloud providers initially offered flatter, less isolated networking models before customer demand for enterprise-grade network isolation grew
- AWS introduced VPC in 2009, letting customers define their own private IP space, subnets, and routing inside AWS’s shared infrastructure, and made it the default (rather than optional) networking model for new accounts by 2013
- GCP and Azure both shipped their own VPC-equivalents (Google’s global VPC model differs notably by spanning regions natively, unlike AWS’s region-scoped VPCs) as the concept became a baseline cloud networking expectation rather than a differentiator
- The rise of Kubernetes and container networking in the mid-to-late 2010s added another layer of IP address planning on top of VPC/subnet design, since cluster CNI plugins often need their own address ranges carved out of the same VPC space
Related Terms
- IP Addressing and Subnetting
- NAT
- Zero Trust Architecture
- Identity and Access Management (IAM)
- Load Balancer
- Cloud Service Models
Example
An organization places its EC2 web servers in a public subnet behind a load balancer, and its RDS databases in a private subnet with no route to the internet at all — even if an attacker compromises the web tier, the database is unreachable except from within the VPC, and a security group rule restricting database access to only the web tier’s security group (rather than an IP range) keeps that boundary tight even as instances are replaced.
Referenced by