TCP Protocol

TCP Protocol

Definition: Transmission Control Protocol (TCP) is a connection-oriented transport protocol that provides reliable, ordered, error-checked delivery of a byte stream between two hosts.

How It Works

Connection establishment (three-way handshake):

  1. Client sends SYN with an initial sequence number (ISN).
  2. Server replies SYN-ACK, acknowledging the client’s ISN and sending its own ISN.
  3. Client replies ACK, acknowledging the server’s ISN. Connection is now established.

Reliable delivery: every byte in the stream is numbered (sequence numbers). The receiver sends ACKs indicating the next byte it expects, letting the sender know what’s been received and what needs retransmitting. A retransmission timer (RTO, dynamically computed from measured round-trip time) fires if an ACK doesn’t arrive in time.

Flow control: the receiver advertises a window size, how many unacknowledged bytes it’s willing to buffer, in every ACK. The sender never sends more than that, preventing it from overwhelming a slow receiver.

Congestion control: separate from flow control, this protects the network itself. Slow start ramps the sending rate up exponentially from a small initial window; congestion avoidance then increases linearly; a detected loss (assumed to mean congestion) cuts the rate back sharply (multiplicative decrease). Modern stacks commonly use CUBIC (Linux default) or BBR (model-based, used by Google/YouTube) instead of the original Reno algorithm.

Connection teardown (four-way handshake): either side can initiate. Side A sends FIN, side B ACKs it (A’s send direction is now closed), B later sends its own FIN when it’s also done, A ACKs that. This is why TCP connections are technically full-duplex and can half-close, one side can stop sending while still receiving.

Under the Hood

TCP segment header (20 bytes minimum, before options):

FieldSizePurpose
Source port / Destination port2 + 2 bytesIdentify sending/receiving application
Sequence number4 bytesPosition of first data byte in this segment
Acknowledgment number4 bytesNext byte the sender expects to receive
Data offset, flags2 bytesHeader length + control bits: SYN, ACK, FIN, RST, PSH, URG
Window size2 bytesReceiver’s current flow-control window
Checksum2 bytesError detection over header + data
Urgent pointer2 bytesRarely used
OptionsvariableMSS, window scaling, SACK permitted, timestamps

Key options that matter in practice: window scaling (extends the 16-bit window field via a scale factor, needed for high-bandwidth-delay-product links to avoid throttling throughput); SACK (Selective ACK, lets a receiver report exactly which non-contiguous byte ranges arrived, so the sender retransmits only the actual gap instead of everything after the first loss); timestamps (used for more precise RTT measurement and to protect against wrapped sequence numbers on fast links).

Sequence numbers are 32-bit and start from a randomized ISN (not zero) specifically to prevent an old, delayed segment from a previous incarnation of the same connection from being misinterpreted as valid data in a new one, and to make blind session-hijacking/spoofing harder.

RST (reset) is TCP’s abrupt-abort signal, sent when a segment arrives for a connection the receiver has no record of, or when an application wants to kill a connection immediately instead of going through the FIN handshake, no further data exchange, no guarantee of graceful delivery of what was in flight.

Connection State Machine

A TCP connection moves through a well-defined set of states, visible via netstat/ss:

StateMeaning
LISTENServer socket waiting for incoming connections
SYN_SENTClient sent SYN, awaiting SYN-ACK
SYN_RECEIVEDServer received SYN, sent SYN-ACK, awaiting final ACK
ESTABLISHEDHandshake complete, data can flow
FIN_WAIT_1 / FIN_WAIT_2Initiated close, awaiting peer’s FIN
CLOSE_WAITPeer closed, this side hasn’t called close() yet
TIME_WAITClosed locally, lingering to catch any delayed duplicate segments

A large number of connections stuck in CLOSE_WAIT usually points to an application bug, the peer closed its end but the code never called close() on the socket. A large number in TIME_WAIT on a busy server is often normal, but can exhaust ephemeral ports under very high connection churn.

History: Congestion Control Evolution

  • Tahoe (1988): the original congestion control algorithm, introduced slow start and congestion avoidance, and reacted to any packet loss by dropping the congestion window all the way back to its initial size.
  • Reno (1990): added fast retransmit and fast recovery, on a single lost packet (detected via duplicate ACKs) it halves the window instead of resetting to minimum, recovering faster from isolated loss while still treating loss as the primary congestion signal.
  • CUBIC (mid-2000s, Linux’s default since 2.6.19): uses a cubic function of time since the last loss event to grow the window, designed to scale better on high-bandwidth, high-latency links than Reno’s linear growth, without being overly aggressive.
  • BBR (Bottleneck Bandwidth and RTT, Google, 2016): a departure from loss-based congestion control entirely, it models the actual bottleneck bandwidth and round-trip time and paces sending to match, rather than waiting for loss as a signal, used heavily by Google/YouTube and increasingly available as a pluggable Linux congestion control module.

This progression reflects a broader shift: early algorithms treated packet loss as the only usable congestion signal, because that’s what the internet mostly gave routers to work with (drop-tail queues), while newer algorithms increasingly use direct measurement and modeling as networks and instrumentation improved.

Why It Matters

TCP is what makes it safe to assume “if my application code sends bytes, they arrive, in order, uncorrupted, or I get told about the failure.” That guarantee is what HTTP, SSH, database wire protocols, and most application-layer protocols are built on top of, none of them have to reimplement retransmission or ordering themselves.

Common Pitfalls

  • Head-of-line blocking: because TCP guarantees strict ordering, a single lost segment stalls delivery of every later segment to the application, even ones that already arrived, until the gap is retransmitted and filled. This is the specific problem HTTP/3’s QUIC (over UDP) was designed to avoid.
  • Handshake latency: a fresh TCP connection costs a full round trip before any data can flow, and HTTPS adds a TLS handshake on top, which is why connection reuse (keep-alive, connection pooling) matters so much for performance.
  • Treating send() success as delivery confirmation, TCP guarantees eventual delivery or an error, not synchronous acknowledgment at the application level.
  • Nagle’s algorithm interacting badly with delayed ACKs, small writes can appear to stall for ~40ms in some naive request/response protocols unless TCP_NODELAY is set.
  • Assuming TCP fixes application-level framing, it delivers an ordered byte stream, not discrete messages, application protocols still need their own length prefixes or delimiters.

Comparison

TCPUDPQUIC
ConnectionConnection-orientedConnectionlessConnection-oriented (over UDP)
ReliabilityGuaranteed, orderedNoneGuaranteed per-stream, no global HoL blocking
Handshake cost1 RTT (+1 for TLS)None0-1 RTT (TLS built in)
Head-of-line blockingYesN/ANo, across independent streams
Typical useWeb, SSH, databasesDNS, video calls, gamingHTTP/3

Debugging Workflow: Reading a tcpdump Capture

Diagnosing a slow or failing TCP connection usually starts with tcpdump -i eth0 host <ip> and port <port> -w capture.pcap and reading the result (in Wireshark, or tcpdump -r capture.pcap -tttt):

  1. Confirm the handshake completes: look for [SYN] from the client, [SYN, ACK] from the server, [ACK] from the client. A [SYN] with no [SYN, ACK] reply means the server isn’t reachable or isn’t listening on that port, a firewall issue or the service is down.
  2. Check for retransmissions: Wireshark flags them explicitly ([TCP Retransmission]); frequent retransmissions of the same segment point to packet loss on the path, not an application bug.
  3. Watch the window size: a receiver repeatedly advertising [TCP ZeroWindow] means the receiving application isn’t reading data fast enough, the sender is being flow-controlled, not congestion-controlled.
  4. Look at the teardown: a clean close shows [FIN, ACK] from both sides. A connection that just stops producing packets with no FIN or RST suggests a silent network failure (dead link, dropped NAT mapping) rather than either side deliberately closing.
  5. Correlate timestamps: a large gap between the handshake completing and the first data segment often points to slow application-level processing (e.g. a slow TLS handshake or slow database query) rather than a network problem at all.

Example

curl -v https://example.com shows the TCP handshake implicitly happening before the TLS handshake and HTTP request. A packet capture (tcpdump or Wireshark) on port 443 shows the literal [SYN], [SYN, ACK], [ACK] sequence at the start of every HTTPS connection, followed later by [FIN, ACK] exchanges at teardown.

FAQ

Why 3 segments to open a connection but 4 to close it? Opening only needs to synchronize sequence numbers in both directions, which a combined SYN-ACK handles in one segment. Closing is per-direction: each side must independently signal “I’m done sending” and get that acknowledged, and the two directions don’t have to close at the same time.

Can TCP guarantee a message arrives, or just the bytes? Just the bytes, in order. TCP has no concept of “message” at all, framing (delimiters, length prefixes) is entirely the application’s responsibility.

Why does a TCP connection sometimes hang instead of failing fast? Without SO_KEEPALIVE or an application-level heartbeat, a peer that silently disappears (crashed, network partition) may not generate any signal at all, TCP alone will just keep retrying retransmissions per its backoff schedule.

Is congestion control the same thing as flow control? No. Flow control protects the receiver from being overwhelmed, using its advertised window. Congestion control protects the network from being overwhelmed, inferring congestion from loss or delay and throttling the sender independently of what the receiver’s window allows.

Dig deeper