TCP Protocol
TCP Protocol
Definition: Transmission Control Protocol (TCP) is a connection-oriented transport protocol that provides reliable, ordered, error-checked delivery of a byte stream between two hosts.
How It Works
Connection establishment (three-way handshake):
- Client sends
SYNwith an initial sequence number (ISN). - Server replies
SYN-ACK, acknowledging the client’s ISN and sending its own ISN. - Client replies
ACK, acknowledging the server’s ISN. Connection is now established.
Reliable delivery: every byte in the stream is numbered (sequence numbers). The receiver sends ACKs indicating the next byte it expects, letting the sender know what’s been received and what needs retransmitting. A retransmission timer (RTO, dynamically computed from measured round-trip time) fires if an ACK doesn’t arrive in time.
Flow control: the receiver advertises a window size, how many unacknowledged bytes it’s willing to buffer, in every ACK. The sender never sends more than that, preventing it from overwhelming a slow receiver.
Congestion control: separate from flow control, this protects the network itself. Slow start ramps the sending rate up exponentially from a small initial window; congestion avoidance then increases linearly; a detected loss (assumed to mean congestion) cuts the rate back sharply (multiplicative decrease). Modern stacks commonly use CUBIC (Linux default) or BBR (model-based, used by Google/YouTube) instead of the original Reno algorithm.
Connection teardown (four-way handshake): either side can initiate. Side A sends FIN, side B ACKs it (A’s send direction is now closed), B later sends its own FIN when it’s also done, A ACKs that. This is why TCP connections are technically full-duplex and can half-close, one side can stop sending while still receiving.
Under the Hood
TCP segment header (20 bytes minimum, before options):
| Field | Size | Purpose |
|---|---|---|
| Source port / Destination port | 2 + 2 bytes | Identify sending/receiving application |
| Sequence number | 4 bytes | Position of first data byte in this segment |
| Acknowledgment number | 4 bytes | Next byte the sender expects to receive |
| Data offset, flags | 2 bytes | Header length + control bits: SYN, ACK, FIN, RST, PSH, URG |
| Window size | 2 bytes | Receiver’s current flow-control window |
| Checksum | 2 bytes | Error detection over header + data |
| Urgent pointer | 2 bytes | Rarely used |
| Options | variable | MSS, window scaling, SACK permitted, timestamps |
Key options that matter in practice: window scaling (extends the 16-bit window field via a scale factor, needed for high-bandwidth-delay-product links to avoid throttling throughput); SACK (Selective ACK, lets a receiver report exactly which non-contiguous byte ranges arrived, so the sender retransmits only the actual gap instead of everything after the first loss); timestamps (used for more precise RTT measurement and to protect against wrapped sequence numbers on fast links).
Sequence numbers are 32-bit and start from a randomized ISN (not zero) specifically to prevent an old, delayed segment from a previous incarnation of the same connection from being misinterpreted as valid data in a new one, and to make blind session-hijacking/spoofing harder.
RST (reset) is TCP’s abrupt-abort signal, sent when a segment arrives for a connection the receiver has no record of, or when an application wants to kill a connection immediately instead of going through the FIN handshake, no further data exchange, no guarantee of graceful delivery of what was in flight.
Connection State Machine
A TCP connection moves through a well-defined set of states, visible via netstat/ss:
| State | Meaning |
|---|---|
LISTEN | Server socket waiting for incoming connections |
SYN_SENT | Client sent SYN, awaiting SYN-ACK |
SYN_RECEIVED | Server received SYN, sent SYN-ACK, awaiting final ACK |
ESTABLISHED | Handshake complete, data can flow |
FIN_WAIT_1 / FIN_WAIT_2 | Initiated close, awaiting peer’s FIN |
CLOSE_WAIT | Peer closed, this side hasn’t called close() yet |
TIME_WAIT | Closed locally, lingering to catch any delayed duplicate segments |
A large number of connections stuck in CLOSE_WAIT usually points to an application bug, the peer closed its end but the code never called close() on the socket. A large number in TIME_WAIT on a busy server is often normal, but can exhaust ephemeral ports under very high connection churn.
History: Congestion Control Evolution
- Tahoe (1988): the original congestion control algorithm, introduced slow start and congestion avoidance, and reacted to any packet loss by dropping the congestion window all the way back to its initial size.
- Reno (1990): added fast retransmit and fast recovery, on a single lost packet (detected via duplicate ACKs) it halves the window instead of resetting to minimum, recovering faster from isolated loss while still treating loss as the primary congestion signal.
- CUBIC (mid-2000s, Linux’s default since 2.6.19): uses a cubic function of time since the last loss event to grow the window, designed to scale better on high-bandwidth, high-latency links than Reno’s linear growth, without being overly aggressive.
- BBR (Bottleneck Bandwidth and RTT, Google, 2016): a departure from loss-based congestion control entirely, it models the actual bottleneck bandwidth and round-trip time and paces sending to match, rather than waiting for loss as a signal, used heavily by Google/YouTube and increasingly available as a pluggable Linux congestion control module.
This progression reflects a broader shift: early algorithms treated packet loss as the only usable congestion signal, because that’s what the internet mostly gave routers to work with (drop-tail queues), while newer algorithms increasingly use direct measurement and modeling as networks and instrumentation improved.
Why It Matters
TCP is what makes it safe to assume “if my application code sends bytes, they arrive, in order, uncorrupted, or I get told about the failure.” That guarantee is what HTTP, SSH, database wire protocols, and most application-layer protocols are built on top of, none of them have to reimplement retransmission or ordering themselves.
Common Pitfalls
- Head-of-line blocking: because TCP guarantees strict ordering, a single lost segment stalls delivery of every later segment to the application, even ones that already arrived, until the gap is retransmitted and filled. This is the specific problem HTTP/3’s QUIC (over UDP) was designed to avoid.
- Handshake latency: a fresh TCP connection costs a full round trip before any data can flow, and HTTPS adds a TLS handshake on top, which is why connection reuse (keep-alive, connection pooling) matters so much for performance.
- Treating
send()success as delivery confirmation, TCP guarantees eventual delivery or an error, not synchronous acknowledgment at the application level. - Nagle’s algorithm interacting badly with delayed ACKs, small writes can appear to stall for ~40ms in some naive request/response protocols unless
TCP_NODELAYis set. - Assuming TCP fixes application-level framing, it delivers an ordered byte stream, not discrete messages, application protocols still need their own length prefixes or delimiters.
Comparison
| TCP | UDP | QUIC | |
|---|---|---|---|
| Connection | Connection-oriented | Connectionless | Connection-oriented (over UDP) |
| Reliability | Guaranteed, ordered | None | Guaranteed per-stream, no global HoL blocking |
| Handshake cost | 1 RTT (+1 for TLS) | None | 0-1 RTT (TLS built in) |
| Head-of-line blocking | Yes | N/A | No, across independent streams |
| Typical use | Web, SSH, databases | DNS, video calls, gaming | HTTP/3 |
Debugging Workflow: Reading a tcpdump Capture
Diagnosing a slow or failing TCP connection usually starts with tcpdump -i eth0 host <ip> and port <port> -w capture.pcap and reading the result (in Wireshark, or tcpdump -r capture.pcap -tttt):
- Confirm the handshake completes: look for
[SYN]from the client,[SYN, ACK]from the server,[ACK]from the client. A[SYN]with no[SYN, ACK]reply means the server isn’t reachable or isn’t listening on that port, a firewall issue or the service is down. - Check for retransmissions: Wireshark flags them explicitly (
[TCP Retransmission]); frequent retransmissions of the same segment point to packet loss on the path, not an application bug. - Watch the window size: a receiver repeatedly advertising
[TCP ZeroWindow]means the receiving application isn’t reading data fast enough, the sender is being flow-controlled, not congestion-controlled. - Look at the teardown: a clean close shows
[FIN, ACK]from both sides. A connection that just stops producing packets with noFINorRSTsuggests a silent network failure (dead link, dropped NAT mapping) rather than either side deliberately closing. - Correlate timestamps: a large gap between the handshake completing and the first data segment often points to slow application-level processing (e.g. a slow TLS handshake or slow database query) rather than a network problem at all.
Example
curl -v https://example.com shows the TCP handshake implicitly happening before the TLS handshake and HTTP request. A packet capture (tcpdump or Wireshark) on port 443 shows the literal [SYN], [SYN, ACK], [ACK] sequence at the start of every HTTPS connection, followed later by [FIN, ACK] exchanges at teardown.
FAQ
Why 3 segments to open a connection but 4 to close it? Opening only needs to synchronize sequence numbers in both directions, which a combined SYN-ACK handles in one segment. Closing is per-direction: each side must independently signal “I’m done sending” and get that acknowledged, and the two directions don’t have to close at the same time.
Can TCP guarantee a message arrives, or just the bytes? Just the bytes, in order. TCP has no concept of “message” at all, framing (delimiters, length prefixes) is entirely the application’s responsibility.
Why does a TCP connection sometimes hang instead of failing fast? Without SO_KEEPALIVE or an application-level heartbeat, a peer that silently disappears (crashed, network partition) may not generate any signal at all, TCP alone will just keep retrying retransmissions per its backoff schedule.
Is congestion control the same thing as flow control? No. Flow control protects the receiver from being overwhelmed, using its advertised window. Congestion control protects the network from being overwhelmed, inferring congestion from loss or delay and throttling the sender independently of what the receiver’s window allows.
Related Terms
Referenced by