AW Dev Rethought

✂️ Perfection is achieved not when there is nothing more to add, but when there is nothing left to take away - Antoine de Saint-Exupéry

Networking Internals: What Actually Happens During a TCP Handshake


Introduction:

The TCP handshake is one of those concepts that every engineer encounters early in their career. Three steps — SYN, SYN-ACK, ACK — establish a connection before data can be sent. It appears in networking textbooks, in interview questions, and in explanations of why the first request to a service is slower than subsequent ones.

Most engineers stop at this surface-level understanding. The handshake happens, the connection is established, and the interesting part — the actual data transfer — begins. What happens during those three steps, why each step is necessary, what state is created and maintained on both sides, and how the handshake influences everything that follows is rarely explored in depth.

That depth matters in production. TCP handshake behaviour directly affects application latency, connection pool design, load balancer configuration, and the performance characteristics of services under load. Engineers who understand what actually happens during a TCP handshake make better decisions about all of these concerns than engineers who know only that three steps are involved.


Why a Handshake Is Necessary at All:

TCP is a connection-oriented protocol that provides reliable, ordered delivery of data. Achieving these guarantees requires both sides of a connection to agree on initial state before data transfer begins — specifically, each side needs to know the other's initial sequence number, which is used to order packets and detect missing ones.

The handshake is the mechanism through which this agreement is reached. It is not bureaucratic overhead — it is the minimum exchange required to establish the shared state that TCP's reliability guarantees depend on. A protocol that skipped the handshake and began sending data immediately would have no way to detect whether the remote end was ready to receive, no agreed basis for packet ordering, and no way to distinguish between a new connection and a retransmitted packet from a previous connection.

Understanding why the handshake exists makes its mechanics intuitive rather than arbitrary.


Step One — SYN:

The client initiates a connection by sending a SYN packet to the server. This packet contains the client's initial sequence number — a randomly chosen value that will be used to number every byte of data the client subsequently sends.

The sequence number is not zero. It is chosen randomly to prevent a class of attacks where an attacker who can predict sequence numbers can inject data into an existing connection. The randomness of the initial sequence number is a security property, not an implementation detail.

When the SYN packet is sent, the client transitions to the SYN_SENT state and starts a timer. If no response is received before the timer expires, the client retransmits the SYN — a behaviour that is visible in network traces as repeated SYN packets when a server is unreachable or overwhelmed.


Step Two — SYN-ACK:

When the server receives the SYN packet, it allocates resources for the half-open connection and responds with a SYN-ACK packet. This packet serves two purposes simultaneously — it acknowledges the client's SYN by incrementing the client's sequence number by one, and it sends the server's own initial sequence number so the client can acknowledge it in return.

The server maintains a data structure called the SYN queue — also called the incomplete connection queue — to track half-open connections that have received a SYN but not yet completed the handshake. This queue has a finite size, and exhausting it is the mechanism behind SYN flood attacks — where an attacker sends large numbers of SYN packets without completing the handshake, filling the queue and preventing legitimate connections from being established.

SYN cookies are the standard mitigation for SYN flood attacks. Instead of allocating queue space when a SYN is received, the server encodes the connection state into the sequence number of the SYN-ACK response. If the ACK arrives and can be validated against the sequence number, the connection is legitimate. If no ACK arrives, no queue space was consumed.


Step Three — ACK:

The client receives the SYN-ACK, acknowledges the server's sequence number by incrementing it by one, and sends the ACK packet. At this point the client considers the connection established and can begin sending data — the ACK and the first data packet can be sent together, which is an optimisation that reduces the latency of the first request.

When the server receives the ACK, it moves the connection from the SYN queue to the accept queue — also called the complete connection queue — which holds fully established connections waiting to be accepted by the application. The application calls accept() to retrieve connections from this queue. If the application is slow to call accept(), the queue fills and new connections are dropped even though the handshake completed successfully.

Both queues — the SYN queue and the accept queue — have configurable size limits that affect how a system behaves under connection load. Tuning these limits is one of the first steps in optimising a server for high connection rates.


The Round-Trip Cost of Connection Establishment:

The TCP handshake requires one full round trip between client and server before any application data can be sent. For a client and server in the same data centre, this round trip may be less than a millisecond. For a client in India connecting to a server in the United States, the round trip may be 150 to 200 milliseconds — added to every new connection regardless of how small the request is.

This latency cost is why connection reuse matters. A connection pool that maintains open connections eliminates the handshake latency for subsequent requests because the connections are already established. HTTP/1.1 keep-alive, HTTP/2 multiplexing, and database connection pools all exist partly to amortise the handshake cost across multiple requests.

TLS adds additional round trips on top of the TCP handshake — the TLS handshake negotiates encryption parameters and exchanges certificates before encrypted data can be sent. TLS 1.3 reduced this to one additional round trip from the two required by TLS 1.2, and TLS session resumption allows subsequent connections to skip part of the TLS handshake entirely. Understanding the cumulative latency of TCP and TLS handshakes explains why connection establishment is often the dominant latency factor for short-lived connections.


Connection State Has Resource Implications:

Every established TCP connection consumes resources on both sides — memory for the socket state, buffer space for send and receive windows, and a file descriptor on the operating system. These resources are finite, and systems under high connection load can exhaust them in ways that are not immediately obvious from application-level metrics.

The TIME_WAIT state — entered by the side that initiates connection closure after the connection is closed — holds connection state for a period typically equal to twice the maximum segment lifetime, often 60 seconds. This prevents delayed packets from a closed connection being misinterpreted as belonging to a new connection on the same port. Under high connection turnover, TIME_WAIT connections can accumulate to the point where the available port range is exhausted and new connections cannot be established.

Understanding connection state and its resource implications is necessary for diagnosing a class of production problems — connection exhaustion, port exhaustion, and file descriptor limits — that appear as mysterious failures to establish new connections in systems that seem otherwise healthy.


Conclusion:

The TCP handshake is not a detail that can be safely ignored after a surface-level understanding is established. Its mechanics — the state it creates, the resources it consumes, the latency it introduces, and the failure modes it enables — directly affect the performance and reliability of every networked system in production.

Engineers who understand what actually happens during a TCP handshake diagnose connection-related production problems more effectively, make better decisions about connection pool sizing and reuse strategies, and understand why optimisations like TLS 1.3 and HTTP/2 improve performance in ways that are invisible to engineers who treat the handshake as a black box. The three steps are the beginning of the understanding, not the end of it.


50 Machine Learning Algorithms Cheatsheet Most downloaded free PDF 50 Machine Learning Algorithms Cheatsheet Blueprints of Intelligence — 50 ML Algorithms decoded Download free Downloaded 395 times

If this article helped you, you can support my work on AW Dev Rethought.


Rethought Relay:
Link copied!

Enjoyed this post?

Stay in the loop

New posts + weekly digest, straight to your inbox.

or

Create a free account

  • Save posts to your vault
  • Like posts & build history
  • New-post alerts

Comments

Add Your Comment

Comment Added!