Skip to content
HN On Hacker News ↗

Postgres with QUIC: Why Stream Multiplexing Beats Traditional TCP Connection Pooling

▲ 19 points • by tenesripranav • 1w ago • HN discussion ↗

Pangram verdict · v3.3

We believe that this entire text is AI.

97 %

AI likelihood · overall

AI
0% human-written 100% AI-generated
SEGMENTS · HUMAN 0 of 1
SEGMENTS · AI 1 of 1
WORD COUNT 1,540
PEAK AI % 97% · §1
Analyzed
Sep 30
backend: pangram/v3.3
Segments scanned
1 windows
avg 1540 words each
Distribution
0 / 100%
human / AI fraction
Verdict
AI
Pangram v3.3

Article text · 1,540 words · 1 segments analyzed

Human AI-generated
§1 AI · 97%

Every backend engineer running Postgres at scale eventually learns the same painful lesson: connection pooling does not fix network physics. We deploy PgBouncer or PgCat, configure transaction pooling, lock down backend connections so the database doesn't run out of memory, and celebrate. But if your application or edge services sit tens of milliseconds away from your database cluster, your queries are still choking on a transport bottleneck we rarely talk about: the client-side TCP connection pool. Recently, I ran an experiment to address this directly: running the Postgres protocol over QUIC (UDP) instead of traditional TCP. We modified PgCat to accept QUIC connections and multiplexed hundreds of virtual database streams over a tiny handful of UDP associations. The outcome was startling. Under heavy load, our QUIC setup pushed 9,718 queries per second at an average latency of 246ms. The identical workload over traditional TCP collapsed into a catastrophic queue pileup of over 40,000 backlogged queries, spiraling to 7,088ms average latency. Here is the real engineering breakdown of why this happens, the arithmetic of latency asymmetry, where QUIC genuinely changes the rules, and where it cannot save you. The Synchronous Postgres Reality To understand the problem, you have to look at the PostgreSQL frontend/backend wire protocol (Protocol 3.0). Postgres is fundamentally synchronous at the connection layer. When a client sends a Query or Execute message down a Postgres TCP connection, that socket is tied up until the server finishes processing and returns the matching CommandComplete and ReadyForQuery message. You cannot interleave two independent queries from different application threads on the same physical connection without strict sequential serialization. The PostgreSQL Wire Protocol Invariant 1 Connection = 1 Active In-Flight Query or Transaction. A client socket is completely blocked from the moment bytes leave the network card until the final row batch and ReadyForQuery flag return across the wire. Because establishing a new Postgres backend process on the database server involves fork-exec overhead, memory allocation (several megabytes per connection for work_mem, cache metadata, and process state), and catalog locks, we can't let 5,000 application threads open 5,000 direct connections to Postgres. The server would instantly melt from context switching and OOM crashes. So we put connection poolers in the middle. But look closely at where the pooler actually sits. The Latency Asymmetry: 30ms WAN vs 2ms LAN In modern distributed architectures, your services don't live in the database rack. You have edge nodes, serverless workers, microservices in different availability zones, or regional clusters in us-east-1 querying a centralized database cluster in us-east-2 or eu-central-1. Let's consider a realistic, standard production deployment: Client to Pooler (WAN / Inter-region): Round-trip time (RTT) is 30ms. Pooler to Postgres (Internal LAN / Same VPC): Round-trip time is <2ms (often sub-millisecond). Query execution time inside the Postgres engine: A fast indexed point lookup or one-liner takes 0.5ms. Figure 1: The Asymmetric Latency Pipeline Application PodClient Pool max_connections = 10 10 Sockets Bound ← 30ms WAN Pipe → Each query occupies 1 connection for 30.5ms total waiting on network wire flight Database VPCPgCat + Postgres <2ms internal LAN latency Cap: 50,000+ QPS ! The Starvation Dilemma: The database engine executes the query in 0.5ms and the internal pooler recycles the connection in 2ms. But the client socket cannot be reused for 30.5ms because the bytes are still flying over the public internet or cross-region fiber! The Math of Client-Side Connection Starvation Here is where basic arithmetic exposes the flaw. Suppose your application client maintains a pool of 10 TCP connections. Why 10? Because in microservice architectures, you might have 50 pods. If every pod creates 200 connections, you quickly overwhelm the pooler and exceed connection limits. Now, calculate the maximum possible throughput your client can achieve: // Little's Law applied to client connection pooling: Throughput (QPS) = Concurrent Connections / Latency per Query Max QPS = 10 connections / 0.030 seconds = 333.3 queries/sec Read that number carefully: 333 queries per second. It does not matter that your Postgres server is an AWS r6i.16xlarge with 64 vCPUs sitting at 3% CPU utilization. It does not matter that PgCat has 100 warm backend connections ready to serve. Your client physically cannot send more than 10 queries every 30 milliseconds. Now, what happens if an inbound burst of 500 HTTP requests hits that pod at once, and each request needs one simple query? The first 10 queries claim the 10 TCP connections and fly out across the 30ms pipe. The remaining 490 queries are forced to wait in an in-memory client queue inside your application runtime (Tokio task queues, connection pool mutexes, Go channel buffers). At 333 QPS drain rate, query #500 will wait in client memory for 1.5 seconds before its bytes even touch the network wire! Your APM dashboard reports that database latency spiked to 1,500ms, even though Postgres executed the query in 0.4ms! Why You Can't Just Open 1,000 TCP Sockets The naive response to this math is: "Just bump client pool size! Change max_connections = 1000!" Every backend engineer who has tried this in production knows the wall you hit immediately: 1. OS Sockets & File Descriptors Every TCP connection consumes a file descriptor. When you scale pods and raise connection pools, you quickly trip EMFILE: too many open files. Even if you tune ulimit -n, the kernel has to track state for tens of thousands of active sockets across epoll queues. 2. Kernel Buffer RAM Bloat TCP is not free memory. The Linux kernel allocates transmit and receive buffers (tcp_wmem and tcp_rmem) for every socket. 2,000 open TCP connections can quietly chew through hundreds of megabytes of kernel slab memory just waiting on keep-alives. 3. Handshake Penalties If idle connections drop or timeout, re-establishing a TCP connection requires a 3-way handshake (1 RTT) plus a TLS handshake (1 to 2 RTTs). Over a 30ms link, creating a new connection takes 60ms to 90ms before a single byte of SQL can be sent. 4. Head-of-Line (HOL) Blocking If any TCP packet in the sequence drops on the public internet, the TCP window freezes. The entire connection halts until that missing packet is retransmitted and acknowledged, even if the application has subsequent data ready to process. We are caught in a trap: we need thousands of concurrent queries in flight over the 30ms pipe to utilize the database, but we cannot afford the architectural overhead of thousands of raw TCP connections. How QUIC Solves the Bottleneck QUIC (RFC 9000) changes the transport mechanics entirely. QUIC runs on top of UDP and provides natively multiplexed, independent bidirectional streams over a single connection. Here is why this shifts the equation for database pooling: Streams are not sockets: In QUIC, opening a stream does not create a Linux kernel socket, does not allocate a file descriptor, and requires zero network handshakes. It is simply an in-memory frame with a 62-bit stream ID sent inside an existing encrypted UDP association. No Head-of-Line Blocking: If Stream #4 drops a packet, only Stream #4 pauses. Stream #1 through Stream #3 and Streams #5 through #40 keep streaming data without waiting. Negligible Client-Side Pooling Overhead: Instead of opening 1,000 TCP sockets, your client can maintain just 10 to 50 QUIC connections, and open 40, 80, or 200 concurrent streams on each one. 0-RTT Resumption (And the Postgres Handshake Reality): In theory, QUIC supports 0-RTT session resumption via TLS 1.3 session tickets. This makes it possible to absorb transient network hiccups or aggressively drop and resume idle connections without waiting 1 RTT for transport handshakes. While we have not implemented 0-RTT resumption in this prototype yet, there is an important database reality to keep in mind: even if your transport layer resumes in zero round trips, you still have to execute the application-level Postgres handshake (sending the StartupMessage with credentials and database name, and awaiting AuthenticationOk and ReadyForQuery) before queries can execute. Even with that constraint, eliminating the initial TCP and TLS transport setup latency is a major win for connection stability. Figure 2: TCP Connection Overhead vs QUIC Stream Multiplexing Traditional TCP Pooling Query 1 1 TCP Socket (FD #12) Query 2 1 TCP Socket (FD #13) Query 3 1 TCP Socket (FD #14) ... Query 1000 1000 Sockets (FD Exhaustion) High memory footprint, socket table limits, and kernel buffer overhead per connection. QUIC Multiplexed Streams Stream ID 0x04 Query 1 (In-memory frame) Stream ID 0x08 Query 2 (In-memory frame) Stream ID 0x0C Query 3 (In-memory frame) 1 Physical UDP Socket ∞ Concurrent Streams Zero kernel socket allocation per query. Concurrency is limited only by buffer limits and pooler capacity. Now, apply Little's Law again: // With 50 QUIC connections and 40 streams each = 2,000 concurrent streams: Max QPS = 2,000 streams / 0.030 seconds = 66,666 queries/sec! By removing the socket tax, the client can keep thousands of queries in flight across the 30ms pipe simultaneously. The client queue drops to zero, and the pooler on the other end is finally saturated with the queries it was designed to handle! The Reality Check: What QUIC Cannot Fix Any engineer telling you QUIC is a universal silver bullet for database performance is selling snake oil. We need to be completely clear about the boundary between transport multiplexing and database execution. The Multi-Statement Transaction Trap