TCP and UDP, blocking sockets, NIO selectors and buffers, Netty, Java 11 HttpClient, HTTP/1.1 vs HTTP/2, TLS, DNS caching, timeouts, pooling, CORS, proxies, WebSockets, backpressure, and debugging hangs, resets and slow calls.
Theory
Q1
What is the difference between TCP and UDP?
basic
TCP is a connection-oriented, reliable, ordered byte stream with flow and congestion control. UDP is connectionless, sends independent datagrams, and gives no delivery, ordering or duplicate protection.
TCP: three-way handshake, retransmission, no message boundaries.
UDP: low overhead, preserves datagram boundaries, suited to DNS, metrics, streaming and QUIC (which builds reliability on top).
Java: Socket/ServerSocket/SocketChannel for TCP, DatagramSocket/DatagramChannel for UDP.
⚠ Follow-up traps
Does one write on a TCP socket arrive as one read? No. TCP has no message boundaries; you must frame messages (length prefix or delimiter).
Is UDP always unreliable in practice? It gives no guarantees; the application must add acks and retries if it needs them.
#tcp#udp#transport
Q2
Describe the TCP three-way handshake and connection teardown.
basic
Handshake: client sends SYN, server replies SYN-ACK, client sends ACK. Teardown uses FIN/ACK in each direction (four segments), so each side closes independently (half-close).
connect() returns after the SYN-ACK/ACK exchange; accept() returns connections already completed by the kernel.
An abrupt close sends RST instead of FIN.
The side that sends the first FIN enters TIME_WAIT.
⚠ Follow-up traps
Who completes the handshake, accept() or the kernel? The kernel; accept() only dequeues an established connection from the accept queue.
What does Socket.shutdownOutput() do? Sends FIN but keeps the input side open (half-close).
#tcp#handshake#teardown
Q3
How do Socket and ServerSocket work in Java?
basic
ServerSocket binds to a port and accept() blocks until a client connects, returning a Socket. Each Socket exposes an InputStream and OutputStream for blocking reads and writes.
try (ServerSocket server = new ServerSocket(8080)) { while (true) { Socket s = server.accept(); new Thread(() -> handle(s)).start(); }}
new ServerSocket(port, backlog) sets the accept-queue length hint.
Always close sockets (try-with-resources) to release the file descriptor.
⚠ Follow-up traps
What does accept() do when the backlog is full? The kernel drops or delays new SYNs/ACKs; clients see connect delays or timeouts.
Does closing the InputStream close the socket? Yes, closing either stream closes the socket.
#socket#serversocket#blocking-io
Q4
What are the socket options SO_TIMEOUT, SO_KEEPALIVE, TCP_NODELAY, SO_REUSEADDR and SO_LINGER?
intermediate
SO_TIMEOUT (setSoTimeout): max blocking time of a read/accept, throws SocketTimeoutException.
SO_KEEPALIVE: OS sends probes on idle connections (default idle 2 hours on Linux).
TCP_NODELAY: disables Nagle's algorithm so small writes go out immediately.
SO_REUSEADDR: lets a server rebind a port with connections in TIME_WAIT.
Does SO_TIMEOUT apply to connect? No; use socket.connect(addr, timeoutMs).
Is SO_REUSEADDR the same as SO_REUSEPORT? No; the latter lets multiple sockets bind the same port for load spreading.
#socket-options#tcp
Q5
What is Nagle's algorithm and when should you disable it?
intermediate
Nagle buffers small writes until the previous data is acknowledged, reducing tiny packets. Combined with delayed ACK on the peer, it can add up to ~40 ms latency to request/response protocols that issue several small writes.
Disable with socket.setTcpNoDelay(true) for latency-sensitive RPC.
Better fix: coalesce writes into one buffer or flush once.
⚠ Follow-up traps
Does TCP_NODELAY make data arrive faster over a slow link? No, it only removes sender-side batching delay.
Is it on by default in Java? No, TCP_NODELAY defaults to false (Netty and many HTTP clients enable it).
#nagle#tcp-nodelay#latency
Q6
How does the thread-per-connection blocking model scale, and what are its limits?
basic
Each connection holds a thread blocked in read. It is simple but each platform thread costs ~1 MB of stack reserve plus scheduling overhead, so tens of thousands of idle connections are impractical.
Thread pools cap resource usage but then idle connections can starve active ones.
Java 21 virtual threads let you keep the blocking style: a blocking socket read parks the virtual thread, not the carrier.
⚠ Follow-up traps
Is the thread stack actually allocated up front? Virtual memory is reserved; physical pages are committed on use.
Do virtual threads make blocking Socket I/O non-blocking? The JDK reimplemented Socket on NIO so reads park the virtual thread; the programming model stays blocking.
#blocking-io#threads#scalability
Q7
What is Java NIO and how does it differ from classic IO?
basic
NIO (java.nio) uses channels, buffers and selectors. Classic IO is stream-oriented and blocking; NIO is buffer-oriented, can be non-blocking, and one thread can multiplex many channels.
Channels (SocketChannel, ServerSocketChannel, FileChannel, DatagramChannel) are bidirectional.
Data is read into / written from ByteBuffers.
Selector reports which channels are ready.
⚠ Follow-up traps
Is NIO always faster than IO? No; for few connections blocking IO is as fast and simpler. NIO wins on connection count.
Are FileChannels selectable? No; only SelectableChannel subclasses such as socket channels.
A buffer has capacity (fixed), limit (first index not to read/write) and position (next index). After writing into a buffer, flip() sets limit=position and position=0 to read it.
clear(): position=0, limit=capacity (ready to write; data not erased).
compact(): copies unread bytes to the start and sets position after them, ready to write more; use after a partial read.
rewind(): position=0, limit unchanged.
ByteBuffer b = ByteBuffer.allocate(8);b.putInt(7); // pos=4b.flip(); // pos=0 limit=4int v = b.getInt(); // 7
⚠ Follow-up traps
What happens if you forget flip() before channel write? It writes from position to limit, which is empty or garbage bytes after position.
Does clear() zero the data? No, only resets indices.
#bytebuffer#nio#buffers
Q9
What is the difference between heap and direct ByteBuffers?
intermediate
A heap buffer wraps a byte[] in the Java heap; a direct buffer lives in native memory. Channel I/O on a heap buffer copies through a temporary direct buffer, so direct buffers avoid one copy for I/O.
Direct allocation is slower and costs more, and it is freed only when the buffer object is GC'd (via a cleaner), limited by -XX:MaxDirectMemorySize.
Use direct buffers for long-lived, reused I/O buffers; pool them.
OutOfMemoryError: Direct buffer memory signals exhaustion.
⚠ Follow-up traps
Are direct buffers counted in -Xmx? No, they are off-heap with a separate limit.
Can GC pressure be low yet direct memory run out? Yes; if the heap rarely collects, cleaners never run.
#bytebuffer#direct-memory#performance
Q10
How does a Selector work and what are the SelectionKey operations?
intermediate
A channel configured non-blocking is registered with a Selector for interest ops (OP_ACCEPT, OP_CONNECT, OP_READ, OP_WRITE). select() blocks until at least one is ready, then you iterate selectedKeys().
Selector sel = Selector.open();ServerSocketChannel ssc = ServerSocketChannel.open();ssc.bind(new InetSocketAddress(8080));ssc.configureBlocking(false);ssc.register(sel, SelectionKey.OP_ACCEPT);while (true) { sel.select(); var it = sel.selectedKeys().iterator(); while (it.hasNext()) { SelectionKey k = it.next(); it.remove(); // handle accept/read }}
On Linux the implementation uses epoll, on macOS kqueue.
⚠ Follow-up traps
Why must you remove the key from selectedKeys()? The selector never clears it; otherwise it is processed again every loop.
Is OP_WRITE interest normally always set? No; sockets are almost always writable, so setting it constantly causes a busy loop. Set it only when a write is pending.
#selector#nio#reactor
Q11
What is the Reactor pattern?
intermediate
One or more event-loop threads wait on a selector, dispatch ready events to handlers, and handlers must not block. Multi-reactor variants use a boss loop to accept and worker loops to read/write.
Netty NioEventLoopGroup, Tomcat NIO connector acceptor/poller and Vert.x use this model.
Blocking work (DB, disk) is offloaded to a separate pool.
⚠ Follow-up traps
What happens if a handler blocks the event loop? All channels bound to that loop stall.
Is the reactor the same as the proactor? No; a proactor (NIO.2 async channels) is notified on completion, not readiness.
#reactor#nio#event-loop
Q12
What is NIO.2 AsynchronousSocketChannel?
advanced
Java 7 AsynchronousSocketChannel/AsynchronousServerSocketChannel start an operation and report completion via Future or CompletionHandler, executed on an AsynchronousChannelGroup thread pool.
Proactor style: you supply the buffer up front; the OS (IOCP on Windows, epoll-based emulation on Linux) completes it.
Rarely used in production; Netty on selectors dominates.
⚠ Follow-up traps
Is it truly kernel-async on Linux? Not for sockets; the JDK emulates it over epoll with internal threads.
Can a CompletionHandler block? It can, but it consumes a pool thread and should be avoided.
#nio2#async#completion-handler
Q13
What is the difference between blocking and non-blocking, and between synchronous and asynchronous I/O?
basic
Blocking vs non-blocking is whether the call waits when no data is ready. Synchronous vs asynchronous is whether the caller itself performs the read after readiness (sync) or is notified once the operation completed (async).
Is NIO with a Selector asynchronous I/O? By strict definition no; it is non-blocking, readiness-based sync I/O.
What does a non-blocking read return when no data is available? 0 bytes (and -1 on end of stream).
#io-models#blocking#async
Q14
What is Netty and why is it used?
intermediate
Netty is an asynchronous event-driven network framework over NIO (or epoll/io_uring natives). It hides selector complexity and gives pipelines, codecs, pooled buffers and backpressure hooks.
Powers Spring WebFlux (Reactor Netty), gRPC-Java, Cassandra and Elasticsearch clients.
⚠ Follow-up traps
Is Netty a web server? It is a toolkit; HTTP is just one codec.
Does Netty use one thread per connection? No; a channel is bound to a single event loop for life, and a loop serves many channels.
#netty#event-loop#framework
Q15
Explain Netty's ChannelPipeline, handlers and ByteBuf.
advanced
The pipeline is an ordered chain of inbound and outbound handlers per channel. Inbound events (read, active) flow head to tail; outbound operations (write, flush) flow tail to head.
Typical chain: SslHandler, HttpServerCodec, HttpObjectAggregator, business handler.
ByteBuf has separate reader and writer indexes (no flip), is reference-counted and usually pooled.
You must release() a ByteBuf you consume (or use SimpleChannelInboundHandler, which does it).
⚠ Follow-up traps
What happens if you forget to release a ByteBuf? Memory leak of pooled buffers; detect with -Dio.netty.leakDetection.level=paranoid.
Where does a ctx.write() go compared with channel.write()?ctx.write starts from the next handler before this one; channel.write starts from the pipeline tail.
#netty#pipeline#bytebuf
Q16
What is Java 11 HttpClient and how is it used?
basic
java.net.http.HttpClient is the standard modern HTTP client: immutable, thread-safe, supports HTTP/1.1 and HTTP/2, WebSocket, sync send and async sendAsync returning CompletableFuture.
Build one client and reuse it; it owns the connection pool and selector thread.
⚠ Follow-up traps
Does it replace HttpURLConnection? It is the recommended replacement; the old class has poor API and weak pooling control.
Is HttpRequest.timeout the connect timeout? No; it bounds time until the response headers (or body handler start) arrive, while connectTimeout is on the client.
#httpclient#java11
Q17
What timeouts does HttpClient support, and what is missing?
intermediate
HttpClient.Builder.connectTimeout limits connection establishment; HttpRequest.Builder.timeout limits the wait for the response (throws HttpTimeoutException). There is no built-in read-idle or total-body timeout.
For total deadlines, use CompletableFuture.orTimeout on sendAsync or cancel the future.
Cancelling the future does not necessarily abort the in-flight request in older JDKs; JDK 16+ improved cancellation.
⚠ Follow-up traps
Is there a default timeout? No; the default is infinite for both.
Does request.timeout cover reading a large streaming body? Not the whole body for streaming handlers; slow-body handling needs your own deadline.
#httpclient#timeouts
Q18
How does HTTP/2 differ from HTTP/1.1?
intermediate
HTTP/2 is a binary, multiplexed protocol: many concurrent streams over one TCP connection, HPACK header compression, stream prioritisation and server push (now deprecated). HTTP/1.1 is textual with one in-flight request per connection.
HTTP/1.1 pipelining exists but suffers head-of-line blocking and is effectively unused; browsers open ~6 connections per host.
HTTP/2 removes application-level HOL blocking but TCP-level HOL blocking remains (solved by HTTP/3 over QUIC).
Browsers require TLS for HTTP/2; h2c is the cleartext variant.
⚠ Follow-up traps
Does HTTP/2 reduce the number of requests? No, it makes them cheaper by sharing a connection.
Is HTTP/2 always faster? Not on lossy networks, where one lost packet stalls all streams.
#http2#http1.1#multiplexing
Q19
How does Java negotiate HTTP/2 and what are its defaults?
intermediate
HttpClient defaults to Version.HTTP_2 and falls back to HTTP/1.1 automatically. Over HTTPS the protocol is chosen through TLS ALPN; over plain HTTP it tries an Upgrade: h2c header.
Force HTTP/1.1 with .version(HttpClient.Version.HTTP_1_1).
Check response.version() to see what was actually used.
Many servers ignore h2c upgrade, so cleartext usually falls back to 1.1.
⚠ Follow-up traps
What is ALPN? A TLS extension in which the client lists protocols (h2, http/1.1) and the server picks one during the handshake.
Does the HTTP/2 client use a connection pool? It keeps one multiplexed connection per origin.
#httpclient#http2#alpn
Q20
How does TLS/HTTPS work and how is it configured in Java?
intermediate
TLS authenticates the server through its certificate chain, negotiates cipher suites and keys (ephemeral ECDHE in TLS 1.2/1.3), then encrypts the stream. Java exposes it through SSLContext, SSLSocket, SSLEngine.
Trust store (javax.net.ssl.trustStore) decides whom to trust; key store holds your private key for mutual TLS.
Defaults: JDK 11+ enables TLS 1.3; the default trust store is cacerts.
Hostname verification must remain enabled; with HttpClient it is on unless the endpoint identification algorithm is cleared.
⚠ Follow-up traps
Why is a trust-all TrustManager dangerous? It accepts any certificate, enabling man-in-the-middle attacks.
Does TLS prove the identity of the client? Only with mutual TLS (client certificates).
#tls#https#sslcontext
Q21
What does a TLS handshake cost, and how is it reduced?
advanced
TLS 1.2 needs 2 round trips after TCP; TLS 1.3 needs 1 (and 0-RTT on resumption). Asymmetric crypto adds CPU cost on the server.
Reduce with connection reuse (keep-alive, HTTP/2), session resumption (tickets), and TLS 1.3.
Terminate TLS at a load balancer to offload CPU.
OCSP stapling avoids a blocking revocation lookup.
⚠ Follow-up traps
Is 0-RTT safe for all requests? No, it is replayable; only idempotent requests should use it.
Is the bulk data encryption expensive? Symmetric AES-GCM with hardware support is cheap; the handshake dominates.
#tls#handshake#session-resumption
Q22
What is the difference between URL and URI in Java?
basic
URI is a pure syntactic parser/value object for identifiers. URL additionally knows a protocol handler and can open a connection; URL.equals/hashCode may perform DNS lookups.
Prefer URI, and convert with uri.toURL() only when needed. Java 20 deprecated new URL(String) constructors in favour of URI.create(...).toURL().
URI.create throws IllegalArgumentException for invalid input; new URI throws a checked URISyntaxException.
⚠ Follow-up traps
Why is URL as a HashMap key problematic?hashCode and equals resolve host names, which blocks and can compare unequal for equal hosts on different IPs.
Does URI encode illegal characters for you? The multi-argument constructors quote illegal characters; the single-string one rejects them.
#url#uri#java.net
Q23
How do you correctly encode URL query parameters and paths?
basic
Use URLEncoder.encode(value, StandardCharsets.UTF_8) for query component values (form encoding, space becomes +) and a URI multi-arg constructor or a builder for paths (space becomes %20).
Encode each component separately, never the whole URL: that would encode ://, / and ?.
Double-encoding (%2520) is a common bug when a framework already encodes.
Spring: UriComponentsBuilder and UriUtils.
⚠ Follow-up traps
Is + a space in a path? No; + means space only in application/x-www-form-urlencoded query/body.
Which charset should be passed? UTF-8; the single-argument URLEncoder.encode(String) is deprecated.
#url-encoding#uri#query
Q24
How does DNS resolution work in Java, and what does InetAddress caching do?
intermediate
InetAddress.getByName calls the OS resolver and caches results inside the JVM. Positive results are cached for 30 seconds by default (when no security manager) and negative results for 10 seconds.
Controlled by security properties networkaddress.cache.ttl and networkaddress.cache.negative.ttl in java.security or set at startup via Security.setProperty.
-1 means cache forever, 0 means no caching.
getAllByName returns all A/AAAA records; resolution order follows java.net.preferIPv4Stack / preferIPv6Addresses.
Is networkaddress.cache.ttl a -D system property? No, it is a security property (-D has no effect); the legacy sun.net.inetaddr.ttl system property also exists.
Does the JVM honour the record's DNS TTL? No, it uses its own fixed TTL.
#dns#inetaddress#caching
Q25
What is the risk of caching DNS forever in the JVM?
advanced
Failover and blue/green deployments that rely on DNS changes will not reach a long-lived JVM, which keeps connecting to the old IP. Cloud services (load balancers, RDS, S3) rotate IPs frequently.
When a security manager is installed (older setups), the default TTL is infinite, a classic production trap.
Set a small positive TTL (for example 30 s) and make clients re-resolve on reconnect.
Pooled connections also pin an old IP until closed; set max connection lifetime.
⚠ Follow-up traps
If the TTL is 30 s, will existing connections move to the new IP? No, only new connections resolve again.
Does a negative TTL of 0 help? It avoids caching failed lookups, but increases load on DNS during outages.
#dns#cache-ttl#failover
Q26
What is the difference between connect timeout and read timeout?
basic
Connect timeout limits TCP connection establishment. Read (socket) timeout limits the wait between bytes while reading, not the total response time.
A request streaming one byte every 29 seconds never trips a 30 s read timeout.
Also needed: connection-acquire (pool) timeout, request/total timeout, and TLS handshake timeout.
Without a connect timeout, the OS default (about 2 minutes via SYN retries on Linux) applies.
⚠ Follow-up traps
Does a read timeout close the connection? The exception is thrown but the socket must be closed by the caller; the connection is in an undefined state and should not be reused.
Is the read timeout the total response deadline? No; use an overall deadline in addition.
#timeouts#connect#read
Q27
How does connection pooling work and why does it matter?
basic
A pool keeps established connections (TCP + TLS done) and leases them to requests, avoiding handshake cost per call and bounding concurrency toward a backend.
Key knobs: max total, max per route/host, acquire timeout, idle eviction, max lifetime, validation.
Apache HttpClient: PoolingHttpClientConnectionManager (default max 20 total, 2 per route in 4.x). OkHttp: ConnectionPool (5 idle, 5 min). JDK HttpClient: internal pool, tuned via system properties such as jdk.httpclient.connectionPoolSize and jdk.httpclient.keepalive.timeout.
A response body must be fully read or closed to return the connection.
⚠ Follow-up traps
Why do requests hang when the body is not closed? The connection is never released, so the pool exhausts and callers wait for a lease.
Is a bigger pool always better? No; it can overload the backend and hit file-descriptor limits.
#connection-pool#performance
Q28
What is HTTP keep-alive and how does it relate to the TCP keep-alive?
basic
HTTP keep-alive (persistent connections, the default in HTTP/1.1) reuses one TCP connection for multiple requests. TCP keep-alive (SO_KEEPALIVE) is an OS probe on idle sockets that detects dead peers. They are unrelated mechanisms.
Servers close idle connections after a timeout (Keep-Alive: timeout=5); a client reusing a just-closed connection gets a reset.
Set the client idle eviction below the server and load balancer idle timeout.
Connection: close disables reuse.
⚠ Follow-up traps
Does SO_KEEPALIVE keep an HTTP connection from timing out at a load balancer? Only if probes fire more often than the LB idle timeout; Linux's default first probe is after 7200 s.
Is keep-alive on by default in HTTP/1.0? No, it needed an explicit Connection: keep-alive.
#keep-alive#http#tcp
Q29
What is TIME_WAIT and why does it exist?
intermediate
After actively closing, a TCP endpoint stays in TIME_WAIT for 2xMSL (60 s on Linux) so late duplicate segments die and the final ACK can be retransmitted. The side that closes first holds the state.
Thousands of TIME_WAIT sockets are normal for a client making many short connections.
They consume ephemeral ports (the 4-tuple stays reserved), not much memory.
Best fix: reuse connections (keep-alive, pooling), not shortening TIME_WAIT.
⚠ Follow-up traps
Does the server or the client enter TIME_WAIT? Whoever sends the first FIN; a server that closes idle connections accumulates them.
Is SO_REUSEADDR a TIME_WAIT bypass? It lets a listener bind despite TIME_WAIT sockets on that port; it does not shorten the state.
#time-wait#tcp#ephemeral-ports
Q30
What are ephemeral ports and how can they be exhausted?
intermediate
The OS assigns a source port from a range (Linux net.ipv4.ip_local_port_range, default 32768-60999, about 28k ports) to each outgoing connection. A connection is identified by (src IP, src port, dst IP, dst port), so exhaustion occurs per destination.
Symptom: BindException: Cannot assign requested address on connect.
Causes: no connection reuse, many TIME_WAIT sockets, NAT gateways with limited ports.
Mitigations: pool connections, widen the port range, add source IPs, tcp_tw_reuse for outgoing connections.
⚠ Follow-up traps
Is the limit 28k connections to a server? 28k concurrent per (source IP, destination IP:port); different destinations reuse ports.
Does tcp_tw_recycle help? It was removed in Linux 4.12 because it broke clients behind NAT.
#ephemeral-ports#port-exhaustion#linux
Q31
How does a forward proxy differ from a reverse proxy, and how do you use a proxy in Java?
intermediate
A forward proxy acts for clients (egress control, caching, anonymity). A reverse proxy acts for servers (TLS termination, routing, load balancing). Java supports proxies through java.net.Proxy, ProxySelector and system properties.
HttpClient c = HttpClient.newBuilder() .proxy(ProxySelector.of(new InetSocketAddress("proxy.corp", 3128))) .build();
System properties: https.proxyHost, https.proxyPort, http.nonProxyHosts.
HTTPS through a forward proxy uses CONNECT host:443 to create a tunnel; the proxy cannot see the content.
⚠ Follow-up traps
Does http.proxyHost also proxy HTTPS? No, HTTPS uses https.proxyHost.
Does HttpClient read system proxy properties by default? No, you must call .proxy(ProxySelector.getDefault()) to use them.
#proxy#forward-proxy#reverse-proxy
Q32
What are Layer 4 and Layer 7 load balancers?
intermediate
L4 balancers (AWS NLB, IPVS) route by IP/port at the TCP/UDP level without parsing HTTP. L7 balancers (ALB, NGINX, Envoy) understand HTTP and can route by path/header, terminate TLS, retry and balance per request.
With HTTP/2 or gRPC over an L4 balancer, all requests of one long-lived connection go to a single backend, so load is uneven; use L7 or client-side balancing.
Algorithms: round robin, least connections, consistent hashing, power of two choices.
⚠ Follow-up traps
Why is one gRPC channel behind an L4 LB unbalanced? Multiplexing keeps all calls on one TCP connection pinned to one backend.
Which hides client IP? L7 proxies; use X-Forwarded-For or PROXY protocol.
#load-balancer#l4#l7
Q33
What is the X-Forwarded-For header and why is it a security concern?
intermediate
Proxies append the client IP to X-Forwarded-For (and set X-Forwarded-Proto/Host) so the origin knows the real client. Clients can send their own forged value.
Only trust the entries added by your own proxies: take the rightmost trusted hop, not the first value.
Spring Boot: server.forward-headers-strategy=native or framework with a trusted proxy.
The standardised alternative is Forwarded (RFC 7239).
⚠ Follow-up traps
Is the leftmost XFF entry the client IP? Only if all hops are trusted; attackers can prepend values.
What is the PROXY protocol? A prefix sent by L4 balancers conveying the original source address.
#proxy#x-forwarded-for#security
Q34
What is CORS and how does it work?
basic
CORS is a browser mechanism that relaxes the same-origin policy: the server declares via Access-Control-Allow-* headers which origins may read its responses. It is enforced by the browser, not the server.
Simple requests are sent directly; the browser checks Access-Control-Allow-Origin on the response.
Access-Control-Allow-Credentials: true requires an explicit origin, not *.
⚠ Follow-up traps
Does CORS protect the API from curl or other servers? No, it only governs browsers.
Does a blocked CORS request reach the server? Simple requests do, the response is merely hidden; side effects still happen.
#cors#browser#security
Q35
Explain the CORS preflight request and its caching.
intermediate
Before a non-simple request the browser sends OPTIONS with Origin, Access-Control-Request-Method and Access-Control-Request-Headers. The server answers with Allow-Origin, Allow-Methods, Allow-Headers, and optionally Access-Control-Max-Age.
Max-Age caches the preflight result (browser caps, for example 2 hours in Chrome).
The preflight must succeed with 2xx and must not require authentication cookies.
Spring: @CrossOrigin or CorsConfigurationSource, applied before Spring Security's filter chain.
⚠ Follow-up traps
Why does the preflight fail with 401 behind Spring Security? The CORS filter was not registered before the auth filter, so OPTIONS is rejected.
Can Allow-Origin contain multiple origins? No, one value only; echo the matching origin and add Vary: Origin.
#cors#preflight#options
Q36
What is the WebSocket protocol and how does the handshake work?
intermediate
WebSocket provides a full-duplex, message-oriented channel over a single TCP connection. It starts as an HTTP/1.1 GET with Upgrade: websocket and Sec-WebSocket-Key; the server replies 101 Switching Protocols with Sec-WebSocket-Accept.
After upgrade, framed messages (text, binary, ping, pong, close) flow in both directions.
Browsers cannot set custom headers on the handshake, so auth uses cookies, query token or a first message.
wss:// is WebSocket over TLS.
⚠ Follow-up traps
Does WebSocket traffic go through normal HTTP load-balancer timeouts? Yes, idle connections are cut; send ping/pong heartbeats.
Does the origin check come from CORS? No, CORS does not apply to WebSockets; the server must validate Origin.
#websocket#upgrade#protocol
Q37
How do you use WebSockets in Java?
intermediate
Client: HttpClient.newWebSocketBuilder() with a WebSocket.Listener. Server: Jakarta WebSocket (@ServerEndpoint) or Spring WebSocketHandler / STOMP.
The JDK client is demand driven: you must call request(n) to receive more messages.
⚠ Follow-up traps
Why does a JDK WebSocket listener receive only one message? The listener must call webSocket.request(1) again; demand is not automatic after the first.
Is sendText safe to call concurrently? No, wait for the previous CompletableFuture to complete before the next send.
#websocket#jakarta#httpclient
Q38
When should you choose WebSockets, SSE or long polling?
intermediate
WebSocket: bidirectional, low-latency (chat, games). Server-Sent Events: one-way server push over plain HTTP with auto-reconnect (notifications, feeds). Long polling: fallback with highest overhead.
SSE works through HTTP/2 and proxies easily but is text only and unidirectional.
WebSockets need sticky sessions or a shared broker to scale beyond one node.
⚠ Follow-up traps
Does SSE support client-to-server messages? No, use separate HTTP requests.
How many SSE connections can a browser open on HTTP/1.1? About 6 per origin, shared with other requests.
#websocket#sse#long-polling
Q39
What is backpressure in network applications?
intermediate
Backpressure is a way for a slow consumer to signal a fast producer to slow down so buffers do not grow unbounded. TCP has it built in: the receiver's window shrinks and the sender's write blocks or becomes not-writable.
Blocking IO: OutputStream.write blocks when the send buffer is full.
Netty: check channel.isWritable() and the write-buffer water marks.
Reactive Streams: subscriber request(n).
⚠ Follow-up traps
What happens if you ignore isWritable() in Netty? Outbound data queues in memory until OutOfMemoryError.
Does an unbounded queue between threads provide backpressure? No, it hides the problem until memory is exhausted.
#backpressure#flow-control#reactive
Q40
How does TCP flow control differ from congestion control?
advanced
Flow control protects the receiver: it advertises a receive window so the sender never exceeds receiver buffers. Congestion control protects the network: the sender limits in-flight data with a congestion window (slow start, congestion avoidance, CUBIC/BBR).
Throughput per connection is bounded by window / RTT (bandwidth-delay product).
A zero window stalls the sender until a window update.
⚠ Follow-up traps
Why is a single TCP stream slow on a high-latency fat link? Window/RTT caps rate unless buffers and window scaling are large enough.
What does SO_RCVBUF set? The receive buffer size and thus max window; set before connect/bind for window scaling to apply.
#tcp#flow-control#congestion
Q41
What are the TCP backlog and accept queue?
advanced
Linux has a SYN queue for half-open connections and an accept queue for completed handshakes awaiting accept(). The Java backlog argument sets the accept-queue limit, capped by net.core.somaxconn (default 4096 on newer kernels, 128 on older).
If the app does not call accept() fast enough, the queue fills and new connections are dropped/retried; clients see connect timeouts.
Check with ss -lnt (Recv-Q is the queue length on a listener, Send-Q the limit).
Always cap the max frame length to prevent memory attacks.
⚠ Follow-up traps
Does InputStream.read(byte[]) fill the whole array? No, it returns what is available; use readFully or loop.
Why enforce a max length? A malicious length field of 2 GB would otherwise cause allocation DoS.
#framing#tcp#protocol-design
Q43
What does `java.net.http.HttpClient` do for threading and what executors does it use?
advanced
Each client has a selector manager thread (HttpClient-N-SelectorManager) for I/O and uses an executor (default: cached thread pool) to run completion stages and handlers. You can supply your own via .executor(...).
sendAsync returns immediately; chained thenApply callbacks run on the client executor unless you use ...Async with another executor.
Blocking in a callback consumes a pool thread.
A client lives until it is unreachable and idle (JDK 21 adds close()/AutoCloseable).
⚠ Follow-up traps
Is creating a new HttpClient per request fine? No; you lose connection reuse and create selector threads each time.
Is HttpClient thread-safe? Yes, it is immutable and shareable.
#httpclient#executor#async
Q44
How do you stream, retry and handle errors with HttpClient?
intermediate
Choose a BodyHandler: ofString, ofByteArray, ofInputStream, ofLines, ofFile, or a subscriber for streaming. HttpClient does not retry on its own beyond the safe internal retry of idempotent requests on connection reset, and non-2xx status codes are not exceptions.
Check statusCode() yourself.
Retry only idempotent requests (GET, PUT, DELETE) with exponential backoff and jitter; for POST use an idempotency key.
ofInputStream must be closed to free the connection.
⚠ Follow-up traps
Does send throw for HTTP 500? No, it returns the response; only I/O failures throw IOException.
Which exception must you handle in send?InterruptedException; restore the interrupt flag when caught.
#httpclient#body-handlers#retry
Q45
How do redirects, cookies and authentication work in HttpClient?
intermediate
Redirects: followRedirects(Redirect.NORMAL) follows them except HTTPS to HTTP downgrades; the default is NEVER. Cookies need cookieHandler(new CookieManager()). Authentication uses .authenticator(...) for Basic/Digest on 401/407.
Setting an Authorization header manually is required for pre-emptive Basic auth, since the authenticator only runs after a challenge.
Redirect.ALWAYS also follows HTTPS to HTTP.
⚠ Follow-up traps
Why does the client return a 302 response instead of the final page? The default is not to follow redirects.
Is 307/308 redirect method-preserving? Yes; 301/302 may change POST to GET in clients.
#httpclient#redirects#cookies
Q46
What is the Happy Eyeballs and IPv4 vs IPv6 handling in Java?
advanced
InetAddress resolution returns both A and AAAA records; Java sorts by preference (java.net.preferIPv6Addresses, default false so IPv4 first) and Socket.connect tries only a single address. The JDK does not implement RFC 8305 Happy Eyeballs; Netty and others add it separately.
A broken IPv6 route with AAAA first causes long connect timeouts.
URL/HttpClient may try only the first resolved address, so failures on a dead first address surface as errors rather than falling back.
⚠ Follow-up traps
Does preferIPv4Stack equal preferIPv6Addresses=false? No; the first removes IPv6 support from the JVM, the second only reorders.
What does InetAddress.getByName("localhost") return? Typically 127.0.0.1 or ::1 depending on settings.
#ipv6#dns#dual-stack
Q47
What are common HTTP status code families and the semantics of idempotent and safe methods?
basic
1xx informational, 2xx success, 3xx redirect, 4xx client error, 5xx server error. Safe methods (GET, HEAD, OPTIONS) do not change state; idempotent methods (those plus PUT, DELETE) can be repeated with the same effect.
502 Bad Gateway: proxy got an invalid upstream response; 503: overloaded or unavailable; 504: upstream timeout.
POST and PATCH are not idempotent by definition, so automatic retries are risky.
429 with Retry-After is the standard throttling signal.
⚠ Follow-up traps
Is DELETE idempotent if the second call returns 404? Yes, idempotence concerns server state, not the response code.
Is GET with a side effect allowed? Violates semantics; crawlers and prefetchers will trigger it.
#http#status-codes#idempotency
Q48
How does HTTP caching work (Cache-Control, ETag, conditional requests)?
intermediate
Cache-Control: max-age lets a cache reuse a response without contacting the origin. After expiry, a conditional request with If-None-Match (ETag) or If-Modified-Since returns 304 Not Modified with no body when unchanged.
no-cache means revalidate before use; no-store means never store.
private restricts to browser caches; public allows shared caches.
Vary lists request headers that make cached variants distinct.
⚠ Follow-up traps
Does no-cache mean not cached? No, that is no-store.
Is a 304 body-less? Yes, the client reuses its stored body.
#http-caching#etag#cache-control
Q49
How do HTTP chunked transfer encoding and Content-Length differ?
intermediate
Content-Length declares the body size up front. Transfer-Encoding: chunked sends the body as length-prefixed chunks ending with a zero-length chunk, used when the size is unknown (streaming responses).
A wrong Content-Length leads to truncated bodies or connection desync.
Disagreement between proxy and backend over these headers underlies HTTP request smuggling.
HTTP/2 has no chunked encoding; it frames data natively.
⚠ Follow-up traps
Can both headers appear? They must not; the framing becomes ambiguous and chunked takes priority in HTTP/1.1.
How does the client know when a keep-alive response ends without Content-Length? Chunked terminator; otherwise it relies on connection close.
#http#chunked#content-length
Q50
What file descriptors and OS limits matter for network servers?
intermediate
Each socket consumes a file descriptor; the per-process limit (ulimit -n, often 1024 by default) caps concurrent connections. Exceeding it throws SocketException: Too many open files.
Raise with ulimit -n, systemd LimitNOFILE, or container settings.
Count with ls /proc/<pid>/fd | wc -l or lsof -p.
Leaks (unclosed sockets, response bodies) show as growing CLOSE_WAIT counts.
⚠ Follow-up traps
Does accept() fail by itself under FD exhaustion? Yes, it returns an error repeatedly; the listener may spin and log in a tight loop.
Where do file-descriptor limits get silently low? Docker/systemd defaults that ignore shell limits.
#file-descriptors#ulimit#linux
Q51
What is CLOSE_WAIT and what does a pile-up of them mean?
intermediate
CLOSE_WAIT means the peer sent FIN and your application has not yet closed its socket. A growing count means your code is leaking connections.
Typical causes: missing close() on error paths, forgetting to close response bodies, a pool not evicting dead connections.
Unlike TIME_WAIT, it never times out by itself; the OS keeps the socket until the process closes it.
Diagnose with ss -tanp state close-wait and a thread/heap dump.
⚠ Follow-up traps
Who is responsible for CLOSE_WAIT sockets, the remote or local? The local application that did not call close().
Do they disappear after a while? No, only when the process closes them or exits.
#close-wait#tcp-states#leaks
Q52
What is the difference between `Connection reset`, `Connection reset by peer`, and `Broken pipe`?
advanced
All indicate the peer sent TCP RST. Connection reset is thrown on a read when an RST arrives; Broken pipe (EPIPE) occurs on a write after the connection was already closed/reset; Connection reset by peer is the OS wording of the same RST.
Common triggers: server closed an idle keep-alive connection, the process crashed, a firewall/NAT dropped state, a LB timeout, or sending data to a closed socket.
Fix: retry idempotent calls on stale pooled connections, evict idle connections earlier than server timeout, check LB idle timeouts.
⚠ Follow-up traps
Does close() always send FIN? Not with SO_LINGER 0 (RST), nor if unread data remains in the receive buffer (RST on Linux).
Is SocketException: Socket closed the same? No, that means your own code closed the socket locally.
#connection-reset#broken-pipe#exceptions
Scenarios
Q53
A service calls a downstream API with no timeouts configured. The downstream stalls. What happens?
basic
Calling threads block indefinitely in read, the worker pool drains, and the service stops serving unrelated requests: a cascading failure from one slow dependency.
Defaults: HttpURLConnection and Socket have infinite connect/read timeouts (connect limited only by the OS); JDK HttpClient has none either.
Fix: set connect, read/request and pool-acquire timeouts, add a circuit breaker and bulkheads.
Confirm with a thread dump: many threads in SocketInputStream.socketRead0 / NioSocketImpl.read.
⚠ Follow-up traps
Will the TCP stack eventually fail it? Only if the peer dies and keep-alive probes fire (hours by default); a live-but-stalled peer never errors.
Is a larger thread pool the fix? No, it only delays exhaustion.
#timeouts#hang#thread-exhaustion
Q54
`socket.setSoTimeout(5000)` is set, but a call still takes 4 minutes. Why?
intermediate
SO_TIMEOUT applies per blocking read, not to the total. If the server trickles a byte every few seconds, each read succeeds before 5 s and the call continues.
Also the timeout does not cover connect, DNS lookup, TLS handshake or pool wait.
Enforce a total deadline: track elapsed time between reads, use Future.get(timeout) and close the socket, or a higher-level client with a request timeout.
Closing the socket from another thread makes the blocked read throw SocketException.
⚠ Follow-up traps
Does SocketTimeoutException mean the server failed? It means no data arrived in time; the server may still be processing and complete the side effect.
Does the timeout apply to writes? No, SO_TIMEOUT applies to reads and accept only.
#timeouts#read-timeout#slow-response
Q55
Connecting to a host that silently drops packets: what do you see with and without a connect timeout?
basic
Without a timeout, connect retries SYNs with exponential backoff and fails with ConnectException: Connection timed out after about 2 minutes on Linux (tcp_syn_retries=6). With connect(addr, 3000) you get SocketTimeoutException: Connect timed out after 3 s.
Silent drops (firewall DROP, security group) cause timeouts; an active reject or closed port gives Connection refused immediately.
Distinguish them: refused means host reachable but nothing listens; timed out means packets vanish.
⚠ Follow-up traps
What does "connection refused" tell you about the network path? The host answered with RST, so routing and firewalls allow the packet.
Does Connect timed out vs Connection timed out matter? The first is the Java-level timeout, the second is the OS-level timeout.
#connect-timeout#firewall
Q56
A client sometimes fails with `Connection reset` on the first request after a quiet period. Diagnose.
intermediate
A pooled keep-alive connection was closed by the server or an intermediary (idle timeout) while the client still considered it open. The next write hits a dead connection and gets RST.
Fix: set client idle eviction/max-idle shorter than the server and LB idle timeout, enable validate-after-inactivity (Apache setValidateAfterInactivity), or retry idempotent requests once.
Example: AWS ALB idle timeout 60 s; server Tomcat keepAliveTimeout 20 s; client must evict before 20 s.
Timing pattern (failures after N seconds of idle) is the signature.
⚠ Follow-up traps
Why not just retry POST? The server might have processed it; retry only if the connection failed before sending or the request is idempotent.
Would TCP keep-alive probes fix this? Only if they keep the connection active within the intermediary's idle timeout.
#keep-alive#stale-connection#connection-reset
Q57
After deploying a new version behind DNS name `api.internal`, some clients keep calling the old IP for hours. Why?
intermediate
The JVM caches the lookup (infinite if a security manager is present or networkaddress.cache.ttl=-1), and pooled connections keep using the old address.
Fix: set networkaddress.cache.ttl to 10-60 s via java.security or Security.setProperty before the first lookup, set max connection lifetime on pools, and verify with jcmd / test.
For long-lived connections (gRPC, DB), use client-side re-resolution or periodic reconnect.
Does setting it after the first lookup work? The property is read once, so set it at startup before any InetAddress use.
Does lowering DNS record TTL alone fix it? No, the JVM ignores record TTLs.
#dns#networkaddress.cache.ttl#failover
Q58
Server logs show thousands of TIME_WAIT sockets on a client box and `Cannot assign requested address`. What do you do?
intermediate
The client opens a new connection per request and closes it, consuming ephemeral ports faster than TIME_WAIT frees them. Fix the root cause by reusing connections.
Use a shared pooled client with keep-alive; HTTP/2 multiplexing reduces connections further.
Check ss -s, sysctl net.ipv4.ip_local_port_range.
Stopgaps: widen port range, net.ipv4.tcp_tw_reuse=1 (client side outgoing only), spread over more destination IPs.
Avoid SO_LINGER 0 hacks that send RST: data loss risk.
⚠ Follow-up traps
Will lowering tcp_fin_timeout shorten TIME_WAIT? No, TIME_WAIT is fixed at 60 s on Linux; tcp_fin_timeout governs FIN_WAIT_2.
Which side should close first to avoid TIME_WAIT on the server? The client; it then holds the state.
#time-wait#port-exhaustion#pooling
Q59
A thread dump shows 200 threads in `PoolingHttpClientConnectionManager.leaseConnection`. What is wrong?
intermediate
Threads wait for a pooled connection, so the pool is exhausted. Either demand exceeds maxPerRoute or connections are leaked because response entities are not consumed/closed.
Check pool stats (getTotalStats(): leased/available/pending).
Always close the response in try-with-resources or use EntityUtils.consume.
Set a connectionRequestTimeout so callers fail fast instead of waiting forever.
Default maxPerRoute is 2 in Apache 4.x, which is rarely enough.
⚠ Follow-up traps
Why does the leak occur only on error paths? Early returns or exceptions skip the close; use try-with-resources.
Is increasing max connections the fix if leased count never decreases? No, a leak will exhaust any size.
#connection-pool#leak#thread-dump
Q60
A POST to a payment API times out on the client. Is it safe to retry?
intermediate
Not blindly. A timeout means the outcome is unknown: the server may have processed it. Retrying a non-idempotent POST can double charge.
Use an idempotency key header that the server de-duplicates on.
Otherwise query the status of the original operation first.
Retry with exponential backoff and jitter, with a cap on attempts and total deadline.
⚠ Follow-up traps
If the connection failed during connect, is retry safe? Yes, the request was never sent.
Does a 503 response guarantee no processing? Not guaranteed; follow the API's contract.
#idempotency#retry#timeouts
Q61
Latency spikes of exactly ~40 ms appear in a request/response protocol that writes header then body separately. Cause?
advanced
Nagle's algorithm holds the second small write until the first is acked, while the receiver delays its ACK (up to ~40 ms) hoping to piggyback it. The interaction stalls the body.
Fix: set TCP_NODELAY, or better write header and body in one buffer / BufferedOutputStream and flush once.
Verify with a packet capture showing the pause between segments.
⚠ Follow-up traps
Would a bigger send buffer fix it? No, the delay is the protocol interaction, not the buffer size.
Is the delay present on loopback? Often yes on Linux, which makes it easy to reproduce locally.
#nagle#delayed-ack#latency
Q62
What is the output and problem in this NIO echo read loop?
advanced
The code forgets to remove the key and ignores the -1 end-of-stream result.
while (true) { selector.select(); for (SelectionKey k : selector.selectedKeys()) { SocketChannel ch = (SocketChannel) k.channel(); ByteBuffer buf = ByteBuffer.allocate(256); ch.read(buf); buf.flip(); ch.write(buf); }}
selectedKeys() is never cleared, so keys are re-processed on every loop (busy spinning on stale keys).
After the peer closes, read returns -1 forever and the key stays ready: CPU spin; you must ch.close() on -1.
write may be partial; unwritten bytes must be kept and OP_WRITE registered.
Allocating a buffer per event is wasteful.
⚠ Follow-up traps
Why does CPU hit 100% when a client disconnects? The channel stays readable (EOF) and the loop never closes it.
Does write always write the whole buffer on a non-blocking channel? No, it may write fewer bytes or zero.
#nio#selector#bug
Q63
Your NIO server hits 100% CPU with `select()` returning immediately with zero ready keys. What is this?
advanced
The historic JDK epoll spin bug (and similar misuse): select() returns 0 repeatedly, burning CPU. Older JDKs on Linux had it; Netty works around it by counting premature returns and rebuilding the selector.
Check for other causes: a registered OP_WRITE interest on an always-writable channel, or OP_CONNECT left set after connect completes.
Use a mature framework (Netty), a recent JDK, and profile the event loop thread.
⚠ Follow-up traps
Why not selectNow() in a loop? It spins by design; select() is meant to block.
How does finishConnect relate? Failing to call it and clear OP_CONNECT keeps the key ready forever.
#nio#epoll-bug#selector
Q64
A Netty handler calls a blocking JDBC query. What breaks and how do you fix it?
intermediate
The query blocks the event-loop thread, so every other channel assigned to that loop stalls (latency spikes, timeouts, heartbeats missed).
Fix: add the handler to a separate EventExecutorGroup (pipeline.addLast(group, handler)) or hand work to a dedicated thread pool and write results back via the channel.
Detect with Netty's BlockHound or logs of long event-loop tasks.
Core rule: event loop threads only do non-blocking work.
⚠ Follow-up traps
Is Thread.sleep in a handler harmless? No, it blocks every channel on that loop.
Can another thread call channel.write? Yes; Netty queues it onto the channel's event loop.
#netty#event-loop#blocking
Q65
Netty server memory grows until OOM during a slow-client download. Why?
advanced
The server writes faster than the client reads; unwritten data accumulates in the channel's outbound buffer because the app ignores channel.isWritable().
Configure WRITE_BUFFER_WATER_MARK (low/high); when the high mark passes, isWritable() turns false and channelWritabilityChanged fires.
Pause producing until writable again, or use ChunkedWriteHandler/ChunkedFile.
Also check for ByteBuf leaks with leak detection.
⚠ Follow-up traps
Does writeAndFlush block when the client is slow? No, it enqueues and returns a future.
Does setting a bigger water mark fix it? It only delays the memory problem.
#netty#backpressure#writability
Q66
A request is slow only on the first call after startup. Which network-related causes exist?
intermediate
First call pays DNS lookup, TCP connect, TLS handshake, class loading of the SSL/HTTP stack and JIT warm-up, plus possibly certificate chain validation and OCSP.
Verify by timing phases (curl -w timing variables, -Djavax.net.debug=ssl:handshake, or client event listeners in OkHttp).
Mitigate with warm-up calls at startup, connection pre-warming, and an app-level health check that hits dependencies.
Entropy-related slow SecureRandom was a classic cause on old VMs.
⚠ Follow-up traps
Why is the second call fast? The connection is pooled and DNS is cached.
Would HTTP/2 help the first call? No, only subsequent calls via multiplexing.
#cold-start#tls#dns
Q67
Calls to a service take exactly 5 seconds intermittently. What is the likely cause?
advanced
A 5 s delay matches the default DNS resolver timeout (glibc timeout:5): a lost UDP DNS packet or a parallel A/AAAA query race in resolv.conf, notoriously seen in Kubernetes with conntrack races on UDP DNS.
Check with time getent hosts name and tcpdump port 53.
Mitigations: options single-request-reopen or use-vc in resolv.conf, NodeLocal DNSCache, lower ndots, caching in JVM.
Excessive ndots:5 makes short names trigger many search-domain queries.
⚠ Follow-up traps
Why does it affect only some calls? Only lookups that hit the lost packet pay the timeout; the cache hides it later.
Does Java use its own resolver? The JDK delegates to the OS (getaddrinfo) by default.
#dns#latency#resolver
Q68
HTTPS call fails with `PKIX path building failed`. What does it mean and how do you fix it?
basic
The JVM cannot build a chain from the server certificate to a trusted root in its trust store: a private/internal CA, missing intermediate certificate, or an outdated cacerts.
Import the CA: keytool -importcert -alias corp-ca -file ca.pem -keystore $JAVA_HOME/lib/security/cacerts, or a dedicated truststore via -Djavax.net.ssl.trustStore.
The server must send the intermediate certificates too.
Do not disable validation; trust-all managers defeat TLS.
Debug with -Djavax.net.debug=ssl:handshake.
⚠ Follow-up traps
Why does it work in the browser but not Java? Browsers fetch missing intermediates (AIA); Java does not by default.
Is the fix to disable hostname verification? No, that is a different check, and disabling it is also unsafe.
#tls#truststore#pkix
Q69
`SSLHandshakeException: No subject alternative names matching IP address` is thrown. Why?
intermediate
Java verifies that the host in the URL matches a certificate Subject Alternative Name. Connecting by IP when the certificate only lists DNS names fails (CN is ignored when SANs are present).
Connect with the DNS name, or issue a certificate with an IP SAN.
For tests, override resolution (hosts file) rather than disabling verification.
A mismatched name behind a proxy or load balancer URL is the common case.
⚠ Follow-up traps
Does the CN still count? Modern Java and HTTP clients only use SANs when they exist.
How does SNI relate? The client sends the host name in ClientHello so the server picks the right cert; with an IP, no SNI is sent.
#tls#san#hostname-verification
Q70
Handshake fails with `handshake_failure` / `protocol_version` after a JDK upgrade. What changed?
advanced
Newer JDKs disable legacy protocols and ciphers via jdk.tls.disabledAlgorithms in java.security: SSLv3, TLS 1.0 and 1.1 (disabled in JDK 8u291 / 11.0.11), weak ciphers, SHA-1 certs.
Best fix: upgrade the server to TLS 1.2+.
Temporary: edit jdk.tls.disabledAlgorithms (security risk), or force protocols with -Dhttps.protocols=TLSv1.2.
Diagnose by -Djavax.net.debug=ssl:handshake and openssl s_client -connect host:443 -tls1_2.
⚠ Follow-up traps
Does TLS 1.3 enabled by default break old servers? Rarely; the client offers 1.2 too and negotiation falls back, but some middleboxes break on 1.3 ClientHello.
Does https.protocols affect HttpClient?HttpClient uses SSLParameters via .sslParameters(...); jdk.tls.client.protocols is the system property.
#tls#jdk-upgrade#cipher-suites
Q71
A health check from a load balancer passes but users see 502 errors intermittently. Where do you look?
intermediate
Classic keep-alive mismatch: the backend closes an idle connection at the same moment the LB reuses it, producing 502. The backend idle timeout must be longer than the LB idle timeout.
Example: ALB idle 60 s requires Tomcat/Netty/NGINX keep-alive > 60 s (for example 65 s).
Also check backend crashes, slow responses beyond LB timeout (504), and request size limits.
Look at LB access logs (elb_status_code vs target_status_code).
⚠ Follow-up traps
Which side should have the longer idle timeout? The backend (the closer-to-origin hop), so the LB always closes first.
Does the health check test keep-alive reuse? No; it opens fresh requests.
#load-balancer#502#keep-alive
Q72
gRPC clients send all traffic to one pod behind a Kubernetes ClusterIP service. Why and what is the fix?
advanced
ClusterIP balances per TCP connection (kube-proxy). gRPC multiplexes all calls over a single long-lived HTTP/2 connection, so one pod receives everything.
Fix: a headless service plus client-side round-robin (dns:/// target with round_robin policy), an L7 proxy/mesh (Envoy, Linkerd), or periodically recycling connections (maxConnectionAge).
Re-resolving DNS interacts with networkaddress.cache.ttl in Java.
⚠ Follow-up traps
Does adding replicas help alone? No, existing connections stay pinned; new pods get no traffic until clients reconnect.
Why is HTTP/1.1 less affected? Clients open several connections and often reconnect more.
#grpc#http2#load-balancing
Q73
Browser console shows `blocked by CORS policy` but the API works in Postman. What is happening?
basic
Postman does not enforce the same-origin policy; the browser does. The response lacks Access-Control-Allow-Origin for the page's origin, or the preflight failed.
Inspect the OPTIONS request in dev tools: status, Allow-Methods, Allow-Headers covering the headers (Authorization, Content-Type).
Fix server-side: allow the specific origin, methods and headers; errors (4xx/5xx) must also carry CORS headers.
A redirect on the preflight also fails.
⚠ Follow-up traps
Is Allow-Origin: * with cookies allowed? No; credentials require an exact origin and Allow-Credentials: true.
Can a frontend proxy avoid CORS? Yes, a same-origin dev proxy removes cross-origin requests entirely.
#cors#browser#debugging
Q74
What happens with this Spring CORS config, and what is wrong?
intermediate
allowedOrigins("*") together with allowCredentials(true) is rejected at startup/runtime in Spring (IllegalArgumentException), because browsers refuse wildcard origin with credentials.
Use allowedOriginPatterns("https://*.example.com") or an explicit list.
With Spring Security, enable http.cors(Customizer.withDefaults()) so the CORS filter runs before authentication.
Reflecting any Origin with credentials enabled is a security hole.
⚠ Follow-up traps
Does @CrossOrigin on a controller work if Security rejects OPTIONS first? No, the filter must handle preflight before authorization.
Is a wildcard subdomain allowed in allowedOrigins? Not in that method; use allowedOriginPatterns.
#cors#spring#security
Q75
WebSocket connections drop every 60 seconds in production but not locally. Why?
intermediate
A proxy or load balancer closes idle connections after its idle timeout (NGINX proxy_read_timeout 60 s, ALB 60 s default). Locally nothing sits in between.
Fix: send application or protocol ping/pong frames more often than the timeout, and raise the proxy timeout for upgrade routes.
The NGINX config must pass Upgrade and Connection headers and use HTTP/1.1 to the backend.
Add client-side reconnect with exponential backoff and jitter.
⚠ Follow-up traps
Does TCP keep-alive replace ping frames? Rarely; its default interval is far above LB timeouts.
Why does 101 never arrive behind NGINX? Upgrade headers are not forwarded unless configured.
#websocket#idle-timeout#load-balancer
Q76
Scaling a WebSocket chat to 3 instances: users on different instances don't receive each other's messages. Why and how to fix?
advanced
Each instance holds only its own connections in memory. A message published on instance A is never delivered to sockets on B.
Fix: a shared pub/sub backbone (Redis pub/sub, Kafka, RabbitMQ with Spring STOMP broker relay) so every instance fans out to its local sessions.
Load balance with least-connections; sticky sessions are needed only for fallback transports like SockJS.
Store presence in a shared store with TTLs; handle reconnect and resume with message IDs.
⚠ Follow-up traps
Do sticky sessions solve cross-user delivery? No, users still connect to different nodes.
What happens on deploy? All sockets drop; clients must reconnect with jitter to avoid a thundering herd.
#websocket#scaling#pubsub
Q77
A client downloads a large response using `HttpResponse.BodyHandlers.ofString()` and the service runs out of memory. Fix?
intermediate
ofString buffers the whole body in memory. For large payloads stream it: ofInputStream, ofLines, or ofFile.
HttpResponse<Path> r = client.send(req, HttpResponse.BodyHandlers.ofFile(Path.of("out.bin")));System.out.println(r.body()); // out.bin
With ofInputStream, close the stream in try-with-resources.
Limit expected sizes (Content-Length check) to avoid untrusted huge bodies.
⚠ Follow-up traps
Does ofLines stream lazily? Yes, a Stream<String> that must be closed.
Does a reactive subscriber give backpressure? Yes, BodySubscriber demand via request(n).
#httpclient#streaming#memory
Q78
A scheduled job fires 500 parallel `client.send` calls to the same host and hangs. What might be limiting it?
advanced
Possible limits: the pool max per route (Apache), the server's HTTP/2 MAX_CONCURRENT_STREAMS (JDK client queues extra streams or opens more connections), the server worker threads, or the ephemeral port/FD limits.
JDK HttpClient with HTTP/1.1 opens up to many connections; use jdk.httpclient.maxstreams and a semaphore on the caller to cap concurrency.
Backpressure in the producer: bound concurrency (for example Semaphore(50)) rather than blasting everything.
Look at the server side queue and thread dump of the client.
⚠ Follow-up traps
Does HTTP/2 mean unlimited parallelism? No, concurrency is capped by the server's advertised stream limit (often 100).
Is more parallelism always faster? No, past the backend capacity it only adds queueing and timeouts.
#httpclient#http2#concurrency
Q79
A reverse proxy returns 413 for uploads and 504 for slow reports. What settings does each point to?
basic
413 Payload Too Large: body exceeds the proxy limit (NGINX client_max_body_size, default 1 MB), or the app server (Spring spring.servlet.multipart.max-file-size). 504 Gateway Timeout: upstream did not answer within the proxy's read timeout (NGINX proxy_read_timeout 60 s).
Fix 504 by making slow work asynchronous (return 202 and poll) rather than only raising timeouts.
Keep timeouts ordered: client > proxy > app > downstream per hop to avoid duplicate work.
⚠ Follow-up traps
If the proxy times out, does the backend stop working? Not necessarily; the request may continue and finish after the client is gone.
Where is the multipart limit enforced first? The earliest hop in the path with a smaller limit.
#proxy#413#504
Q80
Requests via a corporate proxy hang for HTTPS but work for HTTP. Why?
intermediate
HTTPS via a forward proxy uses CONNECT host:443. The JVM may only have http.proxyHost set (not https.proxyHost), or the proxy blocks CONNECT/needs authentication (407), or HttpClient was never given a ProxySelector.
Set both http.* and https.* properties, http.nonProxyHosts.
Java disables Basic auth for HTTPS tunneling by default: jdk.http.auth.tunneling.disabledSchemes="" to allow it.
JDK HttpClient needs .proxy(ProxySelector.getDefault()) and an Authenticator.
⚠ Follow-up traps
Can the proxy read HTTPS payloads? Not through CONNECT unless it performs TLS interception with its own CA.
Why the 407 error with Basic? The tunneling auth scheme is disabled by default since 8u111.
#proxy#connect-tunnel#https
Q81
Your server accepts connections but clients see `Connection refused` during deployment bursts. What is going on?
advanced
If nothing is listening (process restarting), the kernel sends RST: refused. If the process listens but accept queue overflows, behaviour depends on tcp_abort_on_overflow (default 0: drop, causing client timeouts and SYN retries).
Use graceful shutdown with readiness gates (deregister from LB, drain in-flight, then stop).
Increase backlog/somaxconn, and make the accept loop fast.
Check netstat -s | grep -i listen for overflow counters.
⚠ Follow-up traps
Why would refused connections appear on a healthy-looking port? The new instance has not yet bound the port, or an old one closed while the LB still routed.
Does Spring Boot support graceful shutdown? Yes, server.shutdown=graceful.
#backlog#listen-queue#deployment
Q82
A server thread reads with `InputStream.read` and the client crashes without closing. What does the server see?
intermediate
If the client machine died (power loss, cable pull) no FIN/RST is sent, so the server's read blocks indefinitely: a half-open connection. A process crash differs, because the OS sends FIN/RST.
Does Socket.isConnected() reveal a dead peer? No, it only says a connection was made at some time.
Does isClosed()? No, it reflects only the local close state.
#half-open#dead-peer#keep-alive
Q83
A client sends a request and the server replies but the client sees `SocketTimeoutException: Read timed out`. What are the two checks?
intermediate
Check whether the server actually responded in time, and whether the delay is before or after the first byte: use tcpdump/Wireshark on both ends, server access logs with request duration, and curl -w timings.
Server slow (GC pause, DB, thread pool queue): response leaves late.
Network/proxy slow: server log shows fast, client sees slow; look at LB, retransmissions (ss -ti), MTU.
Client slow: its own GC pause or exhausted executor before reading.
⚠ Follow-up traps
Can a client GC pause cause a read timeout? Yes, a long stop-the-world can exceed the timeout before the thread reads.
How to read the timing breakdown from curl?-w "%{time_namelookup} %{time_connect} %{time_appconnect} %{time_starttransfer} %{time_total}".
#read-timeout#debugging#tcpdump
Q84
What does `ss -tan state established '( dport = :443 )'` and `lsof` help find in a connection-leak investigation?
intermediate
They list sockets by state and peer, so you can see which remote endpoints a JVM holds many connections to and in what state (ESTABLISHED, CLOSE_WAIT, TIME_WAIT).
Growth in ESTABLISHED to one host: unreleased pool leases or missing close.
Growth in CLOSE_WAIT: peer closed, app did not.
Combine with jcmd <pid> Thread.print and heap dump histogram of Socket/SocketImpl objects.
⚠ Follow-up traps
Does a heap dump show leaked sockets? Yes, many live NioSocketImpl/SocketChannelImpl instances with their stack-tracked owners.
Why check the process cap? FD limit reached leads to Too many open files.
#ss#lsof#diagnostics
Q85
Throughput over a WAN link is far below bandwidth with a single connection. What can you do?
advanced
Throughput is limited by window/RTT. With 100 Mbit/s and 100 ms RTT the bandwidth-delay product is about 1.25 MB, so a 64 KB window yields about 5 Mbit/s.
Enable window scaling (default on) and let autotuning grow buffers; avoid hard-coding setReceiveBufferSize low values.
Use parallel connections or HTTP/2 multiplexing to fill the pipe, or a congestion algorithm such as BBR.
Check packet loss with ss -ti retransmit counters.
⚠ Follow-up traps
Does setting SO_RCVBUF after connect take effect for window scaling? It must be set before connect (client) or listen (server).
Does adding more bandwidth help if RTT is the limit? Only when the window can cover the larger BDP.
#tcp#bdp#throughput
Q86
What does this blocking server do when two clients connect at once?
basic
It serves them one after the other: the second client connects (the kernel completes the handshake) but is not served until the first finishes, because handle runs on the accept thread.
try (ServerSocket ss = new ServerSocket(9000)) { while (true) { try (Socket s = ss.accept()) { handle(s); // blocks while client 1 is served } }}
Client 2 may appear connected (writes succeed into the kernel buffer) yet gets no response.
Fix: hand each socket to an executor (or virtual threads on Java 21: Executors.newVirtualThreadPerTaskExecutor()).
⚠ Follow-up traps
Does the second client get Connection refused? No, until the accept queue overflows it is queued.
Is an unbounded thread-per-connection fix safe? No, cap concurrency or use virtual threads with a semaphore.
#serversocket#blocking-io#concurrency
Q87
A read loop uses `in.read(buf)` and treats each return as a full message, but messages sometimes arrive merged or split. What happens?
basic
Parsing breaks randomly: TCP can coalesce two writes into one read or split one write into several. It appears to work on localhost with small messages and fails under load or on real networks.
Fix with length-prefixed framing and DataInputStream.readFully:
int len = in.readInt();if (len < 0 || len > 1_000_000) throw new IOException("bad frame");byte[] body = new byte[len];in.readFully(body);
⚠ Follow-up traps
Does it work reliably if every message is small? No, segmentation and Nagle can still merge them.
What does read return at end of stream? -1.
#framing#tcp#bug
Q88
What exception arises when the client closes the connection and the server keeps writing?
intermediate
The first write after the peer's FIN usually succeeds (data goes out, peer answers RST). The next write fails with SocketException: Broken pipe (Linux) or Connection reset by peer/Connection reset.
Seen in web apps as ClientAbortException / AsyncRequestNotUsableException when the browser navigates away mid-response.
Handle by treating them as client disconnects (log at debug), stop producing, and release resources.
⚠ Follow-up traps
Why does the first write after close not fail? TCP is half-duplex-closed from the peer's view; the error only arrives with the RST reply.
Should these exceptions trigger 5xx alerts? No, they are client-caused.
#broken-pipe#connection-reset#write
Q89
In `HttpClient`, you call `sendAsync(...).get(2, TimeUnit.SECONDS)` and it times out. Is the request cancelled?
advanced
get(timeout) only stops waiting and throws TimeoutException; the request keeps running, holding the connection. You must call future.cancel(true) or use orTimeout and cancel.
var f = client.sendAsync(req, BodyHandlers.ofString());try { f.get(2, TimeUnit.SECONDS);} catch (TimeoutException e) { f.cancel(true);}
Cancellation closes the stream for HTTP/2 or the connection for HTTP/1.1 (reliable since JDK 16).
The more direct approach is HttpRequest.timeout.
⚠ Follow-up traps
Does orTimeout cancel the HTTP exchange? It completes the future exceptionally but does not propagate cancellation to the exchange by itself.
Is the server told to stop? Only indirectly through connection close or RST_STREAM.
#httpclient#cancellation#async
Q90
A Spring `RestTemplate` / `WebClient` call works but each request takes a new TCP+TLS handshake. Why?
intermediate
The default SimpleClientHttpRequestFactory (backed by HttpURLConnection) reuses connections only if responses are fully read and closed and keep-alive allowed. Often the response is not consumed, or Connection: close is sent, or a new client is created per call.
Configure a pooled factory (Apache HttpComponentsClientHttpRequestFactory, JDK JdkClientHttpRequestFactory), share one bean.
WebClient (Reactor Netty) pools by default but dedicates a ConnectionProvider; set max idle time below the server idle timeout.
Verify handshakes with packet capture or -Djavax.net.debug=ssl:handshake counts.
⚠ Follow-up traps
Does RestTemplate pool by default? Not with the simple factory; it relies on HttpURLConnection keep-alive cache (http.maxConnections, default 5 per host).
Why create the client as a bean? One pool shared across calls.
#spring#connection-reuse#tls
Q91
A service behind NGINX logs all client IPs as 10.0.0.5. Why and how to correct it?
basic
The remote address seen by the app is the proxy's. Use X-Forwarded-For (set by NGINX with proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for) and tell the framework to trust it.
Spring Boot: server.forward-headers-strategy=framework and configure server.tomcat.remoteip.internal-proxies.
Do not trust the header from untrusted sources, or rate limiting and IP allowlists are bypassed.
⚠ Follow-up traps
Can request.getRemoteAddr() be trusted after enabling it? Only if the chain of proxies is trusted.
What about HTTPS detection? Use X-Forwarded-Proto, or redirect loops/mixed-content links result.
#proxy#x-forwarded-for#client-ip
Q92
A Java client times out reaching an internal service by name but `curl` works from the same host. Why?
advanced
Differences between curl and the JVM: JVM proxy properties (-Dhttp.proxyHost applied to an internal host without nonProxyHosts), IPv6 vs IPv4 preference (AAAA resolves but is unreachable), a stale DNS cache entry, or a different trust/SNI setup.
Compare InetAddress.getAllByName(host) with dig/getent.
Test with -Djava.net.preferIPv4Stack=true and -Dhttp.nonProxyHosts=*.internal.
Run -Djava.net.debug=all-style debugging (-Djdk.httpclient.HttpClient.log=errors,requests,headers) to see the resolved address and proxy.
⚠ Follow-up traps
Does curl try IPv6 and IPv4 in parallel? Yes (Happy Eyeballs), while Java's Socket tries one address per connect.
Where might the JVM pick up a proxy silently? System properties, JAVA_TOOL_OPTIONS, or the default ProxySelector with OS settings.
#dns#ipv6#proxy
Q93
You see `SSLException: Connection reset` during handshake only against one server. What could be the cause?
advanced
The server or a middlebox is rejecting the ClientHello: unsupported protocol version or cipher list, missing/incorrect SNI, client-certificate required (mutual TLS) but not offered, or a WAF resetting suspicious clients.
Check -Djavax.net.debug=ssl:handshake for which messages were sent and the last one before reset.
Send SNI: with SSLParameters.setServerNames; the HttpClient does it automatically from the URI host.
⚠ Follow-up traps
Does Java send SNI when connecting by IP? No, SNI requires a host name.
Does mutual TLS failure always show a clear alert? Often only a reset, depending on server software.
#tls#handshake#sni
Q94
How do you set up mutual TLS in a Java client?
advanced
Load the client key and certificate into a KeyManagerFactory, the trusted CAs into a TrustManagerFactory, build an SSLContext, and hand it to the client.
Passing null for trust managers uses the default cacerts; supply a custom TrustManagerFactory for private CAs.
⚠ Follow-up traps
Which store holds the private key, trust store or key store? Key store; the trust store holds CA/peer certificates.
Do certificate rotations need a restart? The SSLContext caches keys; rebuild the context/client on rotation.
#mtls#keystore#sslcontext
Q95
Clients get 503s whenever one backend is overloaded, even though others are idle. Which load-balancing choice is likely wrong?
intermediate
Plain round-robin ignores per-request cost and backend slowness; a slow instance accumulates in-flight requests. Use least-outstanding-requests or power-of-two-choices, outlier detection that ejects slow/failing hosts, and bounded retry budgets.
Retries amplify overload (retry storms); cap retries and use backoff.
Active health checks should reflect real capacity, not just process liveness.
Slow-start for new instances prevents a cold JVM being flooded.
⚠ Follow-up traps
Does sticky sessions help balance? It does the opposite; it pins load to hosts.
Why can retries cause outages? Each failure multiplies traffic on already-struggling hosts.
#load-balancer#least-connections#health-checks
Q96
Client threads hang forever in `SocketInputStream.socketRead0` during a rolling restart of a database or service. Why, and what prevents it?
advanced
When the peer host vanishes without closing (killed VM, dropped route, NAT entry expired), the client's connection is half-open and a read with no timeout blocks until TCP keep-alive or retransmission failure (up to ~15 minutes for retransmission, hours for keep-alive).
Prevent: always set read timeouts, tune tcp_keepalive_* or TCP_KEEPIDLE options, and tcp_retries2 / TCP_USER_TIMEOUT for faster failure detection.
For JDBC set socketTimeout/tcpKeepAlive; for pools use validation queries.
⚠ Follow-up traps
Why does a write to a half-open connection not fail immediately? The data is buffered and retransmitted until the retry limit gives up.
Does TCP keep-alive detect a dead peer quickly by default? No, first probe after 7200 s.
#hang#half-open#keep-alive
Q97
A scheduled job opens a `new URL(url).openConnection()` per run and eventually hits `Too many open files`. What do you check?
basic
Streams and connections were not closed on all paths (including error streams from getErrorStream()), so FDs leak.
Close getInputStream() or getErrorStream(), use try-with-resources, and call disconnect() if you do not want keep-alive reuse.
Verify growth with lsof -p <pid> | wc -l.
Raise ulimit -n only after the leak is fixed.
⚠ Follow-up traps
Does HTTP 404 throw on getInputStream()? Yes (FileNotFoundException); read getErrorStream() for the body.
Does disconnect() always close the socket? It closes it if the connection is not in the keep-alive cache; otherwise it may be returned for reuse after drained.
#leak#file-descriptors#httpurlconnection
Q98
A UDP-based metrics sender loses data under load. How do you reason about it?
intermediate
UDP offers no delivery guarantee: drops happen when the receiver socket buffer fills (SO_RCVBUF), at routers, or when datagrams exceed the MTU and fragment.
Keep payloads under ~1400 bytes to avoid fragmentation.
Raise SO_RCVBUF and the OS cap net.core.rmem_max; check netstat -su for receive errors.
If loss is unacceptable, use TCP or add sequence numbers, acks and retries.
⚠ Follow-up traps
Does DatagramSocket.send throw when the packet is lost? No, it only reports local errors.
What is the max UDP payload? 65,507 bytes over IPv4, but anything above the MTU fragments.
#udp#datagram#packet-loss
Q99
What happens with `Socket.connect` to `localhost` when the server listens only on IPv6 `::1`?
advanced
localhost may resolve to both 127.0.0.1 and ::1; Java chooses by preferIPv6Addresses (default IPv4 first) and tries only that address, so a service bound to ::1 yields Connection refused on 127.0.0.1.
Bind to 0.0.0.0 / :: (dual-stack), or connect using the explicit loopback the service uses.
Check with ss -ltn for the bound address.
-Djava.net.preferIPv6Addresses=true flips the order.
⚠ Follow-up traps
Does binding new ServerSocket(port) listen on IPv4 and IPv6? On dual-stack systems it binds the wildcard ::, accepting both unless preferIPv4Stack is set.
Does "works in browser" prove the Java path works? No, browsers try both addresses.
#ipv6#localhost#dual-stack
Q100
How would you implement graceful shutdown for a Netty or socket server?
intermediate
Stop accepting new connections, deregister from the load balancer (readiness off), let in-flight requests finish within a deadline, then close channels and release event loops.
shutdownGracefully(quietPeriod, timeout) waits for a quiet period with no tasks before terminating.
For HTTP/1.1 send Connection: close on the final responses; for HTTP/2 send GOAWAY.
Kubernetes: handle SIGTERM and use a preStop sleep to let endpoints deprogram.
⚠ Follow-up traps
Is System.exit on SIGTERM enough? It drops in-flight requests; the shutdown hook must drain.
What is GOAWAY? An HTTP/2 frame telling the client to open no new streams on that connection.
#graceful-shutdown#netty#draining
Q101
A client downloads files and each one stalls at exactly the same offset behind a VPN. What network concept is involved?
advanced
An MTU/PMTU black hole: large packets with the DF bit set are dropped because a tunnel reduces the MTU and ICMP "fragmentation needed" messages are blocked. Small packets (handshakes, headers) pass; the first full-size data segment never arrives.
Diagnose with ping -M do -s 1400 host and varying sizes, tracepath.
Fixes: lower the interface MTU, enable TCP MSS clamping on the gateway, allow ICMP type 3 code 4.
⚠ Follow-up traps
Why does the TLS handshake sometimes hang? The server certificate flight is large and gets dropped.
Does a Java setting fix it? No, it is a network-layer issue.
#mtu#path-mtu#black-hole
Q102
A WebFlux/Reactor Netty client gets `PrematureCloseException: Connection prematurely closed BEFORE response`. Cause?
advanced
Reactor Netty sent a request on a pooled connection that the server had already closed, or the server closed during processing. It is the same stale keep-alive problem as in blocking clients.
Set ConnectionProvider.builder("p").maxIdleTime(Duration.ofSeconds(20)).maxLifeTime(...).evictInBackground(Duration.ofSeconds(30)) below the server idle timeout.
Retry idempotent requests with retryWhen on this exception.
Confirm with server logs that it did not crash or hit max request limits.
⚠ Follow-up traps
Why does evictInBackground matter? Without it, expiry is checked only on acquire, so stale connections linger until used.
Should POSTs be retried automatically? Only with idempotency keys.
#reactor-netty#webclient#stale-connection
Q103
How do you choose between thread-per-request (virtual threads), reactive/Netty, and async HttpClient for an I/O-heavy gateway?
advanced
For Java 21+, virtual threads with blocking code give most of the scalability of reactive with simpler stack traces, and they work with existing blocking libraries. Reactive/Netty still wins for streaming, strict backpressure and very high connection counts with minimal memory.
Virtual threads: simple code, but cap concurrency to protect downstreams (semaphore/bulkhead); watch pinning in synchronized blocks on older JDKs (fixed in 24).
Reactive: operators for backpressure, but harder debugging and colored functions.
Pool limits, timeouts and circuit breakers are needed in all models.
⚠ Follow-up traps
Do virtual threads remove the need for connection pools? No, remote resources like DB or HTTP connections remain finite.
Do virtual threads speed up CPU-bound work? No, only blocking I/O concurrency.
#design#virtual-threads#reactive
Q104
Service A calls B calls C; the user sees timeouts, yet C is fine. How do you set timeouts across the chain?
intermediate
Each upstream timeout must exceed the sum of what it waits for downstream, and a total deadline must propagate. If A's timeout is shorter than B's, A gives up while B (and C) keep working, wasting capacity.
Rule of thumb: A timeout > B timeout x attempts + overhead; budget the deadline (gRPC deadlines, X-Request-Deadline headers) and pass the remaining time down.
Limit retries to one layer to avoid multiplicative retries (3 x 3 x 3 = 27 calls).
Use circuit breakers and bulkheads; trace with correlation IDs to find the slow hop.
⚠ Follow-up traps
What if every layer retries 3 times? C sees 27x load during an incident.
Is a very large timeout safer? No, it ties up threads and defers failure detection.