Streams vs readers, buffering, NIO.2, memory-mapped files, large file processing, charsets, Java serialization and its security risks, JSON vs Protobuf/Avro, zip handling, atomic writes and real-world file failure scenarios.
Theory
Q1
What is the difference between byte streams and character streams?
basic
InputStream/OutputStream move raw bytes; Reader/Writer move chars and perform charset decoding/encoding. Use streams for binary data (images, archives) and readers for text.
InputStreamReader and OutputStreamWriter are the bridges from bytes to chars and need a charset.
FileReader and FileWriter use the platform default charset before Java 18 and UTF-8 from Java 18 (JEP 400).
⚠ Follow-up traps
Can you read a PNG with a Reader? No, decoding arbitrary bytes as text corrupts them irreversibly.
Is char one byte? No, it is a 16-bit UTF-16 code unit.
#streams#readers#io
Q2
Why should you buffer I/O?
basic
Every unbuffered read()/write() on a FileInputStream is a system call. Wrapping in BufferedInputStream/BufferedReader (default 8 KB) batches them and reduces syscalls dramatically.
BufferedWriter needs flush() or close() to push data out.
Bulk reads such as read(byte[]) with a large array already amortize syscalls; extra buffering adds little.
Files.newBufferedReader returns a buffered reader by default.
⚠ Follow-up traps
Does buffering help if you already read 64 KB chunks? Barely; it only adds a copy.
If the JVM exits without close, is buffered data written? Not guaranteed; there is no automatic flush of BufferedWriter.
#buffering#performance
Q3
What does the Decorator pattern look like in java.io?
basic
Streams wrap other streams to add behaviour: new DataInputStream(new BufferedInputStream(new FileInputStream(f))). Each layer adds one concern (buffering, primitive decoding, compression, encryption).
Closing the outermost wrapper closes the whole chain.
Order matters: put GZIPOutputStream between the file and the buffer or the buffer in front, depending on what you want compressed.
⚠ Follow-up traps
If the wrapper constructor throws, is the inner stream leaked? Yes, unless both are declared as separate resources in try-with-resources.
Do I need to close each layer? No, closing the outer one is enough.
#decorator#streams
Q4
How does try-with-resources work?
basic
Resources implementing AutoCloseable declared in the header are closed automatically in reverse order of declaration, whether the block completes normally or throws.
try (var in = Files.newInputStream(src); var out = Files.newOutputStream(dst)) { in.transferTo(out);}
If both the body and close() throw, the body exception is primary and the close exception is attached via addSuppressed.
Java 9+ allows effectively final variables: try (in).
⚠ Follow-up traps
Which closes first, in or out?out, reverse order.
How do you read the suppressed exception?Throwable.getSuppressed().
#try-with-resources#autocloseable
Q5
What are the pitfalls of the old try/finally close pattern?
intermediate
An exception from close() in finally replaces the original exception, hiding the root cause, and nested resources require nested try blocks to avoid leaks.
Try-with-resources preserves the primary exception and suppresses the rest.
A null resource is allowed in the header and simply skipped on close.
⚠ Follow-up traps
Is the resource still open in the catch block? No, it is closed before catch and finally run.
Can close() be called twice safely? For Closeable it must be idempotent, but custom classes may violate this.
#try-with-resources#exceptions
Q6
How do charsets work, and why does the default matter?
basic
A charset maps characters to bytes. Reading with the wrong charset yields garbage (mojibake) or replacement characters. Always specify one explicitly via StandardCharsets.UTF_8.
Before Java 18 the default came from file.encoding/OS locale; since Java 18 it is UTF-8 for APIs like FileReader, String.getBytes().
System.out and console use stdout.encoding (Java 19+), which can still differ.
Files.readString and Files.writeString default to UTF-8 regardless of platform.
⚠ Follow-up traps
Does JEP 400 change Files.readAllLines behaviour? No, those already defaulted to UTF-8.
Can you detect the charset of a file reliably? No, only heuristically (BOM or libraries).
#charset#encoding#utf-8
Q7
How does Java handle malformed input when decoding?
intermediate
InputStreamReader and new String(bytes, cs) replace malformed bytes with U+FFFD silently. Files.readString and Files.newBufferedReader throw MalformedInputException. To control behaviour use CharsetDecoder.onMalformedInput(CodingErrorAction.REPORT).
var dec = StandardCharsets.UTF_8.newDecoder() .onMalformedInput(CodingErrorAction.REPORT);try (var r = new BufferedReader(new InputStreamReader(in, dec))) { /* ... */ }
⚠ Follow-up traps
Why does readAllLines fail on a Latin-1 file? Strict UTF-8 decoding rejects bytes like 0xE9 not forming valid sequences.
Is the BOM stripped automatically for UTF-8? No, Java leaves U+FEFF at the start of the text.
#charset#decoder#malformed
Q8
What is the difference between java.io.File and java.nio.file.Path?
basic
File is a legacy class with poor error reporting (booleans), no symlink support and limited metadata. Path plus Files (NIO.2, Java 7) is the modern API with exceptions, attributes, symlinks, atomic moves and walking.
Convert with file.toPath() and path.toFile().
Path is immutable and filesystem-aware (zip filesystems, in-memory test filesystems).
⚠ Follow-up traps
Does File.delete() tell you why it failed? No, just false; Files.delete throws a specific exception.
Does creating a Path touch the disk? No, it is a pure name until you call a Files method.
resolve joins paths (an absolute argument replaces the base); relativize computes the path between two paths of the same type; normalize removes . and .. lexically.
readAllBytes and readString load the whole file into memory; unsuitable for big files.
Files.lines, Files.walk and Files.list return lazy streams that hold an open file handle.
⚠ Follow-up traps
Do Files.write calls truncate by default? Yes: CREATE, TRUNCATE_EXISTING, WRITE are the defaults.
What is the size limit of readAllBytes? About 2 GB (array limit).
#nio2#files
Q11
Why must Files.lines, walk and list be closed?
intermediate
These streams wrap an open file or directory handle, released only when the stream is closed. Leaking them eventually causes "Too many open files".
try (Stream<String> lines = Files.lines(p)) { long n = lines.filter(l -> l.contains("ERROR")).count();}
Files.lines wraps checked IOException in UncheckedIOException during iteration.
⚠ Follow-up traps
Does a terminal operation close the stream? No, only close() does.
Does Files.walk follow symlinks? Not unless FileVisitOption.FOLLOW_LINKS is passed.
#nio2#streams#leak
Q12
Compare Files.walk, Files.list and walkFileTree.
intermediate
list is one directory level, walk is a lazy depth-first stream to a max depth, and walkFileTree uses a FileVisitor with callbacks and control flow (SKIP_SUBTREE, TERMINATE) and per-file error handling.
Use walkFileTree for recursive delete or copy where you must handle postVisitDirectory and permission errors.
walk aborts the whole stream on an AccessDeniedException (wrapped in UncheckedIOException).
⚠ Follow-up traps
How do you delete a directory tree? Delete files in visitFile and directories in postVisitDirectory, as a directory must be empty.
Is walk order sorted? No, directory order is filesystem-dependent.
#nio2#walk#filevisitor
Q13
How does Files.copy and Files.move behave?
intermediate
By default both fail with FileAlreadyExistsException if the target exists. REPLACE_EXISTING overrides it; ATOMIC_MOVE requests an atomic rename and throws AtomicMoveNotSupportedException if the filesystem cannot guarantee it.
Copying a directory copies only the empty directory, not its content.
Moves across filesystems degrade to copy plus delete, which is not atomic.
COPY_ATTRIBUTES keeps timestamps where supported.
⚠ Follow-up traps
Does ATOMIC_MOVE honour REPLACE_EXISTING? It is implementation-specific; on POSIX rename replaces, on others it may throw.
Does Files.copy(in, path) close the stream? No.
#nio2#copy#move
Q14
How does WatchService work?
intermediate
Register a directory with a WatchService for ENTRY_CREATE, ENTRY_MODIFY, ENTRY_DELETE; then loop on take()/poll() to receive WatchKeys, process pollEvents() and call key.reset().
It uses inotify on Linux, but on macOS the JDK polls (about 10 s delay by default, tunable via SensitivityWatchEventModifier in com.sun).
Watching is not recursive; register each subdirectory yourself.
⚠ Follow-up traps
What if you forget reset()? The key is never signalled again and you stop receiving events.
Do you get one event per write? No, multiple ENTRY_MODIFY events may fire for a single logical write, and OVERFLOW can occur.
#nio2#watchservice
Q15
What is the difference between FileChannel and streams?
intermediate
FileChannel is bidirectional, supports positional reads/writes, transferTo/transferFrom, memory mapping, locking and force(). It works with ByteBuffers, optionally direct ones.
Positional read read(buf, position) is thread-safe and does not move the shared position.
ByteBuffer requires flip() between write and read modes.
Obtain via FileChannel.open(path, options) or stream.getChannel().
⚠ Follow-up traps
Why does reading return nothing after put? You forgot flip(); position equals limit.
Does read fill the buffer fully? Not guaranteed; loop until -1 or buffer full.
#nio#filechannel#bytebuffer
Q16
What is a direct ByteBuffer and when do you use it?
advanced
ByteBuffer.allocateDirect allocates native memory outside the heap, letting the OS do I/O without an intermediate copy from the heap array. It is costly to allocate and freed only when GC collects the buffer object.
Good for long-lived, reused buffers on network or file channels.
Capped by -XX:MaxDirectMemorySize (defaults to max heap); exceeding it throws OutOfMemoryError: Direct buffer memory.
Heap buffers passed to channels are copied to a temporary direct buffer internally.
⚠ Follow-up traps
Are direct buffers visible in heap dumps? Only the small wrapper; the native memory is not counted in heap usage.
Should you allocate a direct buffer per request? No, pool or reuse them.
#nio#direct-buffer#off-heap
Q17
How do memory-mapped files work?
advanced
FileChannel.map(mode, position, size) returns a MappedByteBuffer whose contents are backed by the OS page cache. Reads and writes become memory accesses; the OS faults pages in lazily and writes dirty pages back.
A single mapping is limited to Integer.MAX_VALUE bytes (2 GB) via ByteBuffer; map in windows for larger files.
force() flushes dirty pages. Unmapping is tied to GC (no public unmap) before the Java 22 MemorySegment API with Arena.
⚠ Follow-up traps
Does closing the channel unmap the buffer? No, the mapping stays valid until the buffer is garbage collected.
What happens if another process truncates the file while mapped? Access can crash with InternalError (SIGBUS).
#mmap#mappedbytebuffer#nio
Q18
When is memory mapping better or worse than normal reads?
advanced
It shines for random access to large read-mostly files and for sharing data between processes. It is not magically faster for sequential scans, where buffered FileChannel reads with a large buffer are comparable.
Page faults are hidden I/O stalls, so latency becomes unpredictable and blocking cannot be interrupted.
Many mappings consume virtual address space and are slow to release on 32-bit JVMs.
There is no error handling per read; I/O errors surface as InternalError.
⚠ Follow-up traps
Does mmap avoid the page cache? No, it uses the page cache directly (zero copy is the benefit).
Is it good for tiny files? No, mapping setup cost dominates.
#mmap#performance#trade-offs
Q19
What is FileChannel.transferTo and why is it fast?
intermediate
transferTo(position, count, target) lets the kernel move bytes directly (e.g., sendfile on Linux) between file and socket or file without copying through user space.
It may transfer fewer bytes than requested; loop until done.
Used by servers (Netty FileRegion, Tomcat) for static content delivery.
Does not work with TLS since data must be encrypted in user space (unless kTLS).
⚠ Follow-up traps
Does it always copy everything in one call? No, check the return value and loop.
Is it zero-copy over HTTPS? Generally no.
#nio#zero-copy#transferto
Q20
How do you process a very large file efficiently?
intermediate
Stream it: read line by line (or in chunks) with bounded memory, process incrementally and write results incrementally. Never use readAllBytes, readAllLines or readString on files that may not fit in the heap.
BufferedReader.readLine or Files.lines for text; FileChannel plus a reusable buffer for binary.
For parallelism, split by byte ranges aligned to newline boundaries or use Files.lines(...).parallel() cautiously (the spliterator on file-backed streams splits poorly before Java 9; improved for UTF-8 and ASCII line streams).
Memory use must be independent of file size.
⚠ Follow-up traps
Does Files.lines load the whole file? No, it is lazy.
Is parallel stream always faster for I/O-bound work? No, disk throughput is usually the limit.
#large-files#streaming#memory
Q21
What is RandomAccessFile and when is it appropriate?
intermediate
RandomAccessFile supports seek(pos) with reads and writes at arbitrary offsets in modes r, rw, rws and rwd. It suits fixed-record files and tailing files.
rws forces both content and metadata to disk on each write; rwd forces content only. Both are slow.
It is unbuffered, so wrap or use FileChannel for performance.
⚠ Follow-up traps
Is it thread-safe? No, the file pointer is shared state; use positional channel reads.
Does rw create the file? Yes, if missing.
#randomaccessfile#seek
Q22
How do you guarantee that written data reaches the disk?
advanced
flush() only moves data from Java buffers to the OS. To persist, call FileChannel.force(true) (or FileDescriptor.sync()), which issues fsync. Even then, disk write caches can lie.
Writes with StandardOpenOption.SYNC or DSYNC sync on every write.
For a new file also fsync the parent directory (not possible in pure Java on some platforms), otherwise the directory entry may be lost on crash.
Databases batch writes and fsync once per commit group for throughput.
⚠ Follow-up traps
Does close() fsync? No.
Does flush() on a FileOutputStream do anything? No, it is a no-op because it has no buffer.
#durability#fsync#force
Q23
How do you write a file atomically?
advanced
Write to a temporary file in the same directory (same filesystem), fsync it, then Files.move(tmp, target, ATOMIC_MOVE). Readers see either the old or the new file, never a partial one.
Why not create the temp file in /tmp? It may be a different filesystem, making the move a non-atomic copy.
Is rename atomic on all platforms? On POSIX yes; on Windows replacing may require REPLACE_EXISTING and can fail if the target is open.
#atomic-write#move#durability
Q24
How do temporary files work in Java?
basic
Files.createTempFile and createTempDirectory create a uniquely named file atomically, with owner-only permissions (0600 / 0700) on POSIX. The location is java.io.tmpdir.
Delete explicitly or use StandardOpenOption.DELETE_ON_CLOSE; File.deleteOnExit() keeps a list in memory that grows forever in long-running servers.
The old File.createTempFile created files with default permissions.
⚠ Follow-up traps
Is deleteOnExit safe for a server? No, it leaks memory and does not run on a crash or kill -9.
Can you guess the temp name? The random component prevents predictable-name attacks.
#temp-files#security
Q25
How does file locking work with FileLock?
advanced
FileChannel.lock() (blocking) or tryLock() returns a FileLock, either exclusive or shared, over a region. Locks are held on behalf of the whole JVM, not individual threads, and are advisory on most Unix systems.
Within one JVM a second lock on an overlapping region throws OverlappingFileLockException.
Behaviour on NFS and some network filesystems is unreliable.
Typical use is a lock file to ensure a single application instance or serialize writers across processes.
⚠ Follow-up traps
Does the lock stop non-cooperating processes from reading? Not on Unix; it is advisory.
Does it coordinate threads of the same JVM? No, use normal Java locks for that.
#filelock#concurrency#nio
Q26
How do you handle path traversal when working with user-supplied filenames?
intermediate
Resolve the user input against a base directory, normalize and verify the result still starts with the base. Prefer ignoring client paths and generating your own filenames.
Path base = Path.of("/srv/uploads").toRealPath();Path target = base.resolve(userName).normalize();if (!target.startsWith(base)) { throw new SecurityException("Path traversal: " + userName);}
startsWith on Path compares name elements, so /srv/uploads-evil does not match /srv/uploads (unlike string prefix).
Symlinks inside the base can still escape; check toRealPath on the existing target.
⚠ Follow-up traps
Is checking for ".." in the string enough? No, encodings, absolute paths and symlinks bypass it.
Why not String.startsWith?/srv/uploads2 passes the string check.
#security#path-traversal
Q27
How does Java's built-in serialization work?
basic
A class implementing the marker interface Serializable can be written by ObjectOutputStream.writeObject and rebuilt by ObjectInputStream.readObject. The runtime walks the object graph by reflection, writing class metadata and non-static, non-transient field values.
Shared references and cycles are preserved via back-references within one stream.
Constructors of serializable classes are not called on deserialization; the no-arg constructor of the first non-serializable superclass is.
⚠ Follow-up traps
Is a static field serialized? No, static fields belong to the class.
What if a field type is not Serializable?NotSerializableException at write time.
#serialization#serializable
Q28
What is serialVersionUID and what happens if you omit it?
basic
It is a static final long that versions the serialized form. On read, the stream's UID must equal the local class's UID or InvalidClassException is thrown.
If omitted, the JVM computes a hash from class name, interfaces, fields and methods; almost any change (even adding a method) changes it and breaks compatibility. Compilers may also compute it differently.
Declare it explicitly: private static final long serialVersionUID = 1L;.
Compatible changes: adding fields, adding classes in the hierarchy. Incompatible: removing a field type change, changing a class from Serializable to Externalizable.
⚠ Follow-up traps
Does keeping the same UID guarantee compatibility? No, incompatible structural changes (e.g., field type change) still fail.
What value do removed fields get when reading old data? The default value (null, 0).
#serialization#serialversionuid
Q29
What does transient do, and how do you restore transient state?
basic
transient excludes a field from default serialization. After deserialization it holds the default (null/0/false), so you must recompute it in readObject or readResolve.
Use for caches, derived values, thread handles, passwords and file handles.
Final transient fields cannot be reassigned in readObject without reflection, which makes them awkward.
⚠ Follow-up traps
Is transient honoured by Jackson? Jackson ignores the Java keyword by default unless MapperFeature.PROPAGATE_TRANSIENT_MARKER is on (it still skips transient fields as properties but does not remove getters).
Does static transient mean anything? Nothing extra; static is already excluded.
#serialization#transient
Q30
How do you customize serialization with writeObject and readObject?
intermediate
Declare private methods writeObject(ObjectOutputStream) and readObject(ObjectInputStream) with exact signatures. Call defaultWriteObject()/defaultReadObject() first, then handle extra data.
private void readObject(ObjectInputStream in) throws IOException, ClassNotFoundException { in.defaultReadObject(); if (age < 0) throw new InvalidObjectException("age"); this.cache = new HashMap<>();}
readObject is effectively a public constructor: validate invariants there.
The methods are found by reflection, so a wrong signature or visibility silently disables them.
⚠ Follow-up traps
Do you need to call defaultReadObject first? Normally yes, otherwise the fields are not read and the stream goes out of sync.
Can readObject be public? It would not be invoked; it must be private.
#serialization#readobject#writeobject
Q31
What are readResolve and writeReplace?
advanced
writeReplace substitutes another object before writing; readResolve substitutes the deserialized object with another (often the canonical instance). readResolve preserves singleton and enum-like guarantees.
private Object readResolve() { return INSTANCE; }
Without readResolve a serializable singleton can be duplicated by deserialization.
writeReplace returning a small serialization proxy is the safest pattern (Effective Java).
⚠ Follow-up traps
Why are enum singletons safe? The JVM guarantees one instance per constant on deserialization, serializing only the name.
Does readResolve run before or after readObject? After.
#serialization#readresolve#singleton
Q32
What is the serialization proxy pattern?
advanced
The outer class's writeReplace returns a private static nested SerializationProxy that holds the logical state. The outer class's readObject throws InvalidObjectException to block direct deserialization, and the proxy's readResolve rebuilds the real object via its public constructor/factory.
Invariants are enforced by the normal constructor, defeating forged streams.
Allows fields to be final and decouples the serialized form from the implementation.
java.time classes use it (Ser), as do immutable collections.
⚠ Follow-up traps
Can an attacker still supply the outer class bytes? The outer readObject rejects it.
What is the cost? A bit more code and copying.
#serialization#proxy#invariants
Q33
How does inheritance affect serialization?
intermediate
If a superclass is Serializable, so is the subclass and the superclass fields are serialized. If the superclass is not, its fields are not saved and on deserialization its no-arg constructor runs (which must be accessible, else InvalidClassException).
A subclass can prevent serialization by throwing NotSerializableException from writeObject/readObject.
Superclass state must be re-initialised via that constructor.
⚠ Follow-up traps
Is the subclass constructor called on deserialization? No.
What if the non-serializable parent has no no-arg constructor? Deserialization fails with InvalidClassException.
#serialization#inheritance
Q34
What is Externalizable and how does it differ from Serializable?
intermediate
Externalizable extends Serializable with writeExternal/readExternal, giving full control of the format. The class must have a public no-arg constructor, which is called first and then readExternal fills in the state.
Nothing is saved automatically, including superclass fields.
Can be more compact and faster, but you own versioning and correctness.
⚠ Follow-up traps
What happens without a public no-arg constructor?InvalidClassException (no valid constructor).
Are transient modifiers relevant? No, you decide what to write.
#serialization#externalizable
Q35
What are the security risks of Java deserialization?
advanced
readObject instantiates classes named in the stream and runs their readObject/readResolve/finalize-like code before the caller can check the type. With "gadget" classes on the classpath (commons-collections, Spring, etc.) an attacker can chain calls to achieve remote code execution or DoS.
The cast after readObject happens too late; the damage is done during deserialization.
Famous cases: Apache Commons Collections InvokerTransformer, many app-server CVEs.
Also DoS: deeply nested or huge object graphs, hash collision payloads.
⚠ Follow-up traps
Is it safe if you cast to your expected type? No, the gadget runs during reading.
Is it only an issue for network input? Any untrusted bytes: cookies, caches, message queues, files.
#serialization#security#gadget-chain
Q36
What are deserialization filters (JEP 290/415)?
advanced
ObjectInputFilter (Java 9+, backported to 8u121) lets you allow or reject classes and cap depth, references, array length and stream bytes while reading. Configure per stream or process-wide.
ObjectInputFilter f = ObjectInputFilter.Config.createFilter( "com.acme.dto.*;java.base/*;maxdepth=10;maxbytes=1000000;!*");ois.setObjectInputFilter(f);
Pattern order matters and !* rejects everything not allowed (allow-list).
Global: -Djdk.serialFilter=... or jdk.serialFilter in java.security.
JEP 415 (Java 17) adds context-specific filter factories (ObjectInputFilter.Config.setSerialFilterFactory).
⚠ Follow-up traps
Is a block-list enough? No, new gadgets appear; use an allow-list.
Does the filter run before the object is built? Yes, it is invoked on class resolution and as the stream is parsed.
#serialization#filters#jep-290
Q37
Why is Java serialization generally discouraged for new designs?
intermediate
It is insecure by design, couples the wire format to class internals, is fragile across versions, is Java-only, verbose and slow. Oracle has called it a "horrible mistake" and long-term plans aim to remove it.
Prefer explicit formats: JSON, Protobuf, Avro, with schemas and DTOs.
If you must use it, restrict to trusted data, use filters, declare UID and consider serialization proxies.
⚠ Follow-up traps
Is it still used internally? Yes: RMI, HTTP session replication, some caches.
Is it deprecated? Not formally, though not recommended.
#serialization#design
Q38
How do records interact with serialization?
intermediate
A record implementing Serializable is serialized from its components only, and deserialization calls the canonical constructor, so constructor validation applies. readObject, writeObject, serialVersionUID matching and Externalizable customizations are ignored or restricted.
record Point(int x, int y) implements Serializable { Point { if (x < 0) throw new IllegalArgumentException(); }}
serialVersionUID can be declared, but its matching is not required (the default is 0L unless declared).
This closes the "forged stream bypasses constructor" hole of normal classes.
Jackson 2.12+ supports records natively.
⚠ Follow-up traps
Does writeObject on a record work? No, it is ignored.
Can records implement Externalizable? The Externalizable methods are not used for records.
#records#serialization#java16
Q39
What are the main alternatives to Java serialization, and how do JSON, Protobuf and Avro compare?
intermediate
JSON is human-readable, schema-less and verbose. Protobuf is compact binary with a compiled .proto schema and field numbers. Avro is compact binary where the schema travels with the data (or a registry) and is resolved at read time.
Protobuf: fast, tag-based forward/backward compatibility, great for gRPC.
Avro: no field tags in data, needs writer's schema to read, strong for Kafka plus Schema Registry and big-data files.
⚠ Follow-up traps
Can Avro be read without the writer schema? No, the bytes do not carry field ids.
Is Protobuf always faster than JSON? Usually smaller and faster, but benchmark; JSON libs are well optimised.
#json#protobuf#avro#comparison
Q40
How does schema evolution work in Protobuf versus Avro?
advanced
Protobuf identifies fields by number: add optional fields with new numbers, never reuse or renumber, and mark deleted ones reserved. Unknown fields are skipped. Avro resolves the writer schema against the reader schema by field name: new fields need defaults, and removed fields must have had defaults for backward compatibility.
Changing a Protobuf type (e.g., int32 to string) breaks compatibility; some integer widening is wire compatible.
Avro supports aliases for renaming.
Use a registry (Confluent) with compatibility modes BACKWARD, FORWARD, FULL.
⚠ Follow-up traps
Can you rename a Protobuf field? Yes, only the number is on the wire (JSON mapping uses the name though).
Why must Avro new fields have defaults? So readers with the new schema can read old data lacking them.
#protobuf#avro#schema-evolution
Q41
How does Jackson serialize and deserialize Java objects?
basic
ObjectMapper uses reflection (or generated accessors with modules) to map getters/fields to JSON properties and constructors/setters/@JsonCreator for deserialization. Create one ObjectMapper and reuse it; it is thread-safe once configured.
ObjectMapper om = new ObjectMapper() .registerModule(new JavaTimeModule()) .disable(DeserializationFeature.FAIL_ON_UNKNOWN_PROPERTIES);String json = om.writeValueAsString(dto);Dto back = om.readValue(json, Dto.class);
⚠ Follow-up traps
Is ObjectMapper thread-safe? Yes, after configuration; do not reconfigure it concurrently.
Why is FAIL_ON_UNKNOWN_PROPERTIES on by default? To surface schema mismatches; many services disable it for forward compatibility (Spring Boot already does).
#jackson#json#objectmapper
Q42
What are the main Jackson annotations and pitfalls?
Entities with bidirectional relationships cause infinite recursion; use @JsonManagedReference/@JsonBackReference, DTOs or @JsonIdentityInfo.
Dates: without JavaTimeModule, java.time types fail (2.12+) or serialize as arrays; disable WRITE_DATES_AS_TIMESTAMPS for ISO strings.
Prefer DTOs over serializing JPA entities (lazy proxies).
⚠ Follow-up traps
Why does LocalDate throw InvalidDefinitionException? The JSR-310 module is not registered.
Does @JsonIgnore on a field hide the setter too? It ignores the whole logical property.
#jackson#annotations
Q43
Why is polymorphic deserialization in Jackson a security concern?
advanced
Enabling default typing (enableDefaultTyping, @JsonTypeInfo(use = Id.CLASS)) allows the JSON to name any class to instantiate, enabling gadget attacks similar to Java deserialization (many Jackson CVEs).
Use @JsonTypeInfo(use = Id.NAME) with explicit @JsonSubTypes.
If default typing is needed, use PolymorphicTypeValidator allow-lists (2.10+).
Keep Jackson patched and avoid Object/Serializable typed fields from untrusted JSON.
⚠ Follow-up traps
Is Id.NAME with @JsonSubTypes vulnerable? No, only listed subtypes are allowed.
Is plain JSON to a concrete POJO risky? Far less; only the declared type graph is created.
#jackson#security#polymorphism
Q44
How do you stream large JSON with Jackson?
intermediate
Use the streaming API (JsonParser/JsonGenerator) or ObjectMapper.readerFor(T).readValues(parser) for arrays so only one element is in memory at a time.
try (JsonParser p = om.createParser(in)) { if (p.nextToken() != JsonToken.START_ARRAY) throw new IOException("array"); while (p.nextToken() == JsonToken.START_OBJECT) { Event e = om.readValue(p, Event.class); handle(e); }}
readValue(file, List.class) loads everything.
Jackson also enforces StreamReadConstraints (2.15+) such as max string length to guard against abusive payloads.
⚠ Follow-up traps
Does readValue advance past one element? Yes, leaving the parser at the end token of the object.
Can you stream with JsonNode? Only per element; the tree of the whole doc is in memory.
#jackson#streaming#large-files
Q45
Compare the Gson and Jackson approaches, and when is JSON a poor choice?
intermediate
Jackson is the de-facto standard (Spring), feature-rich, with streaming, modules and afterburner/blackbird optimisations. Gson is simpler, uses fields by default and has fewer features. JSON is a poor fit for high-volume internal messaging, binary payloads (base64 overhead 33%), exact numeric precision and schema enforcement.
Large long values lose precision in JavaScript consumers; send as strings.
BigDecimal needs USE_BIG_DECIMAL_FOR_FLOATS to avoid double conversion on read.
⚠ Follow-up traps
How is byte[] encoded? Base64 string.
Can JSON have comments or trailing commas? Not in the standard; Jackson can opt in via features.
#json#jackson#gson
Q46
How do you work with ZIP files in Java?
basic
Use ZipInputStream/ZipOutputStream for streaming, ZipFile for random access to entries, or the zip FileSystem provider (FileSystems.newFileSystem(path)) to treat an archive as a directory using Files.
GZIPOutputStream compresses a single stream (no entries), commonly used for .gz logs.
ZipFile reads the central directory, so it is faster for selective access; ZipInputStream must scan sequentially.
⚠ Follow-up traps
Does ZipOutputStream.close() write the central directory? Yes; failing to close produces a corrupt archive.
Is Zip charset UTF-8 by default? Yes for names since Java 7 unless a charset is passed.
#zip#zipinputstream#compression
Q47
What is Zip Slip and how do you prevent it?
intermediate
Zip Slip is a path traversal where an archive entry name like ../../etc/cron.d/x makes extraction write outside the target directory. Always resolve each entry against the destination, normalize, and verify it stays inside.
Path dest = destDir.toRealPath();Path out = dest.resolve(entry.getName()).normalize();if (!out.startsWith(dest)) throw new IOException("Bad entry: " + entry.getName());
Also handle symlink entries in tar archives and absolute names.
Create parent directories explicitly, and skip directory entries correctly.
⚠ Follow-up traps
Does ZipFile sanitise names? No, the names are returned raw.
Is the check needed for Files.copy into a zip FileSystem? Not for writing; it is for extraction to the real filesystem.
#zip#security#zip-slip
Q48
What is a zip bomb and how do you defend against it?
advanced
A small archive that expands to gigabytes (or nests archives) to exhaust disk, memory or CPU. Defend by limiting total uncompressed bytes, per-entry size, entry count, nesting depth and the compression ratio while extracting.
Do not trust ZipEntry.getSize(): it can be -1 or forged; count the bytes actually read.
Abort once limits are exceeded and delete partial output.
Run extraction in a quota-limited directory.
⚠ Follow-up traps
Is entry.getSize() safe to check? No, the header can lie in streaming mode.
Does the zip format forbid nested zips? No, so depth limits are needed.
#zip#security#dos
Q49
How do you read and write CSV at scale?
intermediate
Stream rows using a real CSV parser (OpenCSV, Apache Commons CSV, univocity) with a BufferedReader, handling each record and discarding it. split(",") is wrong for quoted fields, embedded commas, newlines inside quotes and escaped quotes.
try (var reader = Files.newBufferedReader(p); CSVParser parser = CSVFormat.DEFAULT.builder().setHeader().setSkipHeaderRecord(true) .build().parse(reader)) { for (CSVRecord r : parser) { process(r.get("id")); }}
readLine-based parsing breaks on quoted multi-line fields.
Batch inserts (e.g., 1000 rows per JDBC batch) rather than per-row.
For export, write through BufferedWriter/response stream and never build a giant List first.
⚠ Follow-up traps
Is CSVParser iteration lazy? Yes, records are produced as you iterate.
How do you guard against CSV injection when exporting? Prefix cells starting with =, +, -, @ with a single quote or tab.
#csv#streaming#parsing
Q50
What problems does Java's old object stream have with large or long-lived streams?
advanced
ObjectOutputStream keeps a strong reference to every object written in its handle table to encode back-references, so writing millions of distinct objects to one stream leaks memory. Call reset() periodically (or use writeUnshared) to clear the table; also mind that mutating an object after writing and re-writing it sends only a back-reference, not the new state.
Same issue on the read side: ObjectInputStream retains deserialized objects.
A new ObjectOutputStream writes a header; appending to an existing file requires a custom subclass overriding writeStreamHeader.
⚠ Follow-up traps
Why is my updated object not saved when re-written? The stream sent a handle to the earlier version; call reset() or writeUnshared.
Can you open a file in append mode and add objects? Not directly, a second header corrupts reading.
#serialization#objectoutputstream#memory
Q51
How do you copy a stream and what does transferTo do?
basic
Since Java 9 InputStream.transferTo(OutputStream) copies everything with an internal buffer (16 KB, later 8 KB variants). Files.copy(InputStream, Path) and Files.copy(Path, OutputStream) cover the common cases.
Before Java 9 write the loop: read into byte[], write the count read.
Always use the returned count from read; writing the full array writes stale bytes.
⚠ Follow-up traps
Does transferTo close either stream? No.
Why not byte[] b = in.readAllBytes() for copies? It holds the entire content in memory.
#streams#copy#transferto
Q52
What are InputStream.read semantics and the common mistakes?
intermediate
read() returns an int 0-255 or -1 at EOF; read(byte[]) returns the number of bytes actually read (possibly fewer than the array size) or -1. Ignoring the return value is the classic bug.
Use readNBytes(n) (Java 9+) or DataInputStream.readFully when you need exactly n bytes.
available() is only an estimate of non-blocking bytes, not the file size or the end marker.
skip(n) may skip fewer bytes than asked.
⚠ Follow-up traps
Why does read() return int? To distinguish the EOF value -1 from the byte 0xFF.
Is available() == 0 end of stream? No, it may just mean no data yet.
#streams#read#pitfalls
Scenarios
Q53
You must find all ERROR lines in a 10 GB log file on a 2 GB heap. How?
basic
Stream line by line and keep no state beyond what you need. Memory stays constant, regardless of file size.
try (BufferedReader r = Files.newBufferedReader(log, StandardCharsets.UTF_8); BufferedWriter w = Files.newBufferedWriter(out)) { String line; while ((line = r.readLine()) != null) { if (line.contains("ERROR")) { w.write(line); w.newLine(); } }}
Files.readAllLines would throw OutOfMemoryError (and exceed the 2 GB array limit).
Run time is bounded by disk throughput; grep may be faster for a one-off.
⚠ Follow-up traps
What if the file contains invalid UTF-8?Files.newBufferedReader throws MalformedInputException; use a decoder with REPLACE or ISO-8859-1.
What if a single line is 1 GB?readLine buffers the whole line; read in chunks instead.
#large-files#logs#streaming
Q54
A 10 GB log must be searched faster. How do you parallelize?
advanced
Split the file into byte ranges, adjust each start forward to the next newline and each end to finish the last line, then process ranges on a thread pool, each with its own FileChannel positional reads or RandomAccessFile. Merge partial results.
Aligning boundaries: seek to the range start, skip to after the next \n (except range 0), and read until passing the range end.
Gains are limited by disk throughput; on SSD/NVMe and cached files it scales with cores, on HDD random reads hurt.
Use FileChannel.read(buf, pos) which is thread-safe.
⚠ Follow-up traps
Why not Files.lines(p).parallel()? Splitting is coarse and line-boundary cost makes it unpredictable; fine for CPU-heavy per-line work only.
Does UTF-8 break at arbitrary byte boundaries? Yes mid-character, but \n (0x0A) is never part of a multi-byte sequence, so aligning on it is safe.
#large-files#parallelism#filechannel
Q55
A user uploads a 5 GB file to your Spring endpoint and the service goes out of memory. What went wrong?
intermediate
The code probably called file.getBytes() or read the whole body into a byte[]/String. Stream the InputStream straight to disk or object storage instead.
Set spring.servlet.multipart.max-file-size and max-request-size; multipart parsers spool above file-size-threshold to disk.
For real big uploads use chunked/resumable uploads or pre-signed S3 URLs so bytes bypass your JVM.
Do not trust file.getOriginalFilename() for paths.
⚠ Follow-up traps
Does MultipartFile hold the file in heap? Small ones may be in memory; above the threshold it is a temp file.
Is Content-Length reliable for limits? No, enforce limits while reading.
#upload#multipart#memory
Q56
A power cut leaves your config file half-written. How do you prevent corrupt writes?
intermediate
Never overwrite in place. Write a temp file in the same directory, force it, then ATOMIC_MOVE over the target. Readers then see either the old or the new complete file.
Files.writeString(target, ...) truncates first, so a crash mid-write leaves an empty or partial file.
Keep one backup generation (config.json.bak) for safety.
Validate (parse/checksum) on load and fall back to the backup.
⚠ Follow-up traps
Is ATOMIC_MOVE enough for durability? It guarantees atomic visibility, but without fsync the data may still be lost after a crash.
What if the tmp file is on another mount? The move throws AtomicMoveNotSupportedException or degrades to a copy.
#atomic-write#durability#corruption
Q57
Two processes append to the same log file. Lines are interleaved. What are your options?
advanced
Open with StandardOpenOption.APPEND, which on POSIX makes each write atomic positioned at EOF, and write each record in a single write call. For multi-write records or non-local filesystems use FileLock, or better route logs through one writer.
BufferedWriter can split a line across multiple write syscalls when its buffer fills; build the complete line (with newline) first and write once.
Atomicity of appends is only guaranteed up to a size (PIPE_BUF applies to pipes, not files); NFS has no atomic append.
Best practice: one logging process or a logging framework with file appender locking.
⚠ Follow-up traps
Does synchronized help across processes? No, only within the JVM.
Do two threads in the same JVM sharing one FileChannel interleave safely? Channel writes are serialized, but your multi-call records are not atomic.
#concurrency#append#filelock
Q58
What does this code print, and why is there a bug?
basic
Output: abcd, then efcd, then 2. The second read fills only two bytes; c and d remain from the previous read. Ignoring the returned count prints stale data.
Use new String(buf, 0, n).
Real streams (sockets, files) may return fewer bytes than requested even before EOF.
⚠ Follow-up traps
What does a third read(buf) return? -1, and the buffer is untouched.
Which method guarantees filling exactly 4 bytes or throwing?DataInputStream.readFully / readNBytes.
Ask the producer for UTF-8 (Excel's "CSV UTF-8" adds a BOM) and decode explicitly.
Strip a leading if present.
Never rely on the platform default (before Java 18 it differed across servers).
⚠ Follow-up traps
Why does it work on your laptop but not in Docker? The container's default locale may be POSIX, so pre-Java 18 default charset is ASCII.
Can you auto-detect? Only heuristically; agree on the encoding in the contract.
#charset#encoding#csv
Q60
After upgrading from Java 11 to 21, a file written with FileWriter reads differently. Why?
advanced
JEP 400 made UTF-8 the default charset in Java 18+, whereas Java 11 used file.encoding from the OS locale (e.g., Windows-1252 or ISO-8859-1). Existing files written with the old platform default are now read as UTF-8, corrupting non-ASCII characters.
Fix by specifying the charset on all constructors/methods, explicitly using the legacy charset for old files.
Compatibility switch: -Dfile.encoding=COMPAT restores the legacy behaviour.
native.encoding system property exposes the OS charset.
⚠ Follow-up traps
Did System.out change too? It uses stdout.encoding (terminal derived) since Java 19, not the default charset.
Would an ASCII-only file be affected? No, ASCII is identical in both.
#charset#java18#jep-400
Q61
A batch job processes thousands of files and crashes with "Too many open files". Diagnose it.
intermediate
File descriptors are leaked: streams, readers or Files.list/walk/lines streams are not closed on every path (especially exceptions). The process hits the ulimit -n limit.
// leakFiles.list(dir).forEach(this::handle);// fixtry (Stream<Path> s = Files.list(dir)) { s.forEach(this::handle); }
Inspect with lsof -p <pid> or ls /proc/<pid>/fd | wc -l.
Raising ulimit hides the leak; fix the code.
GC finalisation/Cleaner may eventually close leaked streams, which makes failures intermittent.
⚠ Follow-up traps
Will GC close forgotten streams? Eventually and unpredictably; do not rely on it.
Do sockets count too? Yes, all descriptors share the limit.
#leak#file-descriptors#try-with-resources
Q62
You delete a file after processing but on Windows it throws. On Linux it works. Why?
intermediate
Windows prevents deleting or renaming a file that is still open without share-delete semantics, so a leaked stream or an active memory mapping makes Files.delete throw FileSystemException ("being used by another process"). On Linux unlinking an open file succeeds; the data persists until the last descriptor closes.
Ensure all handles are closed first; mapped buffers stay mapped until GC.
On Linux, deleting an open log file still consumes disk space until the process closes it (common "disk full but du says empty" case).
⚠ Follow-up traps
Does DELETE_ON_CLOSE work on Windows and Linux? Yes, but semantics are implementation-specific.
Can you read a file after deleting it on Linux? Yes, through the already open descriptor.
#delete#windows#platform
Q63
Design a safe zip upload extraction endpoint.
advanced
Treat the archive as hostile: extract into a fresh temp directory, validate each entry, limit sizes, then process.
Zip Slip: dest.resolve(name).normalize().startsWith(dest).
Zip bomb: cap entry count, per-entry and total uncompressed bytes (counted as read), nesting and ratio.
Reject or ignore symlinks, absolute paths, device-like names; sanitise names for the OS.
Use a dedicated quota-limited directory, run with minimal privileges, delete on failure.
Scan for malware before use and never execute extracted content.
⚠ Follow-up traps
Is it enough to check entry.isDirectory()? No, it only depends on a trailing slash in the name.
Should you trust the extension or MIME type? No, validate content.
#zip#security#upload
Q64
Reading JSON with Jackson fails with `UnrecognizedPropertyException` after the producer adds a field. Fix?
basic
By default FAIL_ON_UNKNOWN_PROPERTIES is true. Make consumers tolerant: disable it globally or annotate the type with @JsonIgnoreProperties(ignoreUnknown = true).
Tolerant reader is a core rule for compatible evolution: ignore unknown fields, never repurpose existing ones.
Spring Boot's auto-configured ObjectMapper already disables the feature, so the failure often comes from a hand-built mapper.
⚠ Follow-up traps
What about a missing field? It becomes null/default unless FAIL_ON_MISSING_CREATOR_PROPERTIES is enabled.
Is silently ignoring typos a risk? Yes, so strict mode for config files and tolerant for inter-service payloads.
#jackson#evolution#compatibility
Q65
You need to send events between services at high volume with evolving schemas. JSON, Protobuf or Avro?
intermediate
For gRPC/RPC choose Protobuf. For Kafka pipelines with many producers and consumers choose Avro (or Protobuf) with a schema registry enforcing compatibility. Choose JSON for public/browser APIs and low volume where debuggability matters.
Measure payload size and CPU; binary formats save bandwidth and parse time.
Registry plus CI compatibility check blocks breaking schema changes before deployment.
Provide a debug tool (e.g., kafka-avro-console-consumer) to offset binary opacity.
⚠ Follow-up traps
Does Protobuf need a registry? Not required; compiled schemas plus field numbers are sufficient, but registries help governance.
Can you change a field's type later? Mostly no, add a new field instead.
#protobuf#avro#json#design
Q66
A cached object stored with Java serialization breaks after a deploy with `InvalidClassException`. Why and how to avoid it?
intermediate
The class changed and, with no explicit serialVersionUID, the computed UID changed, so old cache entries cannot be read by the new version. Declare a fixed UID and only make compatible changes, or version cache keys and flush on deploy.
Better: store JSON/Protobuf in Redis with a version in the key.
During rolling deploys old and new nodes share the cache; incompatible shapes cause errors on one side.
⚠ Follow-up traps
Is adding a method an incompatible change with a declared UID? No, methods are not serialized.
What if you rename a field? It is treated as a removed field plus a new field with default value.
#serialization#serialversionuid#caching
Q67
What does this print?
intermediate
Deserialization prints Base() only. Base is not serializable, so its no-arg constructor runs and b is reset to 5; Child() is not invoked. The result has c=2 (restored) and t=0 (transient, default value, field initialisers not run).
Serialization time prints Base() and Child() once from the original construction.
⚠ Follow-up traps
Is t 9 after deserialization? No, initialisers do not run, so it is 0.
What if Base had only a constructor taking args?InvalidClassException: no valid constructor.
#serialization#transient#constructor
Q68
A deserialization endpoint accepts `ObjectInputStream` from HTTP clients. How do you fix it?
advanced
Best: remove native deserialization from untrusted input and switch to JSON/Protobuf with explicit DTOs. If it cannot be removed immediately, apply an allow-list ObjectInputFilter, limit depth/size, authenticate and sign payloads, and remove gadget libraries or update them.
ObjectInputFilter f = ObjectInputFilter.Config.createFilter( "com.acme.api.Request;java.base/java.lang.*;java.base/java.util.*;maxdepth=5;maxbytes=65536;!*");ois.setObjectInputFilter(f);
Use HMAC signatures if only your own services write the data; never encrypt-only.
Monitor with -Djdk.serialFilter plus logging to learn which classes appear.
⚠ Follow-up traps
Is a ClassLoader allow-list in resolveClass still valid? Yes, the pre-JEP 290 approach works but filters are preferred.
Is a WAF enough? No, payloads are binary and easily obfuscated.
#serialization#security#rce
Q69
A record with validation is deserialized from Jackson JSON with a null name. What happens?
intermediate
Jackson 2.12+ calls the canonical constructor with null for the missing component, so the compact constructor's checks run and throw; Jackson wraps it in ValueInstantiationException.
record User(String name) { User { Objects.requireNonNull(name, "name"); }}// {} -> ValueInstantiationException caused by NullPointerException: name
Without a check, the record silently holds null.
Bean Validation annotations are not enforced by Jackson; validate separately.
⚠ Follow-up traps
Does Jackson need setters for records? No, it uses the canonical constructor.
Does Jackson < 2.12 support records? Not properly; upgrade.
#records#jackson#validation
Q70
Your Java-serialized record deserializes with a different `serialVersionUID` and works. Is that a bug?
advanced
No. For records the serialVersionUID match requirement is waived; the stream is matched by component names and types, then the canonical constructor is invoked. The default UID for a record is 0L unless declared.
Compatibility depends on component names and types: adding a component gives default value (null/0) for old streams, which is passed to the constructor.
Constructor validation may then reject old data.
⚠ Follow-up traps
Can you bypass the constructor with a crafted stream? No, which is the security advantage.
Do removed components break reading? The stream's extra component is ignored.
#records#serialization#serialversionuid
Q71
Your file watcher service misses events or receives duplicates when a large file is copied into the watched folder. What do you do?
intermediate
ENTRY_CREATE fires when the file appears, not when the copy completes, and ENTRY_MODIFY may fire repeatedly. Do not process on first event.
Producers should write to a temp name or directory and atomically rename into the watched folder.
Otherwise debounce: wait until size and mtime are stable, or until a .done marker appears, or try an exclusive lock/open.
Handle OVERFLOW by rescanning the directory; keep a processed-files ledger for idempotency.
⚠ Follow-up traps
Can tryLock prove the writer is finished? Not reliably, locks are advisory on Unix.
Does WatchService watch subdirectories? No, register each one.
#watchservice#race-condition#ingest
Q72
WatchService on macOS reacts after about 10 seconds, but on Linux instantly. Why?
advanced
Linux uses inotify (event-driven). The JDK on macOS has no native FSEvents backend for the default WatchService and falls back to a polling implementation with a default interval of about 10 seconds.
Use the com.sun.nio.file.SensitivityWatchEventModifier.HIGH modifier (about 2 s; deprecated and removed in recent JDKs) or a library using FSEvents (e.g., directory-watcher).
Production services typically run on Linux, but dev environments differ; tests should not assume timing.
⚠ Follow-up traps
Is there a portable API for the interval? No.
Does polling detect every change? It can miss changes that revert between two polls.
#watchservice#platform#polling
Q73
A cron job and your app both rewrite the same file. How do you coordinate?
intermediate
Use an exclusive FileLock on a dedicated lock file (not the data file you atomically replace) that both sides acquire, or use an atomic write so concurrent readers are safe and one writer wins.
Lock the stable lock file because renaming the data file replaces the inode and the lock on the old one no longer protects the new one.
The other process must use the same protocol (advisory locks).
⚠ Follow-up traps
What if the JVM crashes while holding the lock? The OS releases it when the process dies.
What does a second lock() in the same JVM do? Throws OverlappingFileLockException.
#filelock#concurrency#cooperation
Q74
You must make sure only one instance of a desktop or batch app runs. How?
basic
Open a lock file and call tryLock(); if it returns null another process holds it, so exit. Keep the channel and lock referenced for the process lifetime (a GC'd channel releases the lock).
Do not delete the lock file on exit as that causes races; the lock, not the file existence, is the guard.
Not reliable on NFS.
⚠ Follow-up traps
Why store the lock in a static field? Otherwise the channel may be garbage collected and the lock released.
Is checking Files.exists as a guard OK? No, it is racy and leaves stale files after crashes.
#filelock#single-instance
Q75
A temp-file directory fills the disk on a server running for months. Cause?
intermediate
Temp files are created per request but never deleted; deleteOnExit only runs at JVM shutdown (and also leaks its internal path list), so a long-lived server accumulates them.
Delete in finally/try-with-resources, or open with DELETE_ON_CLOSE.
Use a dedicated subdirectory per job and a scheduled cleaner by age.
Configure -Djava.io.tmpdir to a quota-limited volume; tmpfs /tmp counts against RAM.
⚠ Follow-up traps
Does the OS clean /tmp? Possibly at reboot or via tmpfiles rules, not reliably.
Is createTempFile name predictable? No, random.
#temp-files#leak#deleteonexit
Q76
Reading a 3 GB file with `Files.readAllBytes` throws. What error and what are the alternatives?
basic
For files above about 2 GB it throws OutOfMemoryError: Required array size too large since arrays are indexed by int. Under the limit it may still exhaust the heap.
Stream with a buffer and process chunks.
Memory-map in windows when random access is needed.
Use a FileChannel with a reused ByteBuffer.
⚠ Follow-up traps
Does raising -Xmx fix it? Not above the array limit.
Is readString any better? It also loads all bytes and then builds a String, using roughly 2x memory.
#large-files#memory#readallbytes
Q77
You map a 6 GB file with FileChannel.map and get an exception. Fix?
advanced
map throws IllegalArgumentException: Size exceeds Integer.MAX_VALUE. Map the file in multiple windows (e.g., 1 GB each) and index across them, or use the Java 22+ FFM API where FileChannel.map(mode, offset, size, Arena) returns a MemorySegment with long offsets.
Records straddling a window boundary need overlap handling.
Unmapping is GC-driven for MappedByteBuffer; many large windows can hit address-space or OutOfMemoryError: Map failed.
⚠ Follow-up traps
Does the whole file get loaded into RAM? No, only touched pages.
How does the Java 22 API free memory? Closing the Arena unmaps deterministically.
#mmap#large-files#limits
Q78
Two JVMs share a memory-mapped file for IPC. What do you need to be careful of?
advanced
The mapping is shared via the page cache, so writes by one process are visible to the other, but without proper ordering. Use a file layout with a sequence number or header, write payload first and then publish the length/flag with ordered or volatile semantics (VarHandle) and handle torn reads.
Java gives no cross-process mutexes in mapped memory; use atomic fields via VarHandle on a ByteBuffer, or a lock file.
force() is for durability, not visibility between processes.
Libraries like Chronicle Queue and Aeron implement this robustly.
⚠ Follow-up traps
Is force() needed so the other process sees the write? No, the page cache is shared.
Can a reader see a half-written record? Yes, without a commit marker.
#mmap#ipc#concurrency
Q79
A request returns a big report. How do you stream it without building it in memory?
intermediate
Write directly to the response stream (StreamingResponseBody in Spring MVC, or HttpServletResponse.getOutputStream()), fetching rows with a streaming cursor, and flush periodically.
@GetMapping("/export")StreamingResponseBody export() { return out -> { try (var w = new BufferedWriter(new OutputStreamWriter(out, UTF_8))) { repo.streamAll(row -> write(w, row)); } };}
Use DB fetch size/cursor (setFetchSize, Stream return types) in a transaction.
Streaming means errors after the first bytes cannot change the HTTP status; the client sees a truncated response.
Configure async timeouts and a task executor.
⚠ Follow-up traps
Can you still return HTTP 500 midway? No, headers were already sent.
Does a returned List stream lazily? No, it is fully materialised first.
#streaming#export#spring
Q80
Your CSV importer reads a file with `line.split(",")` and mangles a row `1,"Smith, John",NY`. What now?
basic
split(",") yields four fields (1, "Smith, John", NY) because it ignores quotes. Use a proper CSV parser that implements RFC 4180 quoting and escaped double quotes ("").
Also embedded newlines within quotes break readLine loops.
Trailing empty fields are dropped by split(",") without a limit argument (split(",", -1) keeps them).
⚠ Follow-up traps
Does split(",", -1) fix the quotes? No, only the trailing empties.
Is the delimiter always a comma? No (semicolon in many locales); make it configurable.
#csv#parsing#bug
Q81
Import of 50 million CSV rows into a DB takes hours. How do you speed it up?
advanced
Stream the file, batch JDBC inserts (500-5000 rows per executeBatch and commit), use rewriteBatchedStatements=true (MySQL) or reWriteBatchedInserts=true (PostgreSQL), or the DB bulk loader (COPY, LOAD DATA).
Drop or defer secondary indexes/constraints during the load, then rebuild.
Parse and insert in a pipeline: reader thread to bounded queue to writer threads.
Track progress (last committed line) for restartability.
⚠ Follow-up traps
Why not commit per row? Each commit is an fsync; throughput collapses.
What if one bad row fails the batch? Catch BatchUpdateException, quarantine bad rows, and continue.
#csv#batch#jdbc#performance
Q82
What does this print, and how do you fix it?
intermediate
It prints body 1. The body exception is primary, and the close exception is attached as a suppressed exception. With old try/finally, close's exception would have replaced body.
Log e.getSuppressed() in diagnostics; many loggers print them automatically.
⚠ Follow-up traps
If only close() throws? That exception propagates normally.
Is the suppressed list populated if you disable suppression?Throwable constructed with enableSuppression=false ignores it.
#try-with-resources#exceptions#suppressed
Q83
What is wrong with this code that wraps streams?
intermediate
If the GZIPInputStream constructor throws (e.g., file is not gzip, header read fails, ZipException/EOFException), the FileInputStream was already created but never assigned to a resource, so it leaks.
try (var fin = new FileInputStream(path); var gz = new GZIPInputStream(fin)) { /* read */ }
Files.newInputStream plus a separate declaration avoids the leak.
Most JDK stream constructors that read headers can throw.
⚠ Follow-up traps
Does closing gz close fin? Yes, but only if gz was constructed.
Would the GC close it? Eventually via cleaner, nondeterministically.
#try-with-resources#leak#decorator
Q84
Users upload a file named `../../app/config.yml` and overwrite your configuration. Fix the code `Files.copy(in, uploadDir.resolve(name))`.
intermediate
Never use client names as paths. Generate a server-side name (UUID) and store the original name as metadata. If you must use the name, take only Path.of(name).getFileName(), normalize and check startsWith(base).
Reject names with NUL, separators, reserved Windows names.
Store uploads outside the web root and without execute permission.
Validate content type by sniffing, enforce size limits.
⚠ Follow-up traps
Does getFileName() fully protect you? It strips directories but a leading .. alone can still be a name; also reject . and ...
Does startsWith need toRealPath()? Yes, to defeat symlinks inside the upload directory.
#path-traversal#upload#security
Q85
A symlink inside the upload directory points to /etc. How does it affect your safety checks?
advanced
normalize() is purely lexical so uploads/link/passwd passes startsWith(base) but actually reads /etc/passwd. Resolve real paths and use LinkOption.NOFOLLOW_LINKS.
Path real = target.toRealPath();if (!real.startsWith(base.toRealPath())) throw new SecurityException();Files.readString(target); // or open with NOFOLLOW_LINKS
Still racy (TOCTOU): a link can be swapped between check and use; open with NOFOLLOW_LINKS and prefer a directory the app alone controls.
Zip extraction must not create symlinks from entries.
⚠ Follow-up traps
Does Files.isSymbolicLink fully solve it? Only checks the final element, not parents.
What is TOCTOU? Time-of-check to time-of-use race between validating and opening.
#symlink#security#nofollow
Q86
A scheduled export writes `report.csv` and a downstream reader sometimes reads a half-written file. Fix?
intermediate
Write to report.csv.tmp (or a hidden name) in the same directory and Files.move with ATOMIC_MOVE to report.csv once complete. Readers only match final names, never *.tmp.
Alternative: write a report.csv.done marker last.
Do not rely on file size stabilising as proof of completion.
⚠ Follow-up traps
Does a rename keep a reader that already opened the old file consistent? Yes, on POSIX it keeps reading the old inode.
What about S3? Object puts are atomic per key; you can upload directly.
#atomic-write#race-condition#handoff
Q87
You need to rotate a log file while the app writes to it. What is safe?
advanced
Renaming an open file is fine on POSIX since the handle follows the inode, so the writer would keep writing to the renamed file. Either let the logging framework (Logback/Log4j2 rolling appenders) close and reopen, or use copy-truncate (which can lose lines written between the copy and truncation).
With external logrotate, use create plus a signal to reopen, or copytruncate as a last resort.
On Windows, renaming an open file fails.
⚠ Follow-up traps
Why does the app keep filling the old file after rotation? It holds the old descriptor open.
Does Logback lose logs during rolling? Not by default; the appender locks and switches files.
#logging#rotation#file-handles
Q88
A `BufferedWriter` is used for an audit log, but entries vanish on a crash. Why?
intermediate
Entries sit in the 8 KB Java buffer (and then the OS page cache) until flushed. A crash or kill -9 loses buffered data; a power loss loses even flushed-but-not-fsynced data.
Call flush() per record for process-crash safety.
For power-loss durability use FileChannel.force or SYNC open options, balanced against throughput (group commit).
Audit logs often go to a log pipeline or DB with transactional guarantees instead.
⚠ Follow-up traps
Does flush() guarantee on-disk data? No, only hand-off to the OS.
What does a normal System.exit do to the buffer? Nothing; you must close or flush.
#buffering#flush#durability
Q89
A report prints `?` for non-Latin characters when written with `System.out` in a container. Cause?
intermediate
The console encoding is derived from the container's locale (often POSIX/ANSI_X3.4-1968), so unmappable characters become ?. The data in memory is correct.
Set LANG=C.UTF-8, -Dstdout.encoding=UTF-8 (Java 19+) or -Dfile.encoding=UTF-8 on older JVMs, or write bytes via new PrintStream(System.out, true, UTF_8).
For files, always specify the charset when writing.
⚠ Follow-up traps
Does Java 18's UTF-8 default fix System.out? No, it uses the separate stdout encoding.
Is the string itself corrupted? No, encoding happens only at output.
#charset#stdout#containers
Q90
Deserializing an old `ArrayList` from disk after a refactor of your DTO package fails with `ClassNotFoundException`. Why?
intermediate
The stream records fully qualified class names. Renaming or moving a class breaks reads of old data.
Workarounds: keep the old class as a stub, override ObjectInputStream.resolveClass to map old names to new classes, or migrate data with a converter job.
Avoid storing Java-serialized blobs long-term; use schema'd formats.
⚠ Follow-up traps
Does changing the UID fix it? No, the class is not found at all.
Does moving a nested class matter? Yes, the binary name (Outer$Inner) is recorded.
#serialization#refactoring#compatibility
Q91
A `Serializable` class holds a `Thread` and a `Connection`. What happens when serialized, and what's the proper design?
basic
NotSerializableException for the first non-serializable field encountered. Mark such fields transient and re-create them lazily, or better, do not make resource-holding classes serializable; serialize a DTO with the needed state.
Inner classes capture the outer instance, which can drag a non-serializable outer object into the graph.
Lambdas are serializable only if the target interface is and captured values are serializable.
⚠ Follow-up traps
Why does serializing an anonymous inner class fail? It captures this of the enclosing instance.
Are lambdas serialized by default? No, only when cast to a serializable functional interface.
#serialization#transient#design
Q92
After deserialization your `Singleton.getInstance() == deserialized` is false. Fix?
intermediate
Deserialization creates a new instance bypassing the private constructor. Add private Object readResolve() { return INSTANCE; } or, better, implement the singleton as a single-element enum.
Reflection can also break a private-constructor singleton; enums defeat both.
Make fields transient so no forged state leaks through.
⚠ Follow-up traps
Does readResolve return type matter? It must be assignable and the method name/signature exact.
Can an attacker read the instance before readResolve? For crafted streams with references to the object, yes; enums avoid this.
#serialization#singleton#readresolve
Q93
Compare memory and speed: reading a 1 GB file via `FileInputStream.read()` one byte at a time vs buffered. What to expect?
basic
Unbuffered single-byte reads issue about a billion system calls and can be 100-1000x slower than buffered or bulk reads. A BufferedInputStream or read(byte[64k]) reduces it to tens of thousands of calls.
The JIT cannot remove syscalls; the kernel transition dominates.
Optimal buffer size is typically 8-64 KB; very large buffers give diminishing returns.
⚠ Follow-up traps
Does BufferedReader.read() per char hurt? It is a buffered call, cheap compared to syscalls (still a synchronized method).
Does the OS already cache? Yes, but each syscall still costs.
#buffering#performance#syscalls
Q94
Reading a gzip log (`.gz`, 20 GB uncompressed) line by line. How?
basic
Chain the decompressing stream into a reader; decompression is streamed with constant memory.
try (var in = new GZIPInputStream(Files.newInputStream(p), 64 * 1024); var r = new BufferedReader(new InputStreamReader(in, UTF_8))) { r.lines().filter(l -> l.contains("ERROR")).forEach(System.out::println);}
Gzip is not seekable, so parallel splitting needs bgzip or block-based formats (zstd frames, Parquet).
CPU is likely the bottleneck rather than disk.
⚠ Follow-up traps
Does GZIPInputStream handle concatenated gzip members? Yes, since Java 7 it reads multiple members.
Can you jump to the middle of a gzip? No, you must decompress from the start.
#gzip#streaming#large-files
Q95
A Kafka consumer decodes Avro messages and suddenly fails after the producer deployed a new schema. What are the likely causes?
advanced
Most likely an incompatible schema change (new required field without default, type change, removed field without default) or the consumer cannot fetch the writer schema from the registry.
With a registry, enforce BACKWARD (or FULL) compatibility so registration of a breaking schema is rejected at publish time.
Consumers use the writer schema id embedded in the message to resolve against their reader schema.
Deploy consumers first for backward compatible changes, producers first for forward compatible.
⚠ Follow-up traps
Which upgrades first under BACKWARD? Consumers first, as new readers can read old data.
Does adding an optional field with default break old consumers? No, they ignore it if using schema resolution.
#avro#schema-evolution#kafka
Q96
A Protobuf message field was removed and reused by a different type in a new version. What goes wrong?
advanced
Old messages or old senders still emit the old number with its old type. The new reader interprets the bytes using the new type: the data is silently corrupted or fails with parse errors. Field numbers must never be reused.
Same wire type reuse (e.g., int32 to int64) may appear to work but still has semantic risks.
⚠ Follow-up traps
Are unknown fields dropped? Protobuf 3.5+ preserves them on re-serialisation.
Is a default value distinguishable from unset in proto3? Only for optional or message-typed fields.
#protobuf#schema-evolution#reserved
Q97
Your service stores JSON `long` ids and the JavaScript frontend shows wrong values. Why?
intermediate
JavaScript numbers are IEEE 754 doubles with safe integers up to 2^53 - 1. A 64-bit id such as 9007199254740993 is rounded when parsed. Serialize ids as strings (@JsonSerialize(using = ToStringSerializer.class) or JsonFormat.Shape.STRING).
Same problem with BigDecimal money values parsed by JS Number.
Use ISO-8601 strings for dates and instants.
⚠ Follow-up traps
Does the Java side lose data? No, only the JS consumer.
Is this a Jackson bug? No, it is a JSON/JavaScript limitation.
#json#jackson#precision
Q98
You serialize `LocalDateTime` with Jackson and clients get `[2026,10,4,9,30]`. How to emit ISO strings?
basic
JavaTimeModule writes dates as numeric arrays when WRITE_DATES_AS_TIMESTAMPS is enabled (default). Disable the feature to get ISO-8601 text like "2026-10-04T09:30:00".
ObjectMapper om = new ObjectMapper() .registerModule(new JavaTimeModule()) .disable(SerializationFeature.WRITE_DATES_AS_TIMESTAMPS);
Spring Boot sets this by default (spring.jackson.serialization.write-dates-as-timestamps=false).
Prefer Instant/OffsetDateTime over LocalDateTime for wire formats to avoid timezone ambiguity.
⚠ Follow-up traps
Why does LocalDateTime have no zone? It is a local wall-clock value; cross-service data needs an offset or UTC.
Without the module on 2.12+?InvalidDefinitionException.
#jackson#java-time#dates
Q99
Parsing a 2 GB JSON array with `om.readValue(file, new TypeReference<List<Event>>(){})` runs out of memory. What now?
intermediate
readValue builds all Event objects (plus parser overhead) in memory. Iterate using the streaming parser or MappingIterator, handling one event at a time.
try (MappingIterator<Event> it = om.readerFor(Event.class).readValues(file)) { while (it.hasNext()) handle(it.next());}
This works for concatenated objects (NDJSON). For a top-level array use a JsonParser positioned after START_ARRAY.
For files with one huge object, consider the pull parser and extract only needed fields.
⚠ Follow-up traps
Does NDJSON reduce the problem? Yes, each line is an independent record.
Does JsonNode tree parsing help? No, still in memory.
#jackson#streaming#memory
Q100
A security scanner flags `ObjectMapper.enableDefaultTyping()` in your code. Is it real and how do you fix it?
advanced
It is real. Default typing lets payloads name arbitrary classes, which allows gadget chain RCE with vulnerable classes on the classpath. Replace with explicit subtype registration or a restrictive PolymorphicTypeValidator.
Better still: @JsonTypeInfo(use = Id.NAME) and @JsonSubTypes.
Keep Jackson up to date; it blocks known gadget classes.
⚠ Follow-up traps
Does Object-typed property increase risk? Yes, with default typing it accepts any class.
Is the allow-list by package prefix safe? Only if every class in the package is safe to instantiate.
#jackson#security#default-typing
Q101
You extract a tar/zip and file permissions or timestamps are lost. What do you do?
intermediate
ZipInputStream/ZipEntry expose modified time but not Unix permission bits. Apache Commons Compress (ZipArchiveEntry.getUnixMode, TarArchiveEntry.getMode) reads them, and you then apply with Files.setPosixFilePermissions. Never apply setuid bits from untrusted archives.
Set Files.setLastModifiedTime from the entry.
For executables you ship yourself, restore mode 755 explicitly.
Windows ignores POSIX attributes (UnsupportedOperationException for POSIX views).
⚠ Follow-up traps
Does Files.copy into a zip FileSystem keep permissions? The zip provider supports a posix view in newer JDKs only with an option.
Should you honour setuid from an archive? No, mask them.
#zip#permissions#attributes
Q102
A read-only config file on a mounted volume throws `AccessDeniedException` on `Files.write`, but `File.canWrite()` returned true earlier. Why?
intermediate
Checks like canWrite()/Files.isWritable are advisory snapshots: permissions, ACLs, mount options and locks can change or differ from the effective access when you open the file. Just try the operation and handle the exception (EAFP).
AccessDeniedException is a subclass of FileSystemException with the reason; legacy File methods return just false.
Check-then-act sequences are racy (exists then create); use CREATE_NEW for atomic creation.
⚠ Follow-up traps
What does CREATE_NEW do if the file exists? Throws FileAlreadyExistsException atomically.
Does isWritable account for read-only mounts? Sometimes not.
#permissions#toctou#exceptions
Q103
Your batch copies many large files between mounts using `Files.move` and sometimes ends up with duplicates and partial files. Explain.
advanced
Across filesystems Files.move is a copy followed by a delete, not an atomic rename. A crash mid-copy leaves a partial target, and a crash between copy and delete leaves duplicates. ATOMIC_MOVE would fail with AtomicMoveNotSupportedException instead of degrading.
Copy to target.tmp on the destination filesystem, verify size or checksum, atomically rename within that filesystem, then delete the source.
Make the job idempotent: skip when the destination exists with matching checksum.
⚠ Follow-up traps
Can the temp file be on the source mount? No, the final rename would again cross filesystems.
Is REPLACE_EXISTING atomic across mounts? No.
#move#atomic#cross-filesystem
Q104
Which file format and API would you pick to persist a large in-memory cache snapshot (1 GB) so a restart is fast?
advanced
Avoid Java serialization (slow, fragile). Use a compact schema'd binary format (Protobuf/Avro/custom length-prefixed records) written as a stream, with a header (magic, version, count) and a trailing checksum, and load it in a streaming fashion.
Write atomically: temp file, fsync, rename.
Compress with LZ4/zstd if disk-bound.
For random access and zero deserialisation consider memory-mapped fixed layouts (as Chronicle Map or Lucene do).
Keep a version field so a loader can reject or migrate old snapshots, and fall back to a cold start on corruption.
⚠ Follow-up traps
Why a checksum? To detect truncated or corrupted snapshots rather than loading garbage.
Would you snapshot while the cache mutates? Take a consistent copy or use a copy-on-write structure, else the snapshot is torn.