Streaming Large File Uploads in Java Without Killing Your Server
Introduction You've seen the tutorial. A @PostMapping endpoint, a MultipartFile parameter, a call to transferTo() . Done in five minutes. Works great on localhost. Then your first 500MB upload hits production and your pod crashes with an OutOfMemoryError . The problem isn't your code - it's an assumption baked into most tutorials: that files are small enough to buffer entirely in memory or on disk before your handler even runs. In the real world - media platforms, document management, healthcare, fintech - that assumption breaks fast. And the WebFlux version of the tutorial is often no better. It just moves the buffering somewhere less obvious. This is Part 1 of a short series on building a production-grade upload pipeline in Java. Here we'll cover: - Why MultipartFile buffers the whole file before you get a chance to touch it - Why @RequestPart in WebFlux does exactly the same thing, including the version everyone copies - The one WebFlux API that genuinely streams, and what it takes to use it correctly Part 2 picks up from here: storing that stream to S3, detecting file type from a handful of bytes instead of the whole file, and running antivirus scanning safely at scale - including a failure mode worth knowing about before you put an upload endpoint in front of the public internet (a harmless-looking file that takes the whole server down with it). It also does all of this a second time on the servlet stack, for the many teams who can't move to WebFlux - so if you're on Spring MVC, stay with me. Why MultipartFile Is a Trap Spring's MultipartFile (and its equivalents) seem convenient, but under the hood the servlet container buffers the entire request body - either in memory or in a temp file - before your controller method is even called. For a 2GB video upload, this means: - 2GB of heap or disk consumed per concurrent upload - No streaming to storage until the full file is received - Zero opportunity for your own checks to run until you've paid the full I/O cost Be precise about which of the two you get, because they fail differently: - file-size-threshold: 0 (the Tomcat default) spools the whole body to a temp file. Ten concurrent 2GB uploads is 20GB of temp disk you didn't plan for, and nothing reaches your storage until the last byte lands. - A threshold above the file size keeps it in the heap instead. That's the version that kills the JVM - and every other in-flight request with it. One honest caveat before we move on: the container can reject early on its own limits - Tomcat enforces max-file-size during parsing. What you lose is not all rejection, it's your rejection: no quota check, no content-type policy, no persistence until the last byte has been spooled. Part 2 is about those checks, done cheaply. The Big Picture: Where This Series Is Headed Before diving into WebFlux specifics, here's the shape of the full pipeline we're building across this series - so the WebFlux piece below makes sense in context: 1. Accept the upload as a stream - never buffer the whole thing (this post) 2. Peek at the first ~500 bytes to check the file type, then rewind (Part 2) 3. If the type is rejected, bail out immediately - no storage, no AV work wasted 4. If accepted, fan out the single stream to: - storage backend (S3 or disk) - unbounded, we must persist the whole file - antivirus scan - bounded, capped at a sane max so a huge file can't choke it The common thread through every step: read the incoming bytes from the network exactly once, and never require the whole file to be sitting in memory or on disk before you can act on it. Everything in this post is about step 1 - getting a real stream out of the incoming request in the first place. That's the foundation Part 2 builds storage and analysis on top of. The Part Most WebFlux Tutorials Get Wrong Here is the version you've probably seen, and it does not stream: @PostMapping(value = "/upload", consumes = MediaType.MULTIPART_FORM_DATA_VALUE) public Mono upload( @RequestPart("file") Flux fileParts, @RequestPart("metadata") UploadMetadata metadata) { return uploadService.streamToStorage(fileParts, metadata); } It looks right. A Flux is a reactive stream of chunks, and the type says "stream", so surely it streams. It doesn't, and Spring says so in its own documentation: To parse multipart data in streaming fashion, you can use the Flux returned from thePartEventHttpMessageReader instead of using@RequestPart , as that impliesMap -like access to individual parts by name and, hence, requires parsing multipart data in full. That's the whole trick, stated plainly: @RequestPart means "look this part up by name", and you cannot look parts up by name in a body you haven't finished reading. So Spring reads the whole thing first. What that means for each signature you might reach for: | Signature | What actually happens | |---|---| @RequestPart("file") MultipartFile | Servlet stack: whole body buffered before your handler runs | @RequestPart("file") FilePart | WebFlux: whole part buffered in memory, then spilled to a temp file, before your handler runs | @RequestPart("file") Flux | WebFlux: joined into one buffer, capped by max-in-memory-size (256KB by default) | @RequestBody Flux | Actually streams | Three of those four deserve a closer look, because each fails in a different way. FilePart is a temp-file-backed MultipartFile . This is the one that catches people, because FilePart looks like the modern, reactive answer. It isn't: by the time your handler is called, the part has been fully received and written to a .multipart temp file on disk. Spring Boot's own configuration reference admits it, describing the temp-file setting as "Directory used to store file parts larger than maxInMemorySize . **Ignored when using the PartEvent streaming support" - a setting that exists only because @RequestPart writes parts to disk before you see them. The Flux version fails loudly, and confusingly. There's no reactive adapter for a bare DataBuffer , so the decoder falls back to its single-value path and joins your "stream" into one composite buffer bounded by spring.http.codecs.max-in-memory-size . With the default 256KB, a 1MB upload comes back as HTTP 413 - which nobody expects from an endpoint with no size limit configured. Raise that limit to "fix" it and you have bought yourself a full in-heap copy of every upload: the OOM, with extra steps. And DataBufferUtils.join() doesn't help either. The usual "pragmatic bridge" is: InputStream inputStream = DataBufferUtils.join(data) .map(buf -> buf.asInputStream(true)) .block(); Two problems. join() aggregates every buffer into one composite buffer, so memory grows with file size - which is the exact thing we're trying to avoid. And .block() runs on the Netty event-loop thread the handler was invoked on, so it never gets that far: java.lang.IllegalStateException: block()/blockFirst()/blockLast() are blocking, which is not supported in thread reactor-http-nio-4 That's not a tuning problem. It's a "this code cannot run" problem. Spring Boot: WebFlux Reactive Multipart, Done Properly The API that streams is @RequestBody Flux : instead of a map of named parts, you get a flat stream of events as the body is parsed. A form field produces one FormPartEvent ; a file produces one or more FilePartEvent s, each carrying a buffer. The final event of each part has isLast() set, which is what lets you split the stream back into parts without buffering it. @RestController public class UploadController { private final PartEventUploadService uploads; @PostMapping(value = "/api/upload", consumes = MediaType.MULTIPART_FORM_DATA_VALUE) public Mono > upload(@RequestBody Flux events) { return uploads.store(events) .map(result -> ResponseEntity.status(HttpStatus.CREATED).body(result)); } } The service handles one part at a time, in order: public Mono store(Flux events) { AtomicReference metadata = new AtomicReference<>(); return events .windowUntil(PartEvent::isLast) // one window per part .concatMap(window -> handlePart(window, metadata)) // strictly sequential .next() .switchIfEmpty(Mono.error(new InvalidUploadException("No 'file' part in the request"))); } Two details in there are load-bearing. windowUntil(PartEvent::isLast) splits the flat event stream into per-part windows; concatMap then processes them one at a time, so there is nowhere for the body to pile up. And because it is a single pass, part order matters: metadata has to be sent before file , because there is no second pass in which to look it up. That's a real constraint, not a stylistic choice - and it's the price of not buffering. The other detail is the one that leaks memory if you get it wrong. FilePartEvent 's content buffers must be consumed, relayed, or released - every one of them - and the framework will not do it for you. Streaming to storage The storage backend takes the file part's buffers and writes them as they arrive: @Override public Mono store(Flux content, StorageSpec spec) { Path tempFile = root.resolve(key + ".part"); Path targetFile = root.resolve(key); AtomicLong written = new AtomicLong(); MessageDigest digest = sha256(); Flux metered = content .doOnDiscard(DataBuffer.class, DataBufferUtils::release) .doOnNext(buffer -> { long total = written.addAndGet(buffer.readableByteCount()); if (total > maxBytes) { DataBufferUtils.release(buffer); // we are aborting: this one is ours to free throw new UploadTooLargeException(maxBytes, total); } digest.update(buffer.asByteBuffer().duplicate()); }) .map(FileSystemStorageBackend::toHeapBuffer); return Mono.usingWhen( openChannel(tempFile), channel -> DataBufferUtils.write(metered, channel) // write() re-emits each buffer once it is on disk: releasing is our job .doOnNext(DataBufferUtils::release) .doOnDiscard(DataBuffer.class, DataBufferUtils::release) .then(Mono.fromCallable(() -> commit(channel, tempFile, targetFile, digest, written))), channel -> closeQuietly(channel), (channel, cause) -> closeQuietly(channel).then(deleteQuietly(tempFile)).t
Comments
No comments yet. Start the discussion.