Conversation
Decompress gzip uploads in 1 MiB chunks into a spooled temporary file that spills to disk above 1 MiB, instead of materializing the expanded payload as bytes plus a BytesIO copy. Reuse the request-owned UploadFile so FastAPI closes the decompressed file when the request finishes, and close partial output when decompression fails.
9a54345 to
8c00ffd
Compare
There was a problem hiding this comment.
All reported issues were addressed across 3 files
Shadow auto-approve: would not auto-approve because issues were found.
Re-trigger cubic
There was a problem hiding this comment.
0 issues found across 3 files (changes from recent commits).
Shadow auto-approve: would not auto-approve. Auto-approval blocked by 1 unresolved issue from previous reviews.
Re-trigger cubic
…o fix/gzip-decompression-memory
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
There was a problem hiding this comment.
No issues found across 4 files
Shadow auto-approve: would not auto-approve. This PR does not meet the repository auto-approval settings.
Re-trigger cubic
What
Gzip uploads are decompressed in 1 MiB chunks into a temporary file that stays in memory up to 1 MiB and spills to disk above that. Previously the whole expanded payload was held in memory.
UploadFile, so FastAPI closes it when the request finishes. Partial output is closed if the gzip data is corrupt or truncated.Why
ungz_fileheld the expanded payload as onebytesobject for the whole request, so peak memory grew with the uncompressed size.An earlier revision passed large outputs downstream behind a proxy object, to stop
unstructuredfrom copying them back into memory. That proxy madeconvert_to_bytes()reject the file:ValueError: Invalid file-like object typefor gzipped uploads over 1 MiB on the PDF OCR and text-encoding-detection paths. This revision passes the spool itself. Avoiding the downstream copies now comes from unstructured#4419 (open; expected to ship as 0.27.12).Merge dependency
Stacked on #579. Released as 0.1.13, after #579's 0.1.12. The decompression-stage savings apply on any
unstructuredversion. The end-to-end savings below need the unstructured#4419 release (expected 0.27.12); this PR does not change the lock. On 0.22.18 the DOCX/PPTX/PDF partitioners still copy the spool into memory.Impact
Measured on the earlier revision, whose downstream path matches unstructured#4419. Responses were byte-identical:
Risk
A very large expansion now uses temporary disk instead of RAM. Disk capacity is the external bound, and a decompressed-size cap would be a product decision for the upload-limits work.
Validation
test_gzip.py+test_filetypes.pypass onunstructured0.22.18, including an end-to-end gzipped request that asserts the decompressed file is closed after the response (it fails againstmain'sungz_file). The same tests pass with the unstructured#4419 branch installed over the locked dependencies. That includes a regression test that reads a disk-backed decompressed upload throughconvert_to_bytes(); it fails with the error above on the earlier proxy revision.