You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit d00a1ba
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: README.md
+6Lines changed: 6 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -185,6 +185,12 @@ Results include `StartLine`, `EndLine`, `HasMore`, and `TotalLines` when the end
185
185
186
186
## Read PDFs and send pages to vision models
187
187
188
+
Storage-backed PDF tools stage seekable cloud streams asynchronously before parsing. PdfPig's
189
+
synchronous reads and seeks then use a temporary file rather than repeated network ranges. Local
190
+
seekable files and memory streams remain reusable. `PdfSourceStagingMode` can force `TemporaryFile`
191
+
staging, `PdfSourceBufferBytes` bounds the copy buffer, and `PdfTemporaryDirectory` optionally selects
192
+
an existing host directory. Disposal, cancellation and staging failures remove temporary sources.
193
+
188
194
`IFileContextPdf.ReadPdfTextAsync(path)` returns bounded text, `PageCount`, and one-based `PagesWithoutText`. It does not perform OCR. A scanned page can instead be rendered with `RenderPdfPageAsync(path, pageNumber)`, which returns PNG `DataContent`. Use `CountPdfPageImagesAsync` and `ExtractPdfImageAsync` when the original embedded pictures are needed rather than the complete page. The four read-only `file_context_pdf_*` tools expose the same operations from scoped storage.
189
195
190
196
For an authenticated PDF already held as bytes, `FileContextPdfTextExtractor.Extract`, `FileContextPdfImages.RenderPagePng`, and `FileContextPdfImages.ExtractPageImagesPng` work without storing it. PDF source reads default to 100 MiB and accept `FileContextOptions` for a different limit; page rasterization also uses configured pixel and PNG limits. `FileContextImageContent` creates model-visible `DataContent` from PNG bytes or base64 and `UriContent` from an HTTPS URL. A URL reference is not fetched by FileContext, so the model provider must be able to access it. A host must pass image content to its model as image content. A generic OpenAI Chat function result serializes it as text, so hosts must explicitly bridge image tool results into a multimodal model message.
Copy file name to clipboardExpand all lines: docs/Architecture.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -82,7 +82,7 @@ source and retaining it through parsing/rendering. `MaximumConcurrentPdfOperatio
82
82
`FileContextPdfRenderDocument` exposes page count and sequential page rendering from one parsed PDF;
83
83
dispose it after the batch. The low-level synchronous image helpers remain caller-scheduled APIs.
84
84
85
-
PDF text, page rendering and embedded-image APIs accept bounded seekable streams. Storage-backed PDF tools keep seekable provider streams directly and stage non-seekable sources to an automatically deleted temporary file, never a whole-document managed array. Native raster decoding has a separate per-page source-image pixel budget; lowering output scale does not reduce source bitmap allocation.
85
+
PDF text, page rendering and embedded-image APIs accept bounded seekable streams. Storage-backed PDF tools retain seekable local `FileStream`/`MemoryStream` inputs but asynchronously stage other streams, including seekable cloud streams, to an automatically deleted temporary file. PdfPig's synchronous byte reads and seeks then remain local, without blocking on repeated network ranges. `PdfSourceStagingMode.TemporaryFile` also stages local inputs. `PdfSourceBufferBytes` bounds every copy read and `PdfTemporaryDirectory` optionally selects an existing host directory. No whole-document managed array is created. Native raster decoding has a separate per-page source-image pixel budget; lowering output scale does not reduce source bitmap allocation.
86
86
87
87
All potentially large operations are controlled by `IOptions<FileContextOptions>`: PDF source/page/image budgets, full-read bytes, range bytes, files scanned, bytes per searched file, matches per file, total search results, graph documents, graph source bytes, and exported graph characters. Non-seekable cloud streams are supported by sequential streaming.
Copy file name to clipboardExpand all lines: docs/Testing/index.md
+1Lines changed: 1 addition & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -10,6 +10,7 @@ The suite is integration-first:
10
10
- timeout tests cover configured operation expiry, cancellation of every public operation, disabled deadlines, duration validation, and timeout tool results through restored sessions;
11
11
- concurrent storage tests write and range-read eight independent files through one shared adapter/service;
12
12
- a sparse 1 GiB filesystem test reads bounded line windows repeatedly, rejects full-file loading, caps allocations, and proves that an oversized line fails before it can be buffered in memory.
13
+
- PDF cloud-source tests use real files behind an async-only seekable stream, proving parsing uses the staged local file; they cover large inputs, configured buffers/directories, forced staging, local-file reuse, limits, mid-copy cancellation, failure cleanup and private Unix permissions.
13
14
14
15
Every filesystem test owns a unique temporary root and removes it on disposal. Test execution is serialized so process-wide allocation assertions cannot be distorted by another test. No `IStorage`, Agent Framework, Markdown-LD, or LlmTck mocks are used.
0 commit comments