You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 2693685
Browse filesBrowse the repository at this point in the historyBrowse files
| Office document reads | Read bounded DOCX paragraph windows and XLSX cell ranges through native tools |
28
29
29
30
Product code depends only on `ManagedCode.Storage.Core`, so concrete storage providers stay in your application. The integration suite exercises the real filesystem provider; other backends use the same `IStorage` contract.
30
31
@@ -45,7 +46,7 @@ flowchart LR
45
46
Requires **.NET 10**. Add FileContext and the storage provider your application uses:
@@ -134,6 +135,7 @@ Standard `file_access_*` tools come from Agent Framework's `FileAccessProvider`.
134
135
|`file_context_pdf_page_image`| Render one complete PDF page as PNG `DataContent`| Yes |
135
136
|`file_context_pdf_images_info`| Count embedded images on one PDF page | Yes |
136
137
|`file_context_pdf_image`| Extract one embedded PDF image as PNG `DataContent`| Yes |
138
+
|`file_context_docx_text`| Read DOCX paragraph windows with a continuation cursor | Yes |
137
139
|`file_context_info`| Return file presence and metadata without reading content | Yes |
138
140
|`file_context_markdown_graph_search`| Build and ranked-search a Markdown knowledge graph | Yes |
139
141
|`file_context_markdown_graph_export`| Export a graph as Mermaid, DOT, Turtle, or JSON-LD | Yes |
@@ -187,6 +189,8 @@ Results include `StartLine`, `EndLine`, `HasMore`, and `TotalLines` when the end
187
189
188
190
For an authenticated PDF already held as bytes, `FileContextPdfTextExtractor.Extract`, `FileContextPdfImages.RenderPagePng`, and `FileContextPdfImages.ExtractPageImagesPng` work without storing it. Storage reads enforce `MaximumPdfReadBytes` (25 MiB by default); page rasterization caps pixels and PNG size. A host must pass image `DataContent` to its model as image content. A generic OpenAI Chat function result serializes it as text, so hosts must explicitly bridge image tool results into a multimodal model message.
189
191
192
+
`file_context_docx_text(path, startParagraph?, startCharacter?, paragraphCount?)` reads ordinary paragraph and table text from a scoped DOCX package. The result contains numbered paragraph segments and `nextParagraph`/`nextCharacter`; use that cursor to continue a long document. Reads are limited to 50 paragraphs and 20,000 characters per call, with a configurable 25 MiB source limit (`MaximumDocxReadBytes`). It does not OCR embedded images. DOCX and XLSX packages are excluded from generic text reads and grep.
193
+
190
194
## Explore Markdown as a graph
191
195
192
196
Use [ManagedCode.MarkdownLd.Kb](https://github.com/managedcode/markdown-ld-kb) to connect and search concepts across the Markdown documents in your workspace:
Version `1.0.8` is defined centrally in `Directory.Build.props`. Every push to `main` runs the Release workflow: restore, format, build, test with coverage, and pack. For a new package version, it publishes the validated NuGet artifact and creates the matching tag and GitHub release automatically. Already released versions are skipped. To release an update, bump the version, commit, and push; no manual tag is required.
366
+
Version `1.0.9` is defined centrally in `Directory.Build.props`. Every push to `main` runs the Release workflow: restore, format, build, test with coverage, and pack. For a new package version, it publishes the validated NuGet artifact and creates the matching tag and GitHub release automatically. Already released versions are skipped. To release an update, bump the version, commit, and push; no manual tag is required.
363
367
364
368
[MIT licensed](https://github.com/managedcode/FileContext/blob/main/LICENSE) · Built by [ManagedCode](https://github.com/managedcode)
Copy file name to clipboardExpand all lines: docs/Features/file-context.md
+9-3Lines changed: 9 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -4,7 +4,7 @@
4
4
5
5
ManagedCode.FileContext lets an Agent Framework agent work with files from any ManagedCode.Storage provider and query Markdown files as a knowledge graph.
6
6
7
-
In scope: standard file access, bounded line navigation, metadata, Markdown graph retrieval/export, DI, path isolation, and deterministic tool-loop testing. Out of scope: provider credentials, direct model hosting, binary document parsing, and an alternative file protocol.
7
+
In scope: standard file access, bounded line navigation, metadata, native PDF/DOCX/XLSX reads, Markdown graph retrieval/export, DI, path isolation, and deterministic tool-loop testing. Out of scope: provider credentials, direct model hosting, and an alternative file protocol.
8
8
9
9
## Rules
10
10
@@ -23,7 +23,7 @@ In scope: standard file access, bounded line navigation, metadata, Markdown grap
23
23
13. The context provider injects capability instructions and tools, not arbitrary file content as system instructions.
24
24
14. DI supports both the default `IStorage` and a named/keyed `IStorage` registration.
25
25
15. Optional `OperationTimeout` applies to each public storage/context operation, combining with caller cancellation and preserving one deadline across internal steps. It defaults to null; regex matching has its separate `RegexTimeout`.
26
-
16. The NuGet package has version `1.0.8`; publication occurs only from the GitHub Actions release workflow.
26
+
16. The NuGet package has version `1.0.9`; publication occurs only from the GitHub Actions release workflow.
27
27
28
28
## Main flow
29
29
@@ -87,7 +87,7 @@ Independent writes and range reads on eight different files are tested concurren
87
87
9. A real Agent Framework loop receives an LlmTck tool call, executes storage-backed `file_access_read`, proves the file content reaches the second model request, and returns the expected final answer.
88
88
10. LlmTck tool loops exercise every read-only, mutation, and extended tool against the real filesystem provider.
89
89
11. A sparse 1 GiB file supports bounded repeated range reads without proportional allocation; a giant unterminated line fails at the configured byte boundary.
90
-
12. The packed `1.0.8` package installs and runs in a clean smoke project.
90
+
12. The packed `1.0.9` package installs and runs in a clean smoke project.
91
91
92
92
## Definition of done
93
93
@@ -115,4 +115,10 @@ Verification: DocumentCreationTests and DocumentValidationTests reopen real form
115
115
116
116
## PDF reads and vision images
117
117
118
+
DOCX reading uses the native `file_context_docx_text` tool. It reads ordinary paragraph and table
119
+
text from the scoped `.docx` package in bounded windows. Each result includes the next paragraph
120
+
and character offset when more text remains, so an agent can continue without loading a long
121
+
document into one model call. The tool does not execute macros, fetch external links, or perform
122
+
OCR on embedded images. Generic text reads refuse DOCX, and text search skips it.
123
+
118
124
`file_context_pdf_text` reports a bounded text-layer prefix, total page count, and one-based pages with almost no text. It performs no OCR. `file_context_pdf_page_image` renders a complete page as PNG. `file_context_pdf_images_info` counts embedded image objects, and `file_context_pdf_image` returns one object as PNG. The direct `IFileContextPdf` methods and public byte-oriented PDF APIs support the same operations. Storage-scoped PDF reads enforce a byte cap; page rendering enforces pixel and image-byte caps. Image tools return `DataContent`; host chat pipelines must forward it as image content rather than stringify a function result.
0 commit comments