Skip to content

Commit 2693685

Browse files
committed
Add bounded DOCX text tool to FileContext
1 parent 3cc5b64 commit 2693685

21 files changed

Lines changed: 312 additions & 15 deletions

‎CHANGELOG.md‎

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -2,6 +2,11 @@
22

33
All notable changes to ManagedCode.FileContext are documented here.
44

5+
## 1.0.9
6+
7+
- Add a native bounded DOCX text tool with paragraph and character cursors for long documents.
8+
- Read DOCX through scoped storage and the pinned Open XML SDK; exclude binary DOCX packages from generic text reads and search.
9+
510
## 1.0.8
611

712
- Add bounded PDF text/page metadata, full-page PNG rendering, and embedded-image extraction through scoped file tools and public byte APIs.

‎Directory.Build.props‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -12,7 +12,7 @@
1212
<AnalysisMode>Recommended</AnalysisMode>
1313
<TreatWarningsAsErrors>true</TreatWarningsAsErrors>
1414
<NoWarn>$(NoWarn);CS1591;MAAI001</NoWarn>
15-
<Version>1.0.8</Version>
15+
<Version>1.0.9</Version>
1616
<PackageVersion>$(Version)</PackageVersion>
1717
</PropertyGroup>
1818

‎README.md‎

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -25,6 +25,7 @@ Your host supplies an `IStorage` backend and a model client. FileContext supplie
2525
| Multiple files | Issue independent tool calls in one turn, with optional concurrent execution |
2626
| Workspace isolation | Resolve logical paths under a configured storage prefix |
2727
| Observable results | Receive structured `found` / `not_found` metadata results that survive session restoration |
28+
| Office document reads | Read bounded DOCX paragraph windows and XLSX cell ranges through native tools |
2829

2930
Product code depends only on `ManagedCode.Storage.Core`, so concrete storage providers stay in your application. The integration suite exercises the real filesystem provider; other backends use the same `IStorage` contract.
3031

@@ -45,7 +46,7 @@ flowchart LR
4546
Requires **.NET 10**. Add FileContext and the storage provider your application uses:
4647

4748
```bash
48-
dotnet add package ManagedCode.FileContext --version 1.0.8
49+
dotnet add package ManagedCode.FileContext --version 1.0.9
4950
dotnet add package ManagedCode.Storage.FileSystem --version 10.0.7
5051
```
5152

@@ -134,6 +135,7 @@ Standard `file_access_*` tools come from Agent Framework's `FileAccessProvider`.
134135
| `file_context_pdf_page_image` | Render one complete PDF page as PNG `DataContent` | Yes |
135136
| `file_context_pdf_images_info` | Count embedded images on one PDF page | Yes |
136137
| `file_context_pdf_image` | Extract one embedded PDF image as PNG `DataContent` | Yes |
138+
| `file_context_docx_text` | Read DOCX paragraph windows with a continuation cursor | Yes |
137139
| `file_context_info` | Return file presence and metadata without reading content | Yes |
138140
| `file_context_markdown_graph_search` | Build and ranked-search a Markdown knowledge graph | Yes |
139141
| `file_context_markdown_graph_export` | Export a graph as Mermaid, DOT, Turtle, or JSON-LD | Yes |
@@ -187,6 +189,8 @@ Results include `StartLine`, `EndLine`, `HasMore`, and `TotalLines` when the end
187189

188190
For an authenticated PDF already held as bytes, `FileContextPdfTextExtractor.Extract`, `FileContextPdfImages.RenderPagePng`, and `FileContextPdfImages.ExtractPageImagesPng` work without storing it. Storage reads enforce `MaximumPdfReadBytes` (25 MiB by default); page rasterization caps pixels and PNG size. A host must pass image `DataContent` to its model as image content. A generic OpenAI Chat function result serializes it as text, so hosts must explicitly bridge image tool results into a multimodal model message.
189191

192+
`file_context_docx_text(path, startParagraph?, startCharacter?, paragraphCount?)` reads ordinary paragraph and table text from a scoped DOCX package. The result contains numbered paragraph segments and `nextParagraph`/`nextCharacter`; use that cursor to continue a long document. Reads are limited to 50 paragraphs and 20,000 characters per call, with a configurable 25 MiB source limit (`MaximumDocxReadBytes`). It does not OCR embedded images. DOCX and XLSX packages are excluded from generic text reads and grep.
193+
190194
## Explore Markdown as a graph
191195

192196
Use [ManagedCode.MarkdownLd.Kb](https://github.com/managedcode/markdown-ld-kb) to connect and search concepts across the Markdown documents in your workspace:
@@ -359,7 +363,7 @@ dotnet pack src/ManagedCode.FileContext/ManagedCode.FileContext.csproj --configu
359363

360364
## Releases and license
361365

362-
Version `1.0.8` is defined centrally in `Directory.Build.props`. Every push to `main` runs the Release workflow: restore, format, build, test with coverage, and pack. For a new package version, it publishes the validated NuGet artifact and creates the matching tag and GitHub release automatically. Already released versions are skipped. To release an update, bump the version, commit, and push; no manual tag is required.
366+
Version `1.0.9` is defined centrally in `Directory.Build.props`. Every push to `main` runs the Release workflow: restore, format, build, test with coverage, and pack. For a new package version, it publishes the validated NuGet artifact and creates the matching tag and GitHub release automatically. Already released versions are skipped. To release an update, bump the version, commit, and push; no manual tag is required.
363367

364368
[MIT licensed](https://github.com/managedcode/FileContext/blob/main/LICENSE) · Built by [ManagedCode](https://github.com/managedcode)
365369

‎docs/Architecture.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -106,6 +106,8 @@ flowchart LR
106106
`FileContextDocumentService` opens XLSX packages read-only from the scoped storage adapter.
107107
`FileContextWorkbookReader` maps worksheets and cell rectangles into package-owned result records.
108108
It preserves coordinates and stored types without evaluating formulas or accessing external links.
109+
The same service opens DOCX packages read-only and streams bounded paragraph text through
110+
`FileContextDocxReader`, returning a cursor for later paragraph windows.
109111

110112
```mermaid
111113
flowchart LR

‎docs/Features/file-context.md‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44

55
ManagedCode.FileContext lets an Agent Framework agent work with files from any ManagedCode.Storage provider and query Markdown files as a knowledge graph.
66

7-
In scope: standard file access, bounded line navigation, metadata, Markdown graph retrieval/export, DI, path isolation, and deterministic tool-loop testing. Out of scope: provider credentials, direct model hosting, binary document parsing, and an alternative file protocol.
7+
In scope: standard file access, bounded line navigation, metadata, native PDF/DOCX/XLSX reads, Markdown graph retrieval/export, DI, path isolation, and deterministic tool-loop testing. Out of scope: provider credentials, direct model hosting, and an alternative file protocol.
88

99
## Rules
1010

@@ -23,7 +23,7 @@ In scope: standard file access, bounded line navigation, metadata, Markdown grap
2323
13. The context provider injects capability instructions and tools, not arbitrary file content as system instructions.
2424
14. DI supports both the default `IStorage` and a named/keyed `IStorage` registration.
2525
15. Optional `OperationTimeout` applies to each public storage/context operation, combining with caller cancellation and preserving one deadline across internal steps. It defaults to null; regex matching has its separate `RegexTimeout`.
26-
16. The NuGet package has version `1.0.8`; publication occurs only from the GitHub Actions release workflow.
26+
16. The NuGet package has version `1.0.9`; publication occurs only from the GitHub Actions release workflow.
2727

2828
## Main flow
2929

@@ -87,7 +87,7 @@ Independent writes and range reads on eight different files are tested concurren
8787
9. A real Agent Framework loop receives an LlmTck tool call, executes storage-backed `file_access_read`, proves the file content reaches the second model request, and returns the expected final answer.
8888
10. LlmTck tool loops exercise every read-only, mutation, and extended tool against the real filesystem provider.
8989
11. A sparse 1 GiB file supports bounded repeated range reads without proportional allocation; a giant unterminated line fails at the configured byte boundary.
90-
12. The packed `1.0.8` package installs and runs in a clean smoke project.
90+
12. The packed `1.0.9` package installs and runs in a clean smoke project.
9191

9292
## Definition of done
9393

@@ -115,4 +115,10 @@ Verification: DocumentCreationTests and DocumentValidationTests reopen real form
115115

116116
## PDF reads and vision images
117117

118+
DOCX reading uses the native `file_context_docx_text` tool. It reads ordinary paragraph and table
119+
text from the scoped `.docx` package in bounded windows. Each result includes the next paragraph
120+
and character offset when more text remains, so an agent can continue without loading a long
121+
document into one model call. The tool does not execute macros, fetch external links, or perform
122+
OCR on embedded images. Generic text reads refuse DOCX, and text search skips it.
123+
118124
`file_context_pdf_text` reports a bounded text-layer prefix, total page count, and one-based pages with almost no text. It performs no OCR. `file_context_pdf_page_image` renders a complete page as PNG. `file_context_pdf_images_info` counts embedded image objects, and `file_context_pdf_image` returns one object as PNG. The direct `IFileContextPdf` methods and public byte-oriented PDF APIs support the same operations. Storage-scoped PDF reads enforce a byte cap; page rendering enforces pixel and image-byte caps. Image tools return `DataContent`; host chat pipelines must forward it as image content rather than stringify a function result.
Lines changed: 24 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,24 @@
1+
using DocumentFormat.OpenXml.Packaging;
2+
3+
namespace ManagedCode.FileContext;
4+
5+
public sealed partial class FileContextDocumentService
6+
{
7+
public Task<FileContextDocxText> ReadDocxTextAsync(string path, int startParagraph = 1,
8+
int startCharacter = 0, int paragraphCount = 20, CancellationToken cancellationToken = default)
9+
{
10+
RequireExtension(path, ".docx");
11+
ArgumentOutOfRangeException.ThrowIfLessThan(startParagraph, 1);
12+
ArgumentOutOfRangeException.ThrowIfNegative(startCharacter);
13+
ArgumentOutOfRangeException.ThrowIfLessThan(paragraphCount, 1);
14+
ArgumentOutOfRangeException.ThrowIfGreaterThan(paragraphCount,
15+
FileContextDefaults.MaximumDocxParagraphsPerRead);
16+
17+
return ReadSourceAsync(path, (buffer, token) =>
18+
{
19+
using var document = WordprocessingDocument.Open(buffer, false);
20+
return FileContextDocxReader.Read(document, path, startParagraph, startCharacter,
21+
paragraphCount, token);
22+
}, options.MaximumDocxReadBytes, cancellationToken);
23+
}
24+
}

‎src/ManagedCode.FileContext/Documents/FileContextDocumentService.Reading.cs‎

Lines changed: 5 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -23,15 +23,16 @@ private Task<T> ReadWorkbookAsync<T>(string path, Func<SpreadsheetDocument, Canc
2323
{
2424
using var document = SpreadsheetDocument.Open(buffer, false);
2525
return read(document, token);
26-
}, cancellationToken);
26+
}, options.MaximumFullReadBytes, cancellationToken);
2727
}
2828

29-
private Task<T> ReadSourceAsync<T>(string path, Func<MemoryStream, CancellationToken, T> read, CancellationToken cancellationToken) =>
29+
private Task<T> ReadSourceAsync<T>(string path, Func<MemoryStream, CancellationToken, T> read,
30+
long maximumBytes, CancellationToken cancellationToken) =>
3031
FileContextOperation.RunAsync(options.OperationTimeout, async token =>
3132
{
3233
var metadata = await store.GetMetadataAsync(path, token).ConfigureAwait(false)
3334
?? throw new FileNotFoundException("The document was not found in this file context.", path);
34-
if (metadata.Length > (ulong)options.MaximumFullReadBytes)
35+
if (metadata.Length > (ulong)maximumBytes)
3536
{ throw new IOException("Document exceeds the configured source read budget."); }
3637
var source = await store.OpenReadAsync(path, token).ConfigureAwait(false);
3738
await using var lifetime = source.ConfigureAwait(false);
@@ -40,7 +41,7 @@ private Task<T> ReadSourceAsync<T>(string path, Func<MemoryStream, CancellationT
4041
int count;
4142
while ((count = await source.ReadAsync(chunk, token).ConfigureAwait(false)) > 0)
4243
{
43-
if (buffer.Length + count > options.MaximumFullReadBytes)
44+
if (buffer.Length + count > maximumBytes)
4445
{ throw new IOException("Document exceeds the configured source read budget."); }
4546
buffer.Write(chunk, 0, count);
4647
}

‎src/ManagedCode.FileContext/Documents/FileContextDocumentService.Tables.cs‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@ public Task<FileContextTablesInfo> GetTablesInfoAsync(string path, int? headerRo
77
{
88
FileContextTableReader.Validate(path, headerRow);
99
return ReadSourceAsync(path,
10-
(buffer, token) => FileContextTableReader.Read(buffer, path, headerRow, delimiter, options, token), cancellationToken);
10+
(buffer, token) => FileContextTableReader.Read(buffer, path, headerRow, delimiter, options, token),
11+
options.MaximumFullReadBytes, cancellationToken);
1112
}
1213
}
Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,4 @@
1+
namespace ManagedCode.FileContext;
2+
3+
/// <summary>One paragraph or a segment of a long paragraph; character offsets are zero-based.</summary>
4+
public sealed record FileContextDocxParagraph(int Number, int CharacterOffset, string Text);
Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,56 @@
1+
using DocumentFormat.OpenXml;
2+
using DocumentFormat.OpenXml.Packaging;
3+
using DocumentFormat.OpenXml.Wordprocessing;
4+
5+
namespace ManagedCode.FileContext;
6+
7+
internal static class FileContextDocxReader
8+
{
9+
public static FileContextDocxText Read(WordprocessingDocument document, string path,
10+
int startParagraph, int startCharacter, int paragraphCount, CancellationToken cancellationToken)
11+
{
12+
var main = document.MainDocumentPart
13+
?? throw new InvalidDataException("The DOCX has no main document part.");
14+
using var reader = OpenXmlReader.Create(main);
15+
var paragraphs = new List<FileContextDocxParagraph>();
16+
var number = 0;
17+
var remaining = FileContextDefaults.MaximumDocxTextCharacters;
18+
while (reader.Read())
19+
{
20+
cancellationToken.ThrowIfCancellationRequested();
21+
if (!reader.IsStartElement || reader.ElementType != typeof(Paragraph))
22+
{
23+
continue;
24+
}
25+
26+
number++;
27+
if (number < startParagraph)
28+
{
29+
continue;
30+
}
31+
if (paragraphs.Count == paragraphCount || remaining == 0)
32+
{
33+
return new FileContextDocxText(path, paragraphs, number, 0);
34+
}
35+
36+
var paragraph = reader.LoadCurrentElement() as Paragraph
37+
?? throw new InvalidDataException("A DOCX paragraph could not be read.");
38+
var text = paragraph.InnerText;
39+
var offset = number == startParagraph ? startCharacter : 0;
40+
if (offset > text.Length)
41+
{
42+
throw new ArgumentOutOfRangeException(nameof(startCharacter),
43+
"The character offset is beyond this paragraph.");
44+
}
45+
var length = Math.Min(text.Length - offset, remaining);
46+
paragraphs.Add(new FileContextDocxParagraph(number, offset, text.Substring(offset, length)));
47+
remaining -= length;
48+
if (offset + length < text.Length)
49+
{
50+
return new FileContextDocxText(path, paragraphs, number, offset + length);
51+
}
52+
}
53+
54+
return new FileContextDocxText(path, paragraphs, null, 0);
55+
}
56+
}

0 commit comments

Comments
 (0)