Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
44 commits
Select commit Hold shift + click to select a range
02c891a
docs: record audio asset worker brief (MEET-1)
eliotlim Oct 3, 2026
451ea1d
feat(server,sdk): support audio assets (MEET-1)
eliotlim Oct 3, 2026
69f01e0
feat(server,sdk): add asset transcription service and backend (MEET-2)
eliotlim Oct 3, 2026
a535c2f
test(server): keep OIDC gate fixture PAT unexpired (MEET-1)
eliotlim Oct 3, 2026
5c04978
test(server): keep OIDC access-gate PAT fixture valid (MEET-2)
eliotlim Oct 3, 2026
fb16286
docs: report MEET-2 implementation and verification blocker
eliotlim Oct 3, 2026
f089e02
docs: report audio asset verification (MEET-1)
eliotlim Oct 3, 2026
0824d45
chore: drop worker docs (MEET-1)
eliotlim Oct 3, 2026
191543e
docs(sdk): fix asset allowlist comment (MEET-1)
eliotlim Oct 3, 2026
79f576a
docs: record green serialized verification for MEET-2
eliotlim Oct 3, 2026
bbccdbc
chore: drop worker artifacts (MEET-2)
eliotlim Oct 3, 2026
888f895
feat(sdk,mcp): register meeting block contract and validators (MEET-4)
eliotlim Oct 3, 2026
1f97ba0
docs: record MEET-4 verification and representation contract
eliotlim Oct 3, 2026
6e5e12d
chore: merge audio assets for MEET-5
eliotlim Oct 3, 2026
5ac196b
Merge branch 'feat/meet-2-transcribe' into feat/meet-5-recorder-ui
eliotlim Oct 3, 2026
aadffde
feat(ui): add meeting recording, transcription and notes (MEET-5)
eliotlim Oct 3, 2026
17458a3
fix(ui): ignore stale microphone permission failures (MEET-5)
eliotlim Oct 3, 2026
970983e
docs: report meeting recorder completion (MEET-5)
eliotlim Oct 3, 2026
490c499
chore: drop worker docs (MEET-4)
eliotlim Oct 3, 2026
9da6774
fix(ui): address meeting recorder review findings (MEET-5)
eliotlim Oct 3, 2026
4d27b95
chore: drop worker artifacts (MEET-5)
eliotlim Oct 3, 2026
e7da15b
fix(ui): keep focus after stopping a recording (MEET-5)
eliotlim Oct 3, 2026
0a2adfa
feat(server): add managed local whisper transcription (MEET-3)
eliotlim Oct 3, 2026
67c08f0
docs(server): report local whisper verification (MEET-3)
eliotlim Oct 3, 2026
07ced65
fix(server): bound local transcription work (MEET-3)
eliotlim Oct 3, 2026
0a49bb1
chore: drop worker artifacts (MEET-3)
eliotlim Oct 3, 2026
685c25c
feat(desktop): enable meeting microphone access (MEET-8)
eliotlim Oct 4, 2026
36c0ae0
feat(ui): add AI meeting summaries (MEET-6)
eliotlim Oct 4, 2026
33a5586
fix(ui): summary cancel, 403 copy and export paragraphs (MEET-6)
eliotlim Oct 4, 2026
2389a1b
feat(ui): export meeting audio and include document assets (MEET-7)
eliotlim Oct 4, 2026
9d82832
chore(ui): merge meeting summary review fixes (MEET-7)
eliotlim Oct 4, 2026
4923c9a
fix(ui): preserve meeting markdown spacing without audio (MEET-7)
eliotlim Oct 4, 2026
b8ca146
fix(ui): lexicographic audio export names + export cleanups (MEET-7)
eliotlim Oct 4, 2026
c37c1c4
test(web): cover meeting epic and document workflows (MEET-9)
eliotlim Oct 4, 2026
938161c
docs: point local-transcription caveats at MEET-3 (MEET-9)
eliotlim Oct 4, 2026
78dcd9d
Merge branch 'feat/meet-1-audio-assets' into feat/meet-2-transcribe
eliotlim Oct 4, 2026
9082527
Merge branch 'feat/meet-2-transcribe' into feat/meet-4-meeting-catalogue
eliotlim Oct 4, 2026
9b10912
Merge branch 'feat/meet-4-meeting-catalogue' into feat/meet-5-recorde…
eliotlim Oct 4, 2026
23370d4
Merge branch 'feat/meet-5-recorder-ui' into feat/meet-3-local-whisper
eliotlim Oct 4, 2026
759cc8f
Merge branch 'feat/meet-3-local-whisper' into feat/meet-8-desktop-mic
eliotlim Oct 4, 2026
28e025b
Merge branch 'feat/meet-8-desktop-mic' into feat/meet-6-ai-summary
eliotlim Oct 4, 2026
bef72f5
Merge branch 'feat/meet-6-ai-summary' into feat/meet-7-audio-export
eliotlim Oct 4, 2026
629c0c9
Merge branch 'feat/meet-7-audio-export' into feat/meet-9-e2e-docs
eliotlim Oct 4, 2026
9945263
chore: merge main into feat/meet-9-e2e-docs
eliotlim Oct 5, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -309,6 +309,38 @@ The FORM-7 MCP surface in `packages/mcp/src/server.ts` provides `list_forms`,
through the resolved per-page agent-edits policy (suggest by default), and key
regeneration is intentionally author-UI-only.

### Meeting blocks

Meetings are native `type:'meeting'` container blocks registered by
`packages/ui/src/blockeditor/MeetingBlockView.tsx`. The
[representation contract](docs/meeting-block.md) defines audio asset references,
timestamped transcript segments, summary, title, and status props. Manual notes
remain ordinary CRDT child blocks; generated transcript and summary are prop
snapshots. The user workflow is in [meeting notes](docs/meeting-notes.md).

`MeetingRecorder` restarts MediaRecorder every 45 seconds and on pause/resume.
Each uploaded chunk has its own container header, unlike recorder timeslices,
so playback, retries, transcription, and export work on standalone files.
Uploads use page-associated assets; `POST /api/ai/transcribe` receives the asset
and page IDs, enforces access, and records usage. Completed chunks append
transcript segments progressively. Summary generation is an explicit streamed
AI request; cancellation preserves the previous durable summary.

Transcription resolves separately from chat: explicit off rejects; an explicit
OpenAI-compatible transcription provider opts into that endpoint; otherwise the
local resolver runs, followed by the deterministic mock fallback only when chat
provider is mock. An unavailable local engine returns a configuration error,
never implicit cloud fallback. Local Whisper transcription ships by default; **Settings → AI** provides the
model download. See [local transcription setup](docs/local-transcription.md) for
Whisper and FFmpeg installation and runtime requirements.

Audio export downloads a single original file or an ordered timestamped ZIP of
chunks, without remuxing or deleting library assets. Markdown and HTML exports
include transcript, summary, and child notes with audio references; audio bytes
are exported separately. The browser epic proof is
`packages/web/e2e/meeting-epic.spec.ts`, using a WebAudio microphone substitute
with real recording, asset, transcription, generation, and download paths.

### Optional local AI (`packages/server/src/ai/`)

An opt-in, local-only model subsystem (Settings → AI). Pluggable engines
Expand Down
1 change: 1 addition & 0 deletions docs/audits/block-api-coverage.md
Original file line number Diff line number Diff line change
Expand Up @@ -46,6 +46,7 @@
| tooltipcard | ✅ | ✅ | ➖ | ✅ | ➖ | ➖ |
| dbview | ✅ | ✅ | ➖ | ✅ | ➖ | ➖ |
| dbform | ✅ | ✅ | ➖ | ✅ | ➖ | ➖ |
| meeting | ✅ | ✅ | ➖ | ✅ | ➖ | ➖ |
| form | ✅ | ✅ | ➖ | ✅ | ➖ | ➖ |
| openbook.ledger/journal-entry | ✅ | ✅ | ➖ | ✅ | ➖ | ✅ |
| openbook.ledger/trial-balance | ✅ | ✅ | ➖ | ✅ | ➖ | ✅ |
Expand Down
44 changes: 44 additions & 0 deletions docs/local-transcription.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Local transcription

OpenBook transcribes recordings locally by default, without a cloud API key. The server uses the optional whisper.cpp `whisper-cli` executable and FFmpeg. Neither executable nor model weights are bundled with OpenBook; normal CI skips native inference explicitly.

## Setup

1. Install whisper.cpp and FFmpeg on the server host. Follow the [whisper.cpp build instructions](https://github.com/ggml-org/whisper.cpp#quick-start) or use your host package manager.
2. Put `whisper-cli` and `ffmpeg` on the server process's PATH. Alternatively, set executable paths before starting OpenBook:

```sh
export OPENBOOK_WHISPER_BIN=/absolute/path/to/whisper-cli
export OPENBOOK_FFMPEG_BIN=/absolute/path/to/ffmpeg
```

3. In **Settings → AI**, select **Download Whisper base**. This uses the existing authenticated model download and progress flow. Downloads go to `OPENBOOK_MODELS_DIR` when set, otherwise the server data directory's `models` folder (or `~/.openbook/models` without a data directory). The completed model is discovered without restarting.
4. Check that Settings reports the runtime and model ready, then transcribe a recording.

The default model is multilingual Whisper base (`ggml-base.bin`, approximately 142 MiB), with automatic language detection. It trades some accuracy on noisy speech, accents, and difficult multilingual recordings for a smaller download and lower memory use than larger models. Whisper downloads do not change the selected chat model.

Explicit cloud transcription configuration takes precedence over local inference. Otherwise the server tries local transcription, then the existing mock fallback. Local transcription works even with chat disabled; explicitly disabling transcription still disables it. Missing executables or model weights produce an actionable error pointing to Settings → AI.

## Processing and limits

FFmpeg converts recordings to mono 16 kHz PCM WAV. Whisper loads the model for each job and returns text plus segments in seconds; `durationMs` is the rounded maximum segment endpoint in milliseconds. Each job uses a private temporary directory. Cancellation and server shutdown kill active child processes, wait for them to close, and remove scratch files.

Each server permits at most **two local transcription jobs at once**, shared across all clients. There is no queue. Busy requests return HTTP **429** with `Retry-After: 5`. Permits remain held through temporary-file cleanup and are released on success, failure, or cancellation.

The transcription route also allows **six local requests per socket IP per 60-second fixed window**. Excess requests return HTTP **429** with `Retry-After: 60`. Client-supplied forwarding headers do not change this key; clients behind a reverse proxy may share its socket IP budget. The limit applies only when the resolved backend is local; cloud and mock backends are unaffected.

## Native smoke test

Install the optional executables and download `ggml-base.bin` into your model directory, then run from the repository root:

```sh
OPENBOOK_TEST_WHISPER=1 \
OPENBOOK_MODELS_DIR=/absolute/path/to/models \
OPENBOOK_WHISPER_BIN=/absolute/path/to/whisper-cli \
OPENBOOK_FFMPEG_BIN=/absolute/path/to/ffmpeg \
VITEST_MAX_WORKERS=1 \
pnpm --filter @book.dev/server exec vitest run src/transcription.test.ts \
-t 'native whisper'
```

This opt-in test synthesizes a small WAV and exercises the HTTP transcription route without cloud keys, asserting the result shape. It is a runtime integration check, not a speech-accuracy benchmark. Without `OPENBOOK_TEST_WHISPER=1`, the native test is explicitly skipped. The regular subprocess-fixture tests cover concurrency, cancellation, failures, timestamp conversion, and cleanup without installing native dependencies.
44 changes: 44 additions & 0 deletions docs/meeting-block.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Meeting block — THE REPRESENTATION CONTRACT (MEET-4)

`meeting` is a `kit` block with `nature: 'container'`, no required parent, and
`kitValue: false`. It publishes no reactive value. Its `children` are ordinary
manual-note blocks (paragraphs, todos, headings, groups, etc.), not transcript
segments. No dedicated notes slot or child-only type is required.

All top-level props are optional. Missing status means `idle`; missing arrays
mean empty lists; missing summary/title mean empty text; missing startedAt means
not started. These are consumer defaults, not schema-inserted persisted values.
As with other block props, patches shallow-merge; `null` removes a top-level key.
Unknown top-level props remain allowed for forward compatibility.

| Prop | Stored shape and meaning |
| --- | --- |
| `status` | `'idle' \| 'recording' \| 'processing' \| 'done'`. Validation checks the enum, not state transitions. |
| `audioChunks` | Ordered array of `{assetId: string, durationMs: number, startedAtMs?: number}`. `assetId` is a nonempty asset reference (max 512 characters), never inline audio. `durationMs` is the chunk duration; optional `startedAtMs` is an offset from meeting start, **not** Unix time. Array order is capture/playback order. |
| `transcript` | Ordered array of `{startMs: number, endMs: number, text: string}`. Offsets are relative to meeting start; `endMs >= startMs`. Text is plain text. Array order is display order; overlapping segments are allowed. |
| `summary` | Plain string; no rich-text runs or Markdown interpretation is required. |
| `startedAt` | Unix epoch milliseconds, a finite nonnegative number. |
| `title` | Optional plain string. |

All durations and offsets are finite nonnegative numbers in milliseconds;
fractional milliseconds are accepted. Every listed nested field is required
except `startedAtMs`; nested objects reject extra keys and null fields. Empty
arrays and empty transcript text are valid. Cross-field `endMs >= startMs` is
checked at runtime and described in the JSON schema (standard JSON Schema cannot
express a comparison to a sibling field). Asset existence, timeline sorting,
status transitions, and transcription size limits are producer responsibilities.

Audio chunks and transcript segments use structured props, following existing
kit option/rich-run arrays and the form's structured schema prop. They are
machine-produced snapshots, while manual notes need independently editable CRDT
children. Each array is one prop value: updates replace the whole array, with
no per-segment merge guarantee. Recording/transcription producers should batch
updates and serialize writes to avoid concurrent array replacement losing data.
This keeps the first representation small and consistent; large-transcript
storage migration, if needed later, must explicitly version this contract.

MEET-4 registers a minimal meeting shell through `registerArtifactKit` in both
editor and viewer hosts. It displays title, status, and saved note-block count;
notes remain stored but are not yet rendered/editable inside the shell. It has
no slash-menu entry or recording controls. MEET-5 replaces this renderer and
must render the existing child blocks rather than migrate them into props.
63 changes: 63 additions & 0 deletions docs/meeting-notes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# Meeting notes

On a saved page, type `/meeting` and choose **Meeting**. Keep the page connected
to its OpenBook server so recordings can be saved to the library.

## Record and transcribe

Click **Record** and allow microphone access. OpenBook saves roughly 45-second
chunks and transcribes each completed upload; transcript lines appear as chunks
finish, without waiting for the entire meeting. Each chunk is a standalone audio
file, so it can be played or exported independently.

**Pause** closes the current chunk and silences capture. **Resume** starts a new
chunk; paused time is excluded from the recording timeline. **Stop** finishes
capture and lets pending uploads and transcription settle. Keep the page open
until processing completes. Failed uploads retain an in-session audio copy with
save/retry controls; transcription failures preserve uploaded audio and offer
**Retry transcription**. Leaving the page can lose audio that has not uploaded.

Local Whisper transcription ships by default, independently of the chat model.
In **Settings → AI**, select **Download Whisper base** to download the model.
See [local transcription](local-transcription.md) for the required Whisper and
FFmpeg installation, model setup, and runtime checks. If local transcription
is unavailable, recording and manual notes still work; OpenBook does not silently
fall back to a cloud service. Cloud transcription requires explicit opt-in in
**Settings → AI**. Selecting a cloud chat model alone does not opt audio into
cloud transcription.

## Summaries and notes

Configure a generation provider in **Settings → AI**, then click **Generate
summary** once transcript text exists. Summaries never start automatically.
Text appears while generation streams. **Cancel** discards the unfinished
replacement and retains any previous summary. **Regenerate summary** requests a
new version. Long transcripts are clipped to the first 4,000 characters for the
summary prompt, with an instruction to disclose that it covers an excerpt.

Type directly under **Notes**, including while recording. Notes are ordinary
editable blocks inside the meeting, independent of generated transcript and
summary text.

## Export and privacy

Click **Export audio** to download one audio file for a single saved chunk, or a
ZIP containing ordered, timestamped files for multiple chunks. Files retain their
original formats and bytes; the ZIP does not join or re-encode the recording.
Export leaves the original audio in your library. Audio is stored with the page's
assets and follows library/page access rules; it is not uploaded to a third-party
transcription service unless you explicitly configure one. With a remote library,
capture uploads to that library's server, so “local transcription” refers to
processing on the server rather than necessarily on the microphone's device.

Use **Page actions → Export → Markdown (.md)** for transcript timestamps,
summary, manual notes, and audio references. Markdown is not an audio backup;
use **Export audio** for the actual recordings.

## Desktop

Allow microphone access for OpenBook in your OS privacy settings and reopen the
app if permission changes require it. Browser permission instructions may still
appear in the shared UI. Follow the [desktop microphone manual checks](../packages/app/README.md#meeting-microphone-manual-checks)
to verify the actual webview and OS permission behavior; Chromium e2e coverage
uses synthetic audio and cannot establish desktop microphone support.
65 changes: 65 additions & 0 deletions packages/app/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,49 @@
From the workspace root, run it with `pnpm tauri dev`.
Create a desktop build with `pnpm build:desktop`.

## Meeting microphone (macOS)

`src-tauri/Info.plist` is merged into the bundle by Tauri and supplies
`NSMicrophoneUsageDescription`. The existing hardened-runtime entitlements file
now includes `com.apple.security.device.audio-input`. Recording is requested only
from the meeting block's Record action, via `getUserMedia({audio: true})`.

The locked Tauri 2.11.2 / Wry 0.55.1 already installs the Rust
[`WKUIDelegate` media-capture callback](https://github.com/tauri-apps/wry/blob/wry-v0.55.1/src/wkwebview/class/wry_web_view_ui_delegate.rs#L126-L137).
It grants the webview permission layer; macOS TCC still controls device access,
prompts for initial consent, and persists Allow/Deny. Do not replace Wry's UI
delegate or add a second permission store. This version does not expose the newer
`with_permission_handler` API. An OS denial reaches the recorder's existing
graceful microphone error and returns it to idle.

The recorder restarts every 45 seconds for independently playable files, targeting
64 kbit/s: approximately 360,000 bytes, or 480,000 base64 characters per chunk
(container overhead and actual encoder bitrate vary). Safari can fall back to
`audio/mp4`; the server accepts both MP4 and WebM. SDK uploads and playback use
base64 JSON across `tauriFetch`, below the default 10 MiB decoded asset cap and
the server's `ceil(cap * 4 / 3) + 64 KiB` request cap at this target size. The
desktop transport test checks two such chunks byte-for-byte with mocked native
IPC; it does not prove native capture or audible playback.

Manual release-app check (requires an interactive macOS microphone):

1. Run `pnpm build:desktop`, quit any other OpenBook instance, then launch
`packages/app/src-tauri/target/release/bundle/macos/OpenBook.app/Contents/MacOS/OpenBook`.
Use a release bundle: `tauri dev` does not launch the managed server sidecar.
2. Insert a meeting block on a page, press Record, and allow the macOS prompt.
Record for more than 45 seconds, then Stop. Confirm two uploaded audio chunks
and play each; reload the page and play again to verify sidecar persistence.
3. Restart the same app and record again; the OS should retain consent.
4. Disable OpenBook under System Settings → Privacy & Security → Microphone,
relaunch, and press Record. Confirm the microphone error appears, the block
returns to idle, and no capture/upload starts. Re-enable permission to recover.

Distribution must rebuild, sign, and notarize the app with the new entitlement
and purpose string through the existing release process. No signing identity,
hardened-runtime setting, or notarization configuration is changed here. Permission
continuity across locally signed and distributed builds needs a check using the
owner's normal signing identity; signing changes remain owner-gated.

## Sidecar supervision verification

The bundled sidecar is supervised only in a release build (`tauri dev` uses the
Expand Down Expand Up @@ -58,3 +101,25 @@ a fresh bounded run.
Crash-loop exhaustion, the 1/2/4/8/16-second bound, healthy reset, deliberate
stop suppression, and repair reset are deterministic unit tests in
`src-tauri/src/sidecar_supervision.rs` (`cargo test sidecar_supervision`).

## Meeting microphone manual checks

1. Launch the desktop app, open a saved page, type `/meeting`, and select
**Meeting**. Click **Record** and accept the OS microphone permission prompt.
On macOS, check **System Settings → Privacy & Security → Microphone** if the
prompt was previously denied; reopen OpenBook after changing permission.
2. Speak, pause, and confirm a playable audio chunk appears. Resume, speak again,
then stop. Confirm both chunks play and the elapsed time excludes the pause.
3. With a supported transcription backend configured, confirm transcript lines
appear as chunks complete. If no backend is available, confirm audio and notes
remain usable and transcription offers retry. See
[local transcription setup](../../docs/local-transcription.md) for runtime
installation and the model download in **Settings → AI**.
4. Generate a summary, cancel a regeneration, and confirm the previous summary
survives. Type a manual note. Export audio and Markdown; check the ZIP's
individual recordings and the Markdown transcript, summary, and note.
5. Revoke microphone access and reopen the app. Attempt recording: confirm a
permission error appears, Stop stays disabled, and manual notes still work.

These checks require the actual desktop webview and OS permission system. The
web e2e's oscillator stream does not validate desktop entitlements or hardware.
8 changes: 8 additions & 0 deletions packages/app/src-tauri/Info.plist
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>NSMicrophoneUsageDescription</key>
<string>OpenBook records meeting audio only when you press Record.</string>
</dict>
</plist>
2 changes: 2 additions & 0 deletions packages/app/src-tauri/entitlements.plist
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
<key>com.apple.security.device.audio-input</key>
<true/>
<key>com.apple.security.cs.allow-jit</key>
<true/>
<key>com.apple.security.cs.allow-unsigned-executable-memory</key>
Expand Down
7 changes: 7 additions & 0 deletions packages/app/src-tauri/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,13 @@
//! dev` the host is unmanaged and the webview talks to the external `pnpm dev`
//! server over loopback instead. Preferences (publish, token, book folder)
//! persist in `host-config.json` under the app-data dir.
//!
//! Microphone permissions: the locked Wry 0.55.1 already implements
//! `WKUIDelegate::requestMediaCapturePermissionForOrigin` and grants the webview
//! layer's request. macOS still asks for and persists the user's OS permission;
//! Info.plist supplies its purpose string and entitlements.plist enables audio
//! input under the hardened runtime. Keep Wry's delegate (including its other
//! callbacks); replacing it would duplicate upstream behavior. See README.md.

mod ipc;
mod sidecar_supervision;
Expand Down
Loading
Loading