From c37c1c40162a676960789ab49578cd29d870565c Mon Sep 17 00:00:00 2001 From: Eliot Lim Date: Sun, 4 Oct 2026 16:04:17 +0800 Subject: [PATCH 1/2] test(web): cover meeting epic and document workflows (MEET-9) --- ARCHITECTURE.md | 32 ++++++ docs/local-transcription.md | 35 +++++++ docs/meeting-notes.md | 64 ++++++++++++ packages/app/README.md | 22 ++++ packages/web/e2e/fixtures.ts | 20 ++++ packages/web/e2e/meeting-epic.spec.ts | 142 ++++++++++++++++++++++++++ 6 files changed, 315 insertions(+) create mode 100644 docs/local-transcription.md create mode 100644 docs/meeting-notes.md create mode 100644 packages/web/e2e/meeting-epic.spec.ts diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index db88b07b..862b926a 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -309,6 +309,38 @@ The FORM-7 MCP surface in `packages/mcp/src/server.ts` provides `list_forms`, through the resolved per-page agent-edits policy (suggest by default), and key regeneration is intentionally author-UI-only. +### Meeting blocks + +Meetings are native `type:'meeting'` container blocks registered by +`packages/ui/src/blockeditor/MeetingBlockView.tsx`. The +[representation contract](docs/meeting-block.md) defines audio asset references, +timestamped transcript segments, summary, title, and status props. Manual notes +remain ordinary CRDT child blocks; generated transcript and summary are prop +snapshots. The user workflow is in [meeting notes](docs/meeting-notes.md). + +`MeetingRecorder` restarts MediaRecorder every 45 seconds and on pause/resume. +Each uploaded chunk has its own container header, unlike recorder timeslices, +so playback, retries, transcription, and export work on standalone files. +Uploads use page-associated assets; `POST /api/ai/transcribe` receives the asset +and page IDs, enforces access, and records usage. Completed chunks append +transcript segments progressively. Summary generation is an explicit streamed +AI request; cancellation preserves the previous durable summary. + +Transcription resolves separately from chat: explicit off rejects; an explicit +OpenAI-compatible transcription provider opts into that endpoint; otherwise the +local resolver runs, followed by the deterministic mock fallback only when chat +provider is mock. An unavailable local engine returns a configuration error, +never implicit cloud fallback. Local engine wiring and the transcription +settings panel are pending in this branch; see +[local transcription setup](docs/local-transcription.md). + +Audio export downloads a single original file or an ordered timestamped ZIP of +chunks, without remuxing or deleting library assets. Markdown and HTML exports +include transcript, summary, and child notes with audio references; audio bytes +are exported separately. The browser epic proof is +`packages/web/e2e/meeting-epic.spec.ts`, using a WebAudio microphone substitute +with real recording, asset, transcription, generation, and download paths. + ### Optional local AI (`packages/server/src/ai/`) An opt-in, local-only model subsystem (Settings → AI). Pluggable engines diff --git a/docs/local-transcription.md b/docs/local-transcription.md new file mode 100644 index 00000000..070e4b58 --- /dev/null +++ b/docs/local-transcription.md @@ -0,0 +1,35 @@ +# Local transcription setup + +The transcription API resolves a local engine by default and never silently +sends audio to a cloud provider. **This branch exposes the resolver hook but does +not yet wire a local Whisper engine into the server.** Installing Whisper alone +will not enable the meeting block's automatic transcription. The separate +transcription controls in Settings → AI are also pending integration. + +You can prepare and validate Whisper independently using the upstream +[whisper.cpp quick start](https://github.com/ggml-org/whisper.cpp#quick-start). +With Git, CMake, and a C/C++ build toolchain installed: + +```sh +git clone https://github.com/ggml-org/whisper.cpp.git +cd whisper.cpp +sh ./models/download-ggml-model.sh base.en +cmake -B build +cmake --build build -j --config Release +./build/bin/whisper-cli -m models/ggml-base.en.bin -f samples/jfk.wav +``` + +`base.en` is English-only; choose `base` for multilingual audio. Model download +requires a network connection; inference with a downloaded model can run locally. +For an exported meeting chunk, install FFmpeg and convert it to a 16-bit WAV +before invoking the CLI: + +```sh +ffmpeg -i meeting.webm -ar 16000 -ac 1 -c:a pcm_s16le meeting.wav +./build/bin/whisper-cli -m models/ggml-base.en.bin -f meeting.wav +``` + +This standalone CLI check does not write a transcript back into OpenBook. Until +the engine integration lands, keep recorded audio in the library and retry +transcription after configuring a supported backend. See [meeting notes](meeting-notes.md) +for the recording, explicit cloud opt-in, and export behavior. diff --git a/docs/meeting-notes.md b/docs/meeting-notes.md new file mode 100644 index 00000000..86044833 --- /dev/null +++ b/docs/meeting-notes.md @@ -0,0 +1,64 @@ +# Meeting notes + +On a saved page, type `/meeting` and choose **Meeting**. Keep the page connected +to its OpenBook server so recordings can be saved to the library. + +## Record and transcribe + +Click **Record** and allow microphone access. OpenBook saves roughly 45-second +chunks and transcribes each completed upload; transcript lines appear as chunks +finish, without waiting for the entire meeting. Each chunk is a standalone audio +file, so it can be played or exported independently. + +**Pause** closes the current chunk and silences capture. **Resume** starts a new +chunk; paused time is excluded from the recording timeline. **Stop** finishes +capture and lets pending uploads and transcription settle. Keep the page open +until processing completes. Failed uploads retain an in-session audio copy with +save/retry controls; transcription failures preserve uploaded audio and offer +**Retry transcription**. Leaving the page can lose audio that has not uploaded. + +Transcription defaults to **local** resolution, independently of the chat model. +See [local transcription](local-transcription.md) for Whisper installation, +model download, and the current integration limitation. If local transcription +is unavailable, recording and manual notes still work; OpenBook does not silently +fall back to a cloud service. Cloud transcription requires explicit opt-in; the +intended control is **Settings → AI**. In this branch the server accepts a +separate transcription configuration, but its settings panel and local engine +wiring have not landed yet. Selecting a cloud chat model alone does not opt audio +into cloud transcription. + +## Summaries and notes + +Configure a generation provider in **Settings → AI**, then click **Generate +summary** once transcript text exists. Summaries never start automatically. +Text appears while generation streams. **Cancel** discards the unfinished +replacement and retains any previous summary. **Regenerate summary** requests a +new version. Long transcripts are clipped to the first 4,000 characters for the +summary prompt, with an instruction to disclose that it covers an excerpt. + +Type directly under **Notes**, including while recording. Notes are ordinary +editable blocks inside the meeting, independent of generated transcript and +summary text. + +## Export and privacy + +Click **Export audio** to download one audio file for a single saved chunk, or a +ZIP containing ordered, timestamped files for multiple chunks. Files retain their +original formats and bytes; the ZIP does not join or re-encode the recording. +Export leaves the original audio in your library. Audio is stored with the page's +assets and follows library/page access rules; it is not uploaded to a third-party +transcription service unless you explicitly configure one. With a remote library, +capture uploads to that library's server, so “local transcription” refers to +processing on the server rather than necessarily on the microphone's device. + +Use **Page actions → Export → Markdown (.md)** for transcript timestamps, +summary, manual notes, and audio references. Markdown is not an audio backup; +use **Export audio** for the actual recordings. + +## Desktop + +Allow microphone access for OpenBook in your OS privacy settings and reopen the +app if permission changes require it. Browser permission instructions may still +appear in the shared UI. Follow the [desktop microphone manual checks](../packages/app/README.md#meeting-microphone-manual-checks) +to verify the actual webview and OS permission behavior; Chromium e2e coverage +uses synthetic audio and cannot establish desktop microphone support. diff --git a/packages/app/README.md b/packages/app/README.md index f96da3dd..d2fef9d4 100644 --- a/packages/app/README.md +++ b/packages/app/README.md @@ -58,3 +58,25 @@ a fresh bounded run. Crash-loop exhaustion, the 1/2/4/8/16-second bound, healthy reset, deliberate stop suppression, and repair reset are deterministic unit tests in `src-tauri/src/sidecar_supervision.rs` (`cargo test sidecar_supervision`). + +## Meeting microphone manual checks + +1. Launch the desktop app, open a saved page, type `/meeting`, and select + **Meeting**. Click **Record** and accept the OS microphone permission prompt. + On macOS, check **System Settings → Privacy & Security → Microphone** if the + prompt was previously denied; reopen OpenBook after changing permission. +2. Speak, pause, and confirm a playable audio chunk appears. Resume, speak again, + then stop. Confirm both chunks play and the elapsed time excludes the pause. +3. With a supported transcription backend configured, confirm transcript lines + appear as chunks complete. If no backend is available, confirm audio and notes + remain usable and transcription offers retry. See + [local transcription setup](../../docs/local-transcription.md) for the current + local-engine integration limitation. +4. Generate a summary, cancel a regeneration, and confirm the previous summary + survives. Type a manual note. Export audio and Markdown; check the ZIP's + individual recordings and the Markdown transcript, summary, and note. +5. Revoke microphone access and reopen the app. Attempt recording: confirm a + permission error appears, Stop stays disabled, and manual notes still work. + +These checks require the actual desktop webview and OS permission system. The +web e2e's oscillator stream does not validate desktop entitlements or hardware. diff --git a/packages/web/e2e/fixtures.ts b/packages/web/e2e/fixtures.ts index 866c1897..b6827e07 100644 --- a/packages/web/e2e/fixtures.ts +++ b/packages/web/e2e/fixtures.ts @@ -74,6 +74,9 @@ type WorkerFixtures = { }; type TestFixtures = { + /** Opt into deterministic AI through the real server routes. */ + mockAi: boolean; + _mockAiConfig: void; /** * Opt a spec into structural per-test workspace isolation (OB-223). Set once * per file with `test.use({freshWorkspace: true})`: before EACH test the @@ -134,6 +137,23 @@ async function ensureAnyPage(serverUrl: string): Promise { export const test = base.extend({ freshWorkspace: [false, {option: true}], + mockAi: [false, {option: true}], + _mockAiConfig: [ + async ({mockAi, ownerRequest}, use) => { + if (!mockAi) { await use(); return; } + // The owner admin API writes settings key 'ai'. MockEngine keeps real + // upload/transcribe and streamed generation routes deterministic without + // a microphone device, model download, or external paid service. + const configured = await ownerRequest.put('/api/ai/config', {data: {provider: 'mock'}}); + expect(configured.ok()).toBeTruthy(); + try { await use(); } + finally { + const reset = await ownerRequest.put('/api/ai/config', {data: {provider: 'off'}}); + expect(reset.ok()).toBeTruthy(); + } + }, + {auto: true}, + ], ownerGatedRequests: [false, {option: true}], // UI specs that exercise Settings mutations opt into the desktop host's diff --git a/packages/web/e2e/meeting-epic.spec.ts b/packages/web/e2e/meeting-epic.spec.ts new file mode 100644 index 00000000..4da66755 --- /dev/null +++ b/packages/web/e2e/meeting-epic.spec.ts @@ -0,0 +1,142 @@ +import {readFile} from 'node:fs/promises'; +import type {Download, Page} from '@playwright/test'; +import {unzipSync} from 'fflate'; +import {expect, test} from './fixtures'; + +test.use({freshWorkspace: true, mockAi: true}); + +async function insertMeeting(page: Page): Promise { + const text = page.locator('.obe-text').first(); + await expect(text).toBeVisible(); + await text.click(); + await page.keyboard.press('End'); + await page.keyboard.press('Enter'); + await page.keyboard.type('/meeting'); + await page.locator('.obe-slash-item', { + has: page.locator('.obe-slash-label', {hasText: /^Meeting$/}), + }).click(); + await expect(page.getByRole('region', {name: 'Meeting', exact: true})).toBeVisible(); +} + +async function downloadBytes(download: Download) { + expect(await download.failure()).toBeNull(); + const bytes = await readFile(await download.path()); + expect(bytes.length).toBeGreaterThan(0); + return bytes; +} + +test.beforeEach(async ({page, request, dataServer}) => { + const response = await request.post(`${dataServer}/api/pages`, {data: { + name: 'Meeting epic', + data: {editor: 'blocks', blockdoc: {blocks: [ + {id: 'intro', type: 'paragraph', text: [{t: 'Meeting agenda'}]}, + ]}, editorjs: {blocks: []}, values: [], names: []}, + }}); + expect(response.ok()).toBeTruthy(); + const {id} = await response.json() as {id: string}; + await page.addInitScript(() => { + // Chromium's fake-media launch flags do not expose audioinput on every + // build. Keep the real MediaRecorder and supply a WebAudio audio track. + Object.defineProperty(navigator.mediaDevices, 'getUserMedia', {configurable: true, writable: true, value: async () => { + const audio = new AudioContext(); + const oscillator = audio.createOscillator(); + const destination = audio.createMediaStreamDestination(); + oscillator.connect(destination); + oscillator.start(); + await audio.resume(); + destination.stream.getTracks()[0].addEventListener('ended', () => { void audio.close(); }); + return destination.stream; + }}); + }); + await page.goto(`/?page=${id}`); + await insertMeeting(page); +}); + +test('record, progressively transcribe, summarize, note and export a meeting', {tag: ['@meeting', '@p1']}, async ({page, request, dataServer}) => { + test.setTimeout(90_000); + const meeting = page.getByRole('region', {name: 'Meeting', exact: true}); + const transcript = meeting.locator('.obe-meeting-transcript p'); + const transcriptions: string[] = []; + page.on('response', (response) => { + if (response.url().endsWith('/api/ai/transcribe') && response.ok()) transcriptions.push(response.url()); + }); + await meeting.getByRole('button', {name: 'Record', exact: true}).click(); + await expect(meeting.locator('header')).toContainText('Recording'); + // Pause closes a genuine standalone chunk without sleeping for the 45s + // rollover. Resume proves transcript growth while the session is still live. + await meeting.getByRole('button', {name: 'Pause', exact: true}).click(); + await expect(transcript).toHaveCount(1); + await expect(transcript.first()).toContainText(/Mock transcription \([1-9]\d* bytes\)/); + await expect(meeting.locator('header')).toContainText('Paused'); + const firstText = await transcript.first().innerText(); + const singleDownload = page.waitForEvent('download'); + await meeting.getByRole('button', {name: 'Export audio', exact: true}).click(); + const single = await singleDownload; + expect(single.suggestedFilename()).toBe('Meeting-00h00m00s.webm'); + const singleBytes = await downloadBytes(single); + + await meeting.getByRole('button', {name: 'Resume', exact: true}).click(); + await expect(meeting.locator('header')).toContainText('Recording'); + await meeting.getByRole('button', {name: 'Pause', exact: true}).click(); + await expect(transcript).toHaveCount(2); + await expect(transcript.first()).toHaveText(firstText); + await expect(transcript.nth(1)).toContainText(/Mock transcription \([1-9]\d* bytes\)/); + expect(transcriptions).toHaveLength(2); + await expect(meeting.locator('audio')).toHaveCount(2); + await meeting.getByRole('button', {name: 'Stop', exact: true}).click(); + await expect(meeting.locator('header')).toContainText('Done'); + + await expect(meeting.locator('.obe-meeting-summary')).toHaveCount(0); + const generation = page.waitForResponse((response) => response.url().includes('/api/ai/generate') && response.ok()); + await meeting.getByRole('button', {name: 'Generate summary', exact: true}).click(); + expect((await generation).headers()['content-type']).toContain('text/event-stream'); + await expect(meeting.getByRole('button', {name: 'Regenerate summary'})).toBeVisible(); + const summary = await meeting.locator('.obe-meeting-summary').innerText(); + expect(summary).toContain('Mock response to:'); + const note = 'Ellis will send the meeting follow-up.'; + await meeting.locator('.obe-meeting-notes .obe-text').fill(note); + const pageId = new URL(page.url()).searchParams.get('page'); + await expect.poll(async () => { + const stored = await request.get(`${dataServer}/api/pages/${pageId}`); + expect(stored.ok()).toBeTruthy(); + return JSON.stringify((await stored.json() as {data: {blockdoc: unknown}}).data.blockdoc); + }).toContain(note); + + const zipDownload = page.waitForEvent('download'); + await meeting.getByRole('button', {name: 'Export audio', exact: true}).click(); + const zip = await zipDownload; + expect(zip.suggestedFilename()).toMatch(/-audio\.zip$/); + const files = unzipSync(await downloadBytes(zip)); + expect(Object.keys(files)).toHaveLength(2); + expect(Object.keys(files)).toEqual([expect.stringMatching(/^001-\d{2}h\d{2}m\d{2}s\.webm$/), expect.stringMatching(/^002-\d{2}h\d{2}m\d{2}s\.webm$/)]); + expect(files[Object.keys(files)[0]]).toEqual(new Uint8Array(singleBytes)); + for (const bytes of Object.values(files)) expect(bytes.length).toBeGreaterThan(0); + + await page.getByRole('button', {name: 'Page actions'}).click(); + await page.getByRole('menuitem', {name: 'Export', exact: true}).click(); + const markdownDownload = page.waitForEvent('download'); + await page.getByRole('menuitem', {name: 'Markdown (.md)', exact: true}).click(); + const markdown = (await downloadBytes(await markdownDownload)).toString('utf8'); + for (const segment of await transcript.allTextContents()) { + expect(markdown).toContain(segment.replace(/^\S+\s+/, '').trim()); + } + expect(markdown).toContain(summary); + expect(markdown).toContain(note); +}); + +test('microphone denial shows recovery copy and leaves the editor usable', {tag: ['@meeting', '@p1']}, async ({page}) => { + const errors: string[] = []; + page.on('pageerror', (error) => errors.push(error.message)); + await page.evaluate(() => { + navigator.mediaDevices.getUserMedia = async () => { throw new DOMException('Permission denied', 'NotAllowedError'); }; + }); + const meeting = page.getByRole('region', {name: 'Meeting', exact: true}); + await meeting.getByRole('button', {name: 'Record', exact: true}).click(); + await expect(meeting.getByRole('alert')).toHaveText('Microphone access failed. Allow microphone access in browser settings, then try again.'); + await expect(meeting.getByRole('button', {name: 'Record', exact: true})).toBeEnabled(); + await expect(meeting.getByRole('button', {name: 'Stop', exact: true})).toBeDisabled(); + await meeting.locator('.obe-meeting-notes .obe-text').fill('Notes still work without microphone access.'); + await expect(meeting.locator('.obe-meeting-notes')).toContainText('Notes still work without microphone access.'); + await expect(meeting.locator('audio')).toHaveCount(0); + expect(errors).toEqual([]); +}); From 938161c9326ee6543b8ef47cff9e9d8d6b0505b6 Mon Sep 17 00:00:00 2001 From: Eliot Lim Date: Sun, 4 Oct 2026 16:14:07 +0800 Subject: [PATCH 2/2] docs: point local-transcription caveats at MEET-3 (MEET-9) --- ARCHITECTURE.md | 4 +--- docs/meeting-notes.md | 4 +--- 2 files changed, 2 insertions(+), 6 deletions(-) diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 862b926a..01922f93 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -330,9 +330,7 @@ Transcription resolves separately from chat: explicit off rejects; an explicit OpenAI-compatible transcription provider opts into that endpoint; otherwise the local resolver runs, followed by the deterministic mock fallback only when chat provider is mock. An unavailable local engine returns a configuration error, -never implicit cloud fallback. Local engine wiring and the transcription -settings panel are pending in this branch; see -[local transcription setup](docs/local-transcription.md). +never implicit cloud fallback. Local engine wiring and the transcription settings panel land with MEET-3 (PR #370); see [local transcription setup](docs/local-transcription.md). Audio export downloads a single original file or an ordered timestamped ZIP of chunks, without remuxing or deleting library assets. Markdown and HTML exports diff --git a/docs/meeting-notes.md b/docs/meeting-notes.md index 86044833..548d4a1f 100644 --- a/docs/meeting-notes.md +++ b/docs/meeting-notes.md @@ -22,9 +22,7 @@ See [local transcription](local-transcription.md) for Whisper installation, model download, and the current integration limitation. If local transcription is unavailable, recording and manual notes still work; OpenBook does not silently fall back to a cloud service. Cloud transcription requires explicit opt-in; the -intended control is **Settings → AI**. In this branch the server accepts a -separate transcription configuration, but its settings panel and local engine -wiring have not landed yet. Selecting a cloud chat model alone does not opt audio +intended control is **Settings → AI**. The local Whisper engine and its Settings → AI transcription controls ship with MEET-3 (PR #370); until that lands, the server accepts the transcription configuration but has no local engine wired. Selecting a cloud chat model alone does not opt audio into cloud transcription. ## Summaries and notes