diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index db88b07b..453fd80a 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -309,6 +309,38 @@ The FORM-7 MCP surface in `packages/mcp/src/server.ts` provides `list_forms`, through the resolved per-page agent-edits policy (suggest by default), and key regeneration is intentionally author-UI-only. +### Meeting blocks + +Meetings are native `type:'meeting'` container blocks registered by +`packages/ui/src/blockeditor/MeetingBlockView.tsx`. The +[representation contract](docs/meeting-block.md) defines audio asset references, +timestamped transcript segments, summary, title, and status props. Manual notes +remain ordinary CRDT child blocks; generated transcript and summary are prop +snapshots. The user workflow is in [meeting notes](docs/meeting-notes.md). + +`MeetingRecorder` restarts MediaRecorder every 45 seconds and on pause/resume. +Each uploaded chunk has its own container header, unlike recorder timeslices, +so playback, retries, transcription, and export work on standalone files. +Uploads use page-associated assets; `POST /api/ai/transcribe` receives the asset +and page IDs, enforces access, and records usage. Completed chunks append +transcript segments progressively. Summary generation is an explicit streamed +AI request; cancellation preserves the previous durable summary. + +Transcription resolves separately from chat: explicit off rejects; an explicit +OpenAI-compatible transcription provider opts into that endpoint; otherwise the +local resolver runs, followed by the deterministic mock fallback only when chat +provider is mock. An unavailable local engine returns a configuration error, +never implicit cloud fallback. Local Whisper transcription ships by default; **Settings → AI** provides the +model download. See [local transcription setup](docs/local-transcription.md) for +Whisper and FFmpeg installation and runtime requirements. + +Audio export downloads a single original file or an ordered timestamped ZIP of +chunks, without remuxing or deleting library assets. Markdown and HTML exports +include transcript, summary, and child notes with audio references; audio bytes +are exported separately. The browser epic proof is +`packages/web/e2e/meeting-epic.spec.ts`, using a WebAudio microphone substitute +with real recording, asset, transcription, generation, and download paths. + ### Optional local AI (`packages/server/src/ai/`) An opt-in, local-only model subsystem (Settings → AI). Pluggable engines diff --git a/docs/meeting-notes.md b/docs/meeting-notes.md new file mode 100644 index 00000000..ed33a2ec --- /dev/null +++ b/docs/meeting-notes.md @@ -0,0 +1,63 @@ +# Meeting notes + +On a saved page, type `/meeting` and choose **Meeting**. Keep the page connected +to its OpenBook server so recordings can be saved to the library. + +## Record and transcribe + +Click **Record** and allow microphone access. OpenBook saves roughly 45-second +chunks and transcribes each completed upload; transcript lines appear as chunks +finish, without waiting for the entire meeting. Each chunk is a standalone audio +file, so it can be played or exported independently. + +**Pause** closes the current chunk and silences capture. **Resume** starts a new +chunk; paused time is excluded from the recording timeline. **Stop** finishes +capture and lets pending uploads and transcription settle. Keep the page open +until processing completes. Failed uploads retain an in-session audio copy with +save/retry controls; transcription failures preserve uploaded audio and offer +**Retry transcription**. Leaving the page can lose audio that has not uploaded. + +Local Whisper transcription ships by default, independently of the chat model. +In **Settings → AI**, select **Download Whisper base** to download the model. +See [local transcription](local-transcription.md) for the required Whisper and +FFmpeg installation, model setup, and runtime checks. If local transcription +is unavailable, recording and manual notes still work; OpenBook does not silently +fall back to a cloud service. Cloud transcription requires explicit opt-in in +**Settings → AI**. Selecting a cloud chat model alone does not opt audio into +cloud transcription. + +## Summaries and notes + +Configure a generation provider in **Settings → AI**, then click **Generate +summary** once transcript text exists. Summaries never start automatically. +Text appears while generation streams. **Cancel** discards the unfinished +replacement and retains any previous summary. **Regenerate summary** requests a +new version. Long transcripts are clipped to the first 4,000 characters for the +summary prompt, with an instruction to disclose that it covers an excerpt. + +Type directly under **Notes**, including while recording. Notes are ordinary +editable blocks inside the meeting, independent of generated transcript and +summary text. + +## Export and privacy + +Click **Export audio** to download one audio file for a single saved chunk, or a +ZIP containing ordered, timestamped files for multiple chunks. Files retain their +original formats and bytes; the ZIP does not join or re-encode the recording. +Export leaves the original audio in your library. Audio is stored with the page's +assets and follows library/page access rules; it is not uploaded to a third-party +transcription service unless you explicitly configure one. With a remote library, +capture uploads to that library's server, so “local transcription” refers to +processing on the server rather than necessarily on the microphone's device. + +Use **Page actions → Export → Markdown (.md)** for transcript timestamps, +summary, manual notes, and audio references. Markdown is not an audio backup; +use **Export audio** for the actual recordings. + +## Desktop + +Allow microphone access for OpenBook in your OS privacy settings and reopen the +app if permission changes require it. Browser permission instructions may still +appear in the shared UI. Follow the [desktop microphone manual checks](../packages/app/README.md#meeting-microphone-manual-checks) +to verify the actual webview and OS permission behavior; Chromium e2e coverage +uses synthetic audio and cannot establish desktop microphone support. diff --git a/packages/app/README.md b/packages/app/README.md index f135baf6..8d10c9ef 100644 --- a/packages/app/README.md +++ b/packages/app/README.md @@ -101,3 +101,25 @@ a fresh bounded run. Crash-loop exhaustion, the 1/2/4/8/16-second bound, healthy reset, deliberate stop suppression, and repair reset are deterministic unit tests in `src-tauri/src/sidecar_supervision.rs` (`cargo test sidecar_supervision`). + +## Meeting microphone manual checks + +1. Launch the desktop app, open a saved page, type `/meeting`, and select + **Meeting**. Click **Record** and accept the OS microphone permission prompt. + On macOS, check **System Settings → Privacy & Security → Microphone** if the + prompt was previously denied; reopen OpenBook after changing permission. +2. Speak, pause, and confirm a playable audio chunk appears. Resume, speak again, + then stop. Confirm both chunks play and the elapsed time excludes the pause. +3. With a supported transcription backend configured, confirm transcript lines + appear as chunks complete. If no backend is available, confirm audio and notes + remain usable and transcription offers retry. See + [local transcription setup](../../docs/local-transcription.md) for runtime + installation and the model download in **Settings → AI**. +4. Generate a summary, cancel a regeneration, and confirm the previous summary + survives. Type a manual note. Export audio and Markdown; check the ZIP's + individual recordings and the Markdown transcript, summary, and note. +5. Revoke microphone access and reopen the app. Attempt recording: confirm a + permission error appears, Stop stays disabled, and manual notes still work. + +These checks require the actual desktop webview and OS permission system. The +web e2e's oscillator stream does not validate desktop entitlements or hardware. diff --git a/packages/web/e2e/fixtures.ts b/packages/web/e2e/fixtures.ts index 866c1897..b6827e07 100644 --- a/packages/web/e2e/fixtures.ts +++ b/packages/web/e2e/fixtures.ts @@ -74,6 +74,9 @@ type WorkerFixtures = { }; type TestFixtures = { + /** Opt into deterministic AI through the real server routes. */ + mockAi: boolean; + _mockAiConfig: void; /** * Opt a spec into structural per-test workspace isolation (OB-223). Set once * per file with `test.use({freshWorkspace: true})`: before EACH test the @@ -134,6 +137,23 @@ async function ensureAnyPage(serverUrl: string): Promise { export const test = base.extend({ freshWorkspace: [false, {option: true}], + mockAi: [false, {option: true}], + _mockAiConfig: [ + async ({mockAi, ownerRequest}, use) => { + if (!mockAi) { await use(); return; } + // The owner admin API writes settings key 'ai'. MockEngine keeps real + // upload/transcribe and streamed generation routes deterministic without + // a microphone device, model download, or external paid service. + const configured = await ownerRequest.put('/api/ai/config', {data: {provider: 'mock'}}); + expect(configured.ok()).toBeTruthy(); + try { await use(); } + finally { + const reset = await ownerRequest.put('/api/ai/config', {data: {provider: 'off'}}); + expect(reset.ok()).toBeTruthy(); + } + }, + {auto: true}, + ], ownerGatedRequests: [false, {option: true}], // UI specs that exercise Settings mutations opt into the desktop host's diff --git a/packages/web/e2e/meeting-epic.spec.ts b/packages/web/e2e/meeting-epic.spec.ts new file mode 100644 index 00000000..4da66755 --- /dev/null +++ b/packages/web/e2e/meeting-epic.spec.ts @@ -0,0 +1,142 @@ +import {readFile} from 'node:fs/promises'; +import type {Download, Page} from '@playwright/test'; +import {unzipSync} from 'fflate'; +import {expect, test} from './fixtures'; + +test.use({freshWorkspace: true, mockAi: true}); + +async function insertMeeting(page: Page): Promise { + const text = page.locator('.obe-text').first(); + await expect(text).toBeVisible(); + await text.click(); + await page.keyboard.press('End'); + await page.keyboard.press('Enter'); + await page.keyboard.type('/meeting'); + await page.locator('.obe-slash-item', { + has: page.locator('.obe-slash-label', {hasText: /^Meeting$/}), + }).click(); + await expect(page.getByRole('region', {name: 'Meeting', exact: true})).toBeVisible(); +} + +async function downloadBytes(download: Download) { + expect(await download.failure()).toBeNull(); + const bytes = await readFile(await download.path()); + expect(bytes.length).toBeGreaterThan(0); + return bytes; +} + +test.beforeEach(async ({page, request, dataServer}) => { + const response = await request.post(`${dataServer}/api/pages`, {data: { + name: 'Meeting epic', + data: {editor: 'blocks', blockdoc: {blocks: [ + {id: 'intro', type: 'paragraph', text: [{t: 'Meeting agenda'}]}, + ]}, editorjs: {blocks: []}, values: [], names: []}, + }}); + expect(response.ok()).toBeTruthy(); + const {id} = await response.json() as {id: string}; + await page.addInitScript(() => { + // Chromium's fake-media launch flags do not expose audioinput on every + // build. Keep the real MediaRecorder and supply a WebAudio audio track. + Object.defineProperty(navigator.mediaDevices, 'getUserMedia', {configurable: true, writable: true, value: async () => { + const audio = new AudioContext(); + const oscillator = audio.createOscillator(); + const destination = audio.createMediaStreamDestination(); + oscillator.connect(destination); + oscillator.start(); + await audio.resume(); + destination.stream.getTracks()[0].addEventListener('ended', () => { void audio.close(); }); + return destination.stream; + }}); + }); + await page.goto(`/?page=${id}`); + await insertMeeting(page); +}); + +test('record, progressively transcribe, summarize, note and export a meeting', {tag: ['@meeting', '@p1']}, async ({page, request, dataServer}) => { + test.setTimeout(90_000); + const meeting = page.getByRole('region', {name: 'Meeting', exact: true}); + const transcript = meeting.locator('.obe-meeting-transcript p'); + const transcriptions: string[] = []; + page.on('response', (response) => { + if (response.url().endsWith('/api/ai/transcribe') && response.ok()) transcriptions.push(response.url()); + }); + await meeting.getByRole('button', {name: 'Record', exact: true}).click(); + await expect(meeting.locator('header')).toContainText('Recording'); + // Pause closes a genuine standalone chunk without sleeping for the 45s + // rollover. Resume proves transcript growth while the session is still live. + await meeting.getByRole('button', {name: 'Pause', exact: true}).click(); + await expect(transcript).toHaveCount(1); + await expect(transcript.first()).toContainText(/Mock transcription \([1-9]\d* bytes\)/); + await expect(meeting.locator('header')).toContainText('Paused'); + const firstText = await transcript.first().innerText(); + const singleDownload = page.waitForEvent('download'); + await meeting.getByRole('button', {name: 'Export audio', exact: true}).click(); + const single = await singleDownload; + expect(single.suggestedFilename()).toBe('Meeting-00h00m00s.webm'); + const singleBytes = await downloadBytes(single); + + await meeting.getByRole('button', {name: 'Resume', exact: true}).click(); + await expect(meeting.locator('header')).toContainText('Recording'); + await meeting.getByRole('button', {name: 'Pause', exact: true}).click(); + await expect(transcript).toHaveCount(2); + await expect(transcript.first()).toHaveText(firstText); + await expect(transcript.nth(1)).toContainText(/Mock transcription \([1-9]\d* bytes\)/); + expect(transcriptions).toHaveLength(2); + await expect(meeting.locator('audio')).toHaveCount(2); + await meeting.getByRole('button', {name: 'Stop', exact: true}).click(); + await expect(meeting.locator('header')).toContainText('Done'); + + await expect(meeting.locator('.obe-meeting-summary')).toHaveCount(0); + const generation = page.waitForResponse((response) => response.url().includes('/api/ai/generate') && response.ok()); + await meeting.getByRole('button', {name: 'Generate summary', exact: true}).click(); + expect((await generation).headers()['content-type']).toContain('text/event-stream'); + await expect(meeting.getByRole('button', {name: 'Regenerate summary'})).toBeVisible(); + const summary = await meeting.locator('.obe-meeting-summary').innerText(); + expect(summary).toContain('Mock response to:'); + const note = 'Ellis will send the meeting follow-up.'; + await meeting.locator('.obe-meeting-notes .obe-text').fill(note); + const pageId = new URL(page.url()).searchParams.get('page'); + await expect.poll(async () => { + const stored = await request.get(`${dataServer}/api/pages/${pageId}`); + expect(stored.ok()).toBeTruthy(); + return JSON.stringify((await stored.json() as {data: {blockdoc: unknown}}).data.blockdoc); + }).toContain(note); + + const zipDownload = page.waitForEvent('download'); + await meeting.getByRole('button', {name: 'Export audio', exact: true}).click(); + const zip = await zipDownload; + expect(zip.suggestedFilename()).toMatch(/-audio\.zip$/); + const files = unzipSync(await downloadBytes(zip)); + expect(Object.keys(files)).toHaveLength(2); + expect(Object.keys(files)).toEqual([expect.stringMatching(/^001-\d{2}h\d{2}m\d{2}s\.webm$/), expect.stringMatching(/^002-\d{2}h\d{2}m\d{2}s\.webm$/)]); + expect(files[Object.keys(files)[0]]).toEqual(new Uint8Array(singleBytes)); + for (const bytes of Object.values(files)) expect(bytes.length).toBeGreaterThan(0); + + await page.getByRole('button', {name: 'Page actions'}).click(); + await page.getByRole('menuitem', {name: 'Export', exact: true}).click(); + const markdownDownload = page.waitForEvent('download'); + await page.getByRole('menuitem', {name: 'Markdown (.md)', exact: true}).click(); + const markdown = (await downloadBytes(await markdownDownload)).toString('utf8'); + for (const segment of await transcript.allTextContents()) { + expect(markdown).toContain(segment.replace(/^\S+\s+/, '').trim()); + } + expect(markdown).toContain(summary); + expect(markdown).toContain(note); +}); + +test('microphone denial shows recovery copy and leaves the editor usable', {tag: ['@meeting', '@p1']}, async ({page}) => { + const errors: string[] = []; + page.on('pageerror', (error) => errors.push(error.message)); + await page.evaluate(() => { + navigator.mediaDevices.getUserMedia = async () => { throw new DOMException('Permission denied', 'NotAllowedError'); }; + }); + const meeting = page.getByRole('region', {name: 'Meeting', exact: true}); + await meeting.getByRole('button', {name: 'Record', exact: true}).click(); + await expect(meeting.getByRole('alert')).toHaveText('Microphone access failed. Allow microphone access in browser settings, then try again.'); + await expect(meeting.getByRole('button', {name: 'Record', exact: true})).toBeEnabled(); + await expect(meeting.getByRole('button', {name: 'Stop', exact: true})).toBeDisabled(); + await meeting.locator('.obe-meeting-notes .obe-text').fill('Notes still work without microphone access.'); + await expect(meeting.locator('.obe-meeting-notes')).toContainText('Notes still work without microphone access.'); + await expect(meeting.locator('audio')).toHaveCount(0); + expect(errors).toEqual([]); +});