Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -309,6 +309,38 @@ The FORM-7 MCP surface in `packages/mcp/src/server.ts` provides `list_forms`,
through the resolved per-page agent-edits policy (suggest by default), and key
regeneration is intentionally author-UI-only.

### Meeting blocks

Meetings are native `type:'meeting'` container blocks registered by
`packages/ui/src/blockeditor/MeetingBlockView.tsx`. The
[representation contract](docs/meeting-block.md) defines audio asset references,
timestamped transcript segments, summary, title, and status props. Manual notes
remain ordinary CRDT child blocks; generated transcript and summary are prop
snapshots. The user workflow is in [meeting notes](docs/meeting-notes.md).

`MeetingRecorder` restarts MediaRecorder every 45 seconds and on pause/resume.
Each uploaded chunk has its own container header, unlike recorder timeslices,
so playback, retries, transcription, and export work on standalone files.
Uploads use page-associated assets; `POST /api/ai/transcribe` receives the asset
and page IDs, enforces access, and records usage. Completed chunks append
transcript segments progressively. Summary generation is an explicit streamed
AI request; cancellation preserves the previous durable summary.

Transcription resolves separately from chat: explicit off rejects; an explicit
OpenAI-compatible transcription provider opts into that endpoint; otherwise the
local resolver runs, followed by the deterministic mock fallback only when chat
provider is mock. An unavailable local engine returns a configuration error,
never implicit cloud fallback. Local Whisper transcription ships by default; **Settings → AI** provides the
model download. See [local transcription setup](docs/local-transcription.md) for
Whisper and FFmpeg installation and runtime requirements.

Audio export downloads a single original file or an ordered timestamped ZIP of
chunks, without remuxing or deleting library assets. Markdown and HTML exports
include transcript, summary, and child notes with audio references; audio bytes
are exported separately. The browser epic proof is
`packages/web/e2e/meeting-epic.spec.ts`, using a WebAudio microphone substitute
with real recording, asset, transcription, generation, and download paths.

### Optional local AI (`packages/server/src/ai/`)

An opt-in, local-only model subsystem (Settings → AI). Pluggable engines
Expand Down
63 changes: 63 additions & 0 deletions docs/meeting-notes.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# Meeting notes

On a saved page, type `/meeting` and choose **Meeting**. Keep the page connected
to its OpenBook server so recordings can be saved to the library.

## Record and transcribe

Click **Record** and allow microphone access. OpenBook saves roughly 45-second
chunks and transcribes each completed upload; transcript lines appear as chunks
finish, without waiting for the entire meeting. Each chunk is a standalone audio
file, so it can be played or exported independently.

**Pause** closes the current chunk and silences capture. **Resume** starts a new
chunk; paused time is excluded from the recording timeline. **Stop** finishes
capture and lets pending uploads and transcription settle. Keep the page open
until processing completes. Failed uploads retain an in-session audio copy with
save/retry controls; transcription failures preserve uploaded audio and offer
**Retry transcription**. Leaving the page can lose audio that has not uploaded.

Local Whisper transcription ships by default, independently of the chat model.
In **Settings → AI**, select **Download Whisper base** to download the model.
See [local transcription](local-transcription.md) for the required Whisper and
FFmpeg installation, model setup, and runtime checks. If local transcription
is unavailable, recording and manual notes still work; OpenBook does not silently
fall back to a cloud service. Cloud transcription requires explicit opt-in in
**Settings → AI**. Selecting a cloud chat model alone does not opt audio into
cloud transcription.

## Summaries and notes

Configure a generation provider in **Settings → AI**, then click **Generate
summary** once transcript text exists. Summaries never start automatically.
Text appears while generation streams. **Cancel** discards the unfinished
replacement and retains any previous summary. **Regenerate summary** requests a
new version. Long transcripts are clipped to the first 4,000 characters for the
summary prompt, with an instruction to disclose that it covers an excerpt.

Type directly under **Notes**, including while recording. Notes are ordinary
editable blocks inside the meeting, independent of generated transcript and
summary text.

## Export and privacy

Click **Export audio** to download one audio file for a single saved chunk, or a
ZIP containing ordered, timestamped files for multiple chunks. Files retain their
original formats and bytes; the ZIP does not join or re-encode the recording.
Export leaves the original audio in your library. Audio is stored with the page's
assets and follows library/page access rules; it is not uploaded to a third-party
transcription service unless you explicitly configure one. With a remote library,
capture uploads to that library's server, so “local transcription” refers to
processing on the server rather than necessarily on the microphone's device.

Use **Page actions → Export → Markdown (.md)** for transcript timestamps,
summary, manual notes, and audio references. Markdown is not an audio backup;
use **Export audio** for the actual recordings.

## Desktop

Allow microphone access for OpenBook in your OS privacy settings and reopen the
app if permission changes require it. Browser permission instructions may still
appear in the shared UI. Follow the [desktop microphone manual checks](../packages/app/README.md#meeting-microphone-manual-checks)
to verify the actual webview and OS permission behavior; Chromium e2e coverage
uses synthetic audio and cannot establish desktop microphone support.
22 changes: 22 additions & 0 deletions packages/app/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -101,3 +101,25 @@ a fresh bounded run.
Crash-loop exhaustion, the 1/2/4/8/16-second bound, healthy reset, deliberate
stop suppression, and repair reset are deterministic unit tests in
`src-tauri/src/sidecar_supervision.rs` (`cargo test sidecar_supervision`).

## Meeting microphone manual checks

1. Launch the desktop app, open a saved page, type `/meeting`, and select
**Meeting**. Click **Record** and accept the OS microphone permission prompt.
On macOS, check **System Settings → Privacy & Security → Microphone** if the
prompt was previously denied; reopen OpenBook after changing permission.
2. Speak, pause, and confirm a playable audio chunk appears. Resume, speak again,
then stop. Confirm both chunks play and the elapsed time excludes the pause.
3. With a supported transcription backend configured, confirm transcript lines
appear as chunks complete. If no backend is available, confirm audio and notes
remain usable and transcription offers retry. See
[local transcription setup](../../docs/local-transcription.md) for runtime
installation and the model download in **Settings → AI**.
4. Generate a summary, cancel a regeneration, and confirm the previous summary
survives. Type a manual note. Export audio and Markdown; check the ZIP's
individual recordings and the Markdown transcript, summary, and note.
5. Revoke microphone access and reopen the app. Attempt recording: confirm a
permission error appears, Stop stays disabled, and manual notes still work.

These checks require the actual desktop webview and OS permission system. The
web e2e's oscillator stream does not validate desktop entitlements or hardware.
20 changes: 20 additions & 0 deletions packages/web/e2e/fixtures.ts
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,9 @@ type WorkerFixtures = {
};

type TestFixtures = {
/** Opt into deterministic AI through the real server routes. */
mockAi: boolean;
_mockAiConfig: void;
/**
* Opt a spec into structural per-test workspace isolation (OB-223). Set once
* per file with `test.use({freshWorkspace: true})`: before EACH test the
Expand Down Expand Up @@ -134,6 +137,23 @@ async function ensureAnyPage(serverUrl: string): Promise<void> {

export const test = base.extend<TestFixtures, WorkerFixtures>({
freshWorkspace: [false, {option: true}],
mockAi: [false, {option: true}],
_mockAiConfig: [
async ({mockAi, ownerRequest}, use) => {
if (!mockAi) { await use(); return; }
// The owner admin API writes settings key 'ai'. MockEngine keeps real
// upload/transcribe and streamed generation routes deterministic without
// a microphone device, model download, or external paid service.
const configured = await ownerRequest.put('/api/ai/config', {data: {provider: 'mock'}});
expect(configured.ok()).toBeTruthy();
try { await use(); }
finally {
const reset = await ownerRequest.put('/api/ai/config', {data: {provider: 'off'}});
expect(reset.ok()).toBeTruthy();
}
},
{auto: true},
],
ownerGatedRequests: [false, {option: true}],

// UI specs that exercise Settings mutations opt into the desktop host's
Expand Down
142 changes: 142 additions & 0 deletions packages/web/e2e/meeting-epic.spec.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,142 @@
import {readFile} from 'node:fs/promises';
import type {Download, Page} from '@playwright/test';
import {unzipSync} from 'fflate';
import {expect, test} from './fixtures';

test.use({freshWorkspace: true, mockAi: true});

async function insertMeeting(page: Page): Promise<void> {
const text = page.locator('.obe-text').first();
await expect(text).toBeVisible();
await text.click();
await page.keyboard.press('End');
await page.keyboard.press('Enter');
await page.keyboard.type('/meeting');
await page.locator('.obe-slash-item', {
has: page.locator('.obe-slash-label', {hasText: /^Meeting$/}),
}).click();
await expect(page.getByRole('region', {name: 'Meeting', exact: true})).toBeVisible();
}

async function downloadBytes(download: Download) {
expect(await download.failure()).toBeNull();
const bytes = await readFile(await download.path());
expect(bytes.length).toBeGreaterThan(0);
return bytes;
}

test.beforeEach(async ({page, request, dataServer}) => {
const response = await request.post(`${dataServer}/api/pages`, {data: {
name: 'Meeting epic',
data: {editor: 'blocks', blockdoc: {blocks: [
{id: 'intro', type: 'paragraph', text: [{t: 'Meeting agenda'}]},
]}, editorjs: {blocks: []}, values: [], names: []},
}});
expect(response.ok()).toBeTruthy();
const {id} = await response.json() as {id: string};
await page.addInitScript(() => {
// Chromium's fake-media launch flags do not expose audioinput on every
// build. Keep the real MediaRecorder and supply a WebAudio audio track.
Object.defineProperty(navigator.mediaDevices, 'getUserMedia', {configurable: true, writable: true, value: async () => {
const audio = new AudioContext();
const oscillator = audio.createOscillator();
const destination = audio.createMediaStreamDestination();
oscillator.connect(destination);
oscillator.start();
await audio.resume();
destination.stream.getTracks()[0].addEventListener('ended', () => { void audio.close(); });
return destination.stream;
}});
});
await page.goto(`/?page=${id}`);
await insertMeeting(page);
});

test('record, progressively transcribe, summarize, note and export a meeting', {tag: ['@meeting', '@p1']}, async ({page, request, dataServer}) => {
test.setTimeout(90_000);
const meeting = page.getByRole('region', {name: 'Meeting', exact: true});
const transcript = meeting.locator('.obe-meeting-transcript p');
const transcriptions: string[] = [];
page.on('response', (response) => {
if (response.url().endsWith('/api/ai/transcribe') && response.ok()) transcriptions.push(response.url());
});
await meeting.getByRole('button', {name: 'Record', exact: true}).click();
await expect(meeting.locator('header')).toContainText('Recording');
// Pause closes a genuine standalone chunk without sleeping for the 45s
// rollover. Resume proves transcript growth while the session is still live.
await meeting.getByRole('button', {name: 'Pause', exact: true}).click();
await expect(transcript).toHaveCount(1);
await expect(transcript.first()).toContainText(/Mock transcription \([1-9]\d* bytes\)/);
await expect(meeting.locator('header')).toContainText('Paused');
const firstText = await transcript.first().innerText();
const singleDownload = page.waitForEvent('download');
await meeting.getByRole('button', {name: 'Export audio', exact: true}).click();
const single = await singleDownload;
expect(single.suggestedFilename()).toBe('Meeting-00h00m00s.webm');
const singleBytes = await downloadBytes(single);

await meeting.getByRole('button', {name: 'Resume', exact: true}).click();
await expect(meeting.locator('header')).toContainText('Recording');
await meeting.getByRole('button', {name: 'Pause', exact: true}).click();
await expect(transcript).toHaveCount(2);
await expect(transcript.first()).toHaveText(firstText);
await expect(transcript.nth(1)).toContainText(/Mock transcription \([1-9]\d* bytes\)/);
expect(transcriptions).toHaveLength(2);
await expect(meeting.locator('audio')).toHaveCount(2);
await meeting.getByRole('button', {name: 'Stop', exact: true}).click();
await expect(meeting.locator('header')).toContainText('Done');

await expect(meeting.locator('.obe-meeting-summary')).toHaveCount(0);
const generation = page.waitForResponse((response) => response.url().includes('/api/ai/generate') && response.ok());
await meeting.getByRole('button', {name: 'Generate summary', exact: true}).click();
expect((await generation).headers()['content-type']).toContain('text/event-stream');
await expect(meeting.getByRole('button', {name: 'Regenerate summary'})).toBeVisible();
const summary = await meeting.locator('.obe-meeting-summary').innerText();
expect(summary).toContain('Mock response to:');
const note = 'Ellis will send the meeting follow-up.';
await meeting.locator('.obe-meeting-notes .obe-text').fill(note);
const pageId = new URL(page.url()).searchParams.get('page');
await expect.poll(async () => {
const stored = await request.get(`${dataServer}/api/pages/${pageId}`);
expect(stored.ok()).toBeTruthy();
return JSON.stringify((await stored.json() as {data: {blockdoc: unknown}}).data.blockdoc);
}).toContain(note);

const zipDownload = page.waitForEvent('download');
await meeting.getByRole('button', {name: 'Export audio', exact: true}).click();
const zip = await zipDownload;
expect(zip.suggestedFilename()).toMatch(/-audio\.zip$/);
const files = unzipSync(await downloadBytes(zip));
expect(Object.keys(files)).toHaveLength(2);
expect(Object.keys(files)).toEqual([expect.stringMatching(/^001-\d{2}h\d{2}m\d{2}s\.webm$/), expect.stringMatching(/^002-\d{2}h\d{2}m\d{2}s\.webm$/)]);
expect(files[Object.keys(files)[0]]).toEqual(new Uint8Array(singleBytes));
for (const bytes of Object.values(files)) expect(bytes.length).toBeGreaterThan(0);

await page.getByRole('button', {name: 'Page actions'}).click();
await page.getByRole('menuitem', {name: 'Export', exact: true}).click();
const markdownDownload = page.waitForEvent('download');
await page.getByRole('menuitem', {name: 'Markdown (.md)', exact: true}).click();
const markdown = (await downloadBytes(await markdownDownload)).toString('utf8');
for (const segment of await transcript.allTextContents()) {
expect(markdown).toContain(segment.replace(/^\S+\s+/, '').trim());
}
expect(markdown).toContain(summary);
expect(markdown).toContain(note);
});

test('microphone denial shows recovery copy and leaves the editor usable', {tag: ['@meeting', '@p1']}, async ({page}) => {
const errors: string[] = [];
page.on('pageerror', (error) => errors.push(error.message));
await page.evaluate(() => {
navigator.mediaDevices.getUserMedia = async () => { throw new DOMException('Permission denied', 'NotAllowedError'); };
});
const meeting = page.getByRole('region', {name: 'Meeting', exact: true});
await meeting.getByRole('button', {name: 'Record', exact: true}).click();
await expect(meeting.getByRole('alert')).toHaveText('Microphone access failed. Allow microphone access in browser settings, then try again.');
await expect(meeting.getByRole('button', {name: 'Record', exact: true})).toBeEnabled();
await expect(meeting.getByRole('button', {name: 'Stop', exact: true})).toBeDisabled();
await meeting.locator('.obe-meeting-notes .obe-text').fill('Notes still work without microphone access.');
await expect(meeting.locator('.obe-meeting-notes')).toContainText('Notes still work without microphone access.');
await expect(meeting.locator('audio')).toHaveCount(0);
expect(errors).toEqual([]);
});
Loading