Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 44 additions & 0 deletions docs/local-transcription.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,44 @@
# Local transcription

OpenBook transcribes recordings locally by default, without a cloud API key. The server uses the optional whisper.cpp `whisper-cli` executable and FFmpeg. Neither executable nor model weights are bundled with OpenBook; normal CI skips native inference explicitly.

## Setup

1. Install whisper.cpp and FFmpeg on the server host. Follow the [whisper.cpp build instructions](https://github.com/ggml-org/whisper.cpp#quick-start) or use your host package manager.
2. Put `whisper-cli` and `ffmpeg` on the server process's PATH. Alternatively, set executable paths before starting OpenBook:

```sh
export OPENBOOK_WHISPER_BIN=/absolute/path/to/whisper-cli
export OPENBOOK_FFMPEG_BIN=/absolute/path/to/ffmpeg
```

3. In **Settings → AI**, select **Download Whisper base**. This uses the existing authenticated model download and progress flow. Downloads go to `OPENBOOK_MODELS_DIR` when set, otherwise the server data directory's `models` folder (or `~/.openbook/models` without a data directory). The completed model is discovered without restarting.
4. Check that Settings reports the runtime and model ready, then transcribe a recording.

The default model is multilingual Whisper base (`ggml-base.bin`, approximately 142 MiB), with automatic language detection. It trades some accuracy on noisy speech, accents, and difficult multilingual recordings for a smaller download and lower memory use than larger models. Whisper downloads do not change the selected chat model.

Explicit cloud transcription configuration takes precedence over local inference. Otherwise the server tries local transcription, then the existing mock fallback. Local transcription works even with chat disabled; explicitly disabling transcription still disables it. Missing executables or model weights produce an actionable error pointing to Settings → AI.

## Processing and limits

FFmpeg converts recordings to mono 16 kHz PCM WAV. Whisper loads the model for each job and returns text plus segments in seconds; `durationMs` is the rounded maximum segment endpoint in milliseconds. Each job uses a private temporary directory. Cancellation and server shutdown kill active child processes, wait for them to close, and remove scratch files.

Each server permits at most **two local transcription jobs at once**, shared across all clients. There is no queue. Busy requests return HTTP **429** with `Retry-After: 5`. Permits remain held through temporary-file cleanup and are released on success, failure, or cancellation.

The transcription route also allows **six local requests per socket IP per 60-second fixed window**. Excess requests return HTTP **429** with `Retry-After: 60`. Client-supplied forwarding headers do not change this key; clients behind a reverse proxy may share its socket IP budget. The limit applies only when the resolved backend is local; cloud and mock backends are unaffected.

## Native smoke test

Install the optional executables and download `ggml-base.bin` into your model directory, then run from the repository root:

```sh
OPENBOOK_TEST_WHISPER=1 \
OPENBOOK_MODELS_DIR=/absolute/path/to/models \
OPENBOOK_WHISPER_BIN=/absolute/path/to/whisper-cli \
OPENBOOK_FFMPEG_BIN=/absolute/path/to/ffmpeg \
VITEST_MAX_WORKERS=1 \
pnpm --filter @book.dev/server exec vitest run src/transcription.test.ts \
-t 'native whisper'
```

This opt-in test synthesizes a small WAV and exercises the HTTP transcription route without cloud keys, asserting the result shape. It is a runtime integration check, not a speech-accuracy benchmark. Without `OPENBOOK_TEST_WHISPER=1`, the native test is explicitly skipped. The regular subprocess-fixture tests cover concurrency, cancellation, failures, timestamp conversion, and cleanup without installing native dependencies.
2 changes: 2 additions & 0 deletions packages/sdk/src/ai.ts
Original file line number Diff line number Diff line change
Expand Up @@ -212,6 +212,8 @@ export interface AiUsageResponse {
}

export interface AiStatus {
/** Local audio capability is independent of the chat provider. */
transcription?: {model: string; modelPresent: boolean; runtimeAvailable: boolean; ready: boolean; downloadUrl: string; detail?: string};
config: AiConfig;
/** The engine can generate text right now. */
ready: boolean;
Expand Down
23 changes: 21 additions & 2 deletions packages/server/src/ai/routes.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,11 +6,16 @@ import type {PageStore} from '../store';
import type {AppEnv} from '../appEnv';
import {isLocalInstanceOwner, requireAuthenticatedRead, requireCreate, requireInstanceAdmin, requireInstanceOwner} from '../access';
import {AgentRunner, type AgentMessage} from './agent';
import {TranscriptionConfigError, type AiService} from './service';
import {ModelDownloadConfigError, TranscriptionConfigError, type AiService} from './service';
import {McpConfigError, type ExternalAgentTool, type McpClientManager} from './mcpClients';
import {FixedWindowLimiter, clientIpKey} from '../agentTokens';
import {LocalTranscriptionBusyError} from './whisper';
import type {TokenUsage} from './providers';
import type {AiUsageLog, UsageKind} from './usage';

export const LOCAL_TRANSCRIPTION_RATE_LIMIT = 6;
const LOCAL_TRANSCRIPTION_RATE_WINDOW_MS = 60_000;

/**
* The `/api/ai/*` surface. Generation endpoints stream tokens as SSE
* (`data: {"token": "..."}` frames, closed by `data: {"done": true}`);
Expand All @@ -22,6 +27,7 @@ import type {AiUsageLog, UsageKind} from './usage';
* logging failure never breaks the request.
*/
export function mountAiRoutes(app: Hono<AppEnv>, ai: AiService, store: PageStore, onPagesChanged?: () => Promise<void>, aiUsage?: AiUsageLog, mcp?: McpClientManager): void {
const localTranscriptionLimiter = new FixedWindowLimiter(LOCAL_TRANSCRIPTION_RATE_LIMIT, LOCAL_TRANSCRIPTION_RATE_WINDOW_MS);
/**
* Log a single generate/complete usage row against the effective provider/model.
* Best-effort (the logger swallows its own errors); does nothing without a logger
Expand Down Expand Up @@ -195,12 +201,20 @@ export function mountAiRoutes(app: Hono<AppEnv>, ai: AiService, store: PageStore
try {
const {engine, provider, model} = await ai.transcriptionBackend();
await requirePaidInferenceAccess(c, store, provider === 'openai-compat' ? 'openai' : 'off');
if (provider === 'local' && localTranscriptionLimiter.exceeded(clientIpKey(c))) {
c.header('Retry-After', String(LOCAL_TRANSCRIPTION_RATE_WINDOW_MS / 1000));
return c.json({error: 'Too many local transcription requests. Please retry shortly.'}, 429);
}
const result = await engine.transcribe(asset.bytes, {mime: asset.mime, signal: c.req.raw.signal});
// Audio backends do not report token counts. Keep unknown cloud cost null.
await aiUsage?.log({provider, model, kind: 'transcribe', principal, usage: {inputTokens: 0, outputTokens: 0}});
return c.json(result);
} catch (err) {
if (err instanceof HTTPException) throw err;
if (err instanceof LocalTranscriptionBusyError) {
c.header('Retry-After', String(err.retryAfterSeconds));
return c.json({error: err.message}, 429);
}
if (err instanceof TranscriptionConfigError) return c.json({error: err.message}, 400);
return c.json({error: 'Transcription failed. Check the provider in Settings → AI and retry.'}, 502);
}
Expand All @@ -211,7 +225,12 @@ export function mountAiRoutes(app: Hono<AppEnv>, ai: AiService, store: PageStore
// surface) — only the trusted instance owner may supply it.
await requireInstanceOwner(c, store);
const {url} = (await c.req.json().catch(() => ({}))) as {url?: string};
return c.json(await ai.startDownload(url));
try {
return c.json(await ai.startDownload(url));
} catch (err) {
if (err instanceof ModelDownloadConfigError) return c.json({error: err.message}, 400);
throw err;
}
});

// The agent harness: runs the tool loop against the library and streams
Expand Down
19 changes: 15 additions & 4 deletions packages/server/src/ai/service.ts
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ import {providerSettings, type AiConfig, type AiProvider, type AiProviderSetting
import type {Db} from '../db';
import {createEngine, MockEngine, OpenAiCompatEngine, type TranscriptionEngine, type AiEngine, type GenerateOptions} from './providers';
import {assembleSearchResults, bm25Scores, buildIndex, cosine, pageRowsToDocs, parseTaskList, type Bm25Index} from './search';
import {WHISPER_MODEL, WHISPER_MODEL_URL} from './whisper';
import {SkillStore} from './skills';

/**
Expand Down Expand Up @@ -96,6 +97,7 @@ interface DownloadState {
}

export class TranscriptionConfigError extends Error {}
export class ModelDownloadConfigError extends Error {}

export class AiService {
private config: AiConfig = DEFAULT_CONFIG;
Expand All @@ -114,6 +116,7 @@ export class AiService {
private readonly modelsDir: string,
/** MEET-3: lazily resolve the managed local audio backend. */
private readonly localTranscription?: () => Promise<TranscriptionEngine | null>,
private readonly localLifecycle?: {status(): Promise<NonNullable<AiStatus['transcription']>>; dispose(): Promise<void>},
) {
this.skills = new SkillStore(db);
}
Expand Down Expand Up @@ -162,9 +165,9 @@ export class AiService {
return {engine: new OpenAiCompatEngine(audio.baseUrl?.trim() || 'https://api.openai.com', audio.model?.trim() || 'whisper-1', 'openai', audio.apiKey), provider: 'openai-compat', model: audio.model?.trim() || 'whisper-1'};
}
const local = await this.localTranscription?.();
if (local) return {engine: local, provider: 'local', model: audio?.model ?? 'local'};
if (local) return {engine: local, provider: 'local', model: WHISPER_MODEL};
if (config.provider === 'mock') return {engine: new MockEngine(), provider: 'mock', model: 'mock'};
throw new TranscriptionConfigError('Local transcription is unavailable. Configure transcription in Settings → AI.');
throw new TranscriptionConfigError('Local transcription is unavailable. Download the Whisper model and check the local runtime in Settings → AI.');
}

async status(): Promise<AiStatus> {
Expand All @@ -183,6 +186,7 @@ export class AiService {
}
return {
config: this.config,
transcription: await this.localLifecycle?.status(),
ready,
embeddings,
detail,
Expand Down Expand Up @@ -336,7 +340,13 @@ export class AiService {
async startDownload(url = DEFAULT_MODEL_URL): Promise<DownloadState> {
await this.loadConfig();
if (this.download && !this.download.done && !this.download.error) return this.download;
const fileName = decodeURIComponent(new URL(url).pathname.split('/').pop() || 'model.gguf');
let fileName: string;
try {
fileName = decodeURIComponent(new URL(url).pathname.split('/').pop() || 'model.gguf');
} catch {
throw new ModelDownloadConfigError('Invalid model download URL or filename encoding.');
}
if (fileName !== path.basename(fileName) || fileName === '.' || fileName === '..' || fileName.includes('\\')) throw new ModelDownloadConfigError('Invalid model filename');
mkdirSync(this.modelsDir, {recursive: true});
const dest = path.join(this.modelsDir, fileName);
const state: DownloadState = {url, received: 0, total: null, done: false};
Expand Down Expand Up @@ -367,7 +377,7 @@ export class AiService {
state.done = true;
}
// Auto-select the downloaded model for the llama provider.
if (this.config.provider === 'llama' && !this.config.model) {
if (url !== WHISPER_MODEL_URL && this.config.provider === 'llama' && !this.config.model) {
await this.setConfig({...this.config, model: fileName});
}
} catch (err) {
Expand All @@ -381,5 +391,6 @@ export class AiService {

async dispose(): Promise<void> {
await this.engine?.dispose().catch(() => undefined);
await this.localLifecycle?.dispose();
}
}
104 changes: 104 additions & 0 deletions packages/server/src/ai/whisper.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
import {mkdtemp, readFile, readdir, rm, writeFile} from 'node:fs/promises';
import {tmpdir} from 'node:os';
import path from 'node:path';
import {afterEach, beforeEach, describe, expect, it} from 'vitest';
import {LocalWhisper, LocalTranscriptionBusyError, parseWhisperOutput, runWhisperProcess, WHISPER_MODEL} from './whisper';

let dir: string;
beforeEach(async () => { dir = await mkdtemp(path.join(tmpdir(), 'meet3-test-')); });
afterEach(async () => { await rm(dir, {recursive: true, force: true}); });

async function script(name: string, source: string): Promise<string> {
const file = path.join(dir, name);
await writeFile(file, `#!${process.execPath}\n${source}`, {mode: 0o700});
return file;
}

describe('optional local whisper runtime', () => {
it('returns null for missing model or executables, and discovers a model without restart', async () => {
const local = new LocalWhisper(dir, process.execPath, process.execPath);
expect(await local.resolve()).toBeNull();
expect(await local.status()).toMatchObject({modelPresent: false, runtimeAvailable: true, ready: false});
await writeFile(path.join(dir, WHISPER_MODEL), 'test model');
expect(await local.resolve()).toBe(local);
const missing = new LocalWhisper(dir, path.join(dir, 'missing'), process.execPath);
expect(await missing.resolve()).toBeNull();
expect(await missing.status()).toMatchObject({modelPresent: true, runtimeAvailable: false, detail: expect.stringContaining('Settings → AI')});
await local.dispose();
expect(await local.resolve()).toBeNull();
});

it('decodes exact input bytes, converts millisecond offsets, and cleans scratch files', async () => {
const marker = path.join(dir, 'scratch');
const ffmpeg = await script('ffmpeg', 'const fs = require(\'node:fs\'); const args = process.argv.slice(2); const input = args[args.indexOf(\'-i\') + 1]; if (fs.readFileSync(input).toString() !== \'recording\') process.exit(2); fs.writeFileSync(args.at(-1), \'wav\');');
const whisper = await script('whisper', `const fs = require('node:fs'); const args = process.argv.slice(2); const output = args[args.indexOf('-of') + 1]; fs.writeFileSync(${JSON.stringify(marker)}, require('node:path').dirname(output)); fs.writeFileSync(output + '.json', JSON.stringify({transcription: [{offsets: {from: 250, to: 1500}, text: ' Hello world.'}]}));`);
const local = new LocalWhisper(dir, whisper, ffmpeg);
expect(await local.transcribe(new TextEncoder().encode('recording'), {filename: '../../escape.webm', mime: 'audio/webm'}))
.toEqual({text: 'Hello world.', segments: [{start: 0.25, end: 1.5, text: 'Hello world.'}], durationMs: 1500});
await expect(readdir(await readFile(marker, 'utf8'))).rejects.toThrow();
});

it.each(['abort', 'dispose'] as const)('%s kills in-flight inference and removes scratch files', async (action) => {
const marker = path.join(dir, 'started');
const ffmpeg = await script('ffmpeg', 'process.exit(0);');
const whisper = await script('whisper', `const fs = require('node:fs'); const args = process.argv.slice(2); fs.writeFileSync(${JSON.stringify(marker)}, JSON.stringify({pid: process.pid, dir: require('node:path').dirname(args[args.indexOf('-of') + 1])})); setInterval(() => {}, 1000);`);
const local = new LocalWhisper(dir, whisper, ffmpeg);
const controller = new AbortController();
const promise = local.transcribe(new Uint8Array([1]), {signal: controller.signal});
const rejected = expect(promise).rejects.toMatchObject({name: 'AbortError'});
await expect.poll(async () => readFile(marker, 'utf8').catch(() => '')).not.toBe('');
const started = JSON.parse(await readFile(marker, 'utf8')) as {pid: number; dir: string};
if (action === 'abort') controller.abort();
else await local.dispose();
await rejected;
expect(() => process.kill(started.pid, 0)).toThrow();
await expect(readdir(started.dir)).rejects.toThrow();
});

it.each(['abort', 'failure'] as const)('limits jobs to two and releases permits after %s', async (outcome) => {
const mode = path.join(dir, 'mode');
await writeFile(mode, 'wait');
const ffmpeg = await script('ffmpeg', 'process.exit(0);');
const whisper = await script('whisper', `const fs = require('node:fs'); const args = process.argv.slice(2); const output = args[args.indexOf('-of') + 1]; fs.writeFileSync(${JSON.stringify(dir)} + '/' + process.pid + '.started', ''); setInterval(() => { const mode = fs.readFileSync(${JSON.stringify(mode)}, 'utf8'); if (mode === 'fail') process.exit(1); if (mode === 'ok') { fs.writeFileSync(output + '.json', JSON.stringify({transcription: []})); process.exit(0); } }, 10);`);
const local = new LocalWhisper(dir, whisper, ffmpeg);
const controller = new AbortController();
const jobs = Promise.allSettled([0, 1].map(() => local.transcribe(new Uint8Array([1]), {signal: controller.signal})));
try {
await expect.poll(async () => (await readdir(dir)).filter((f) => f.endsWith('.started')).length).toBe(2);
await expect(local.transcribe(new Uint8Array([1]))).rejects.toBeInstanceOf(LocalTranscriptionBusyError);
if (outcome === 'abort') controller.abort();
else await writeFile(mode, 'fail');
const results = await jobs;
for (const result of results) {
expect(result.status).toBe('rejected');
if (result.status === 'rejected') expect(result.reason.message).toContain(outcome === 'abort' ? 'abort' : 'Local audio processing failed');
}
await writeFile(mode, 'ok');
const recovered = await Promise.all([0, 1].map(() => local.transcribe(new Uint8Array([1]))));
expect(recovered).toEqual([{text: '', segments: [], durationMs: 0}, {text: '', segments: [], durationMs: 0}]);
} finally {
await local.dispose();
await jobs;
}
});

it.each([1001, 1000.6])('keeps duration integer milliseconds without a seconds round-trip (%s)', (end) => {
expect(parseWhisperOutput({transcription: [
{offsets: {from: 0, to: end}, text: 'First'},
{offsets: {from: 0, to: 900}, text: 'Second'},
]}).durationMs).toBe(1001);
});

it('rejects pre-aborted work, spawn failures, failed processes and malformed output', async () => {
const controller = new AbortController();
controller.abort();
await expect(new LocalWhisper(dir).transcribe(new Uint8Array(), {signal: controller.signal})).rejects.toMatchObject({name: 'AbortError'});
const signal = new AbortController().signal;
await expect(runWhisperProcess(path.join(dir, 'missing'), [], signal)).rejects.toThrow();
await expect(runWhisperProcess(process.execPath, ['-e', 'process.exit(1)'], signal)).rejects.toThrow('Local audio processing failed');
for (const value of [null, {}, {transcription: [{text: 'bad', offsets: {from: -1, to: 100}}]}]) {
expect(() => parseWhisperOutput(value)).toThrow();
}
expect(parseWhisperOutput({transcription: []})).toEqual({text: '', segments: [], durationMs: 0});
});
});
Loading
Loading