Generate AI-powered natural-language summaries of git diffs for code review in any git repository.
Use any LLM provider supported by the Vercel AI SDK (OpenAI, Anthropic, Google Gemini, Amazon Bedrock, Mistral, Cohere, Groq, xAI, DeepSeek, and OpenAI-compatible gateways).
- Node.js 22.22.1+
- An LLM provider credential (see Provider configuration)
- No git binary required, on any platform — smart-diff reads the git repository directly via
@scolladon/tsgit, a pure-TypeScript git implementation with zero native dependencies
npm install @mcarvin/smart-diffAll provider packages are optional — only install the one(s) you need:
# OpenAI
npm install @ai-sdk/openai
# Anthropic
npm install @ai-sdk/anthropic
# Google Gemini
npm install @ai-sdk/google
# Amazon Bedrock
npm install @ai-sdk/amazon-bedrock
# OpenAI-compatible gateway (Azure, Ollama, Together, etc.)
npm install @ai-sdk/openai-compatible
# Others: @ai-sdk/mistral @ai-sdk/cohere @ai-sdk/groq @ai-sdk/xai @ai-sdk/deepseekIf the package for the selected provider is missing at runtime, smart-diff throws a clear error telling you which one to install.
smart-diff is "configured" when isLlmProviderConfigured() returns true — i.e. at least one supported provider can be resolved from env vars — or you pass your own llmModelProvider factory. Otherwise summarizeGitDiff / generateSummary throw with LLM_GATEWAY_REQUIRED_MESSAGE.
LLM_PROVIDER explicitly selects a provider. When unset, the resolver auto-detects in this order: LLM_BASE_URL/OPENAI_BASE_URL → openai-compatible, OPENAI_API_KEY/LLM_API_KEY → openai, then ANTHROPIC_API_KEY, GOOGLE_GENERATIVE_AI_API_KEY (or GOOGLE_API_KEY), MISTRAL_API_KEY, COHERE_API_KEY, GROQ_API_KEY, XAI_API_KEY, DEEPSEEK_API_KEY, AWS_ACCESS_KEY_ID/AWS_PROFILE → bedrock, and finally OPENAI_DEFAULT_HEADERS/LLM_DEFAULT_HEADERS → openai.
Provider (LLM_PROVIDER) |
Package | Credential env vars | Default model |
|---|---|---|---|
openai |
@ai-sdk/openai |
OPENAI_API_KEY or LLM_API_KEY |
gpt-4o-mini |
openai-compatible |
@ai-sdk/openai-compatible |
LLM_BASE_URL or OPENAI_BASE_URL (required); OPENAI_API_KEY/LLM_API_KEY or custom headers |
gpt-4o-mini |
anthropic |
@ai-sdk/anthropic |
ANTHROPIC_API_KEY |
claude-haiku-4-5-20251001 |
google |
@ai-sdk/google |
GOOGLE_GENERATIVE_AI_API_KEY or GOOGLE_API_KEY |
gemini-2.0-flash |
bedrock |
@ai-sdk/amazon-bedrock |
AWS_ACCESS_KEY_ID / AWS_PROFILE (auto-detects); full AWS credential chain supported |
anthropic.claude-3-5-haiku-20241022-v1:0 |
mistral |
@ai-sdk/mistral |
MISTRAL_API_KEY |
mistral-small-latest |
cohere |
@ai-sdk/cohere |
COHERE_API_KEY |
command-r-08-2024 |
groq |
@ai-sdk/groq |
GROQ_API_KEY |
llama-3.1-8b-instant |
xai |
@ai-sdk/xai |
XAI_API_KEY |
grok-2-latest |
deepseek |
@ai-sdk/deepseek |
DEEPSEEK_API_KEY |
deepseek-chat |
LLM_*wins overOPENAI_*where both exist.
| Variable | Purpose |
|---|---|
LLM_PROVIDER |
Explicit provider id from the table above. |
LLM_MODEL |
Overrides the per-provider default model id. |
OPENAI_BASE_URL / LLM_BASE_URL |
Base URL for an OpenAI-compatible gateway; presence alone auto-selects the openai-compatible provider. |
OPENAI_DEFAULT_HEADERS / LLM_DEFAULT_HEADERS |
JSON object of extra headers merged onto OpenAI / OpenAI-compatible requests (e.g. RBAC tokens, raw Authorization). LLM_* overrides OPENAI_* key-by-key. |
LLM_PROVIDER_NAME |
Display name used when openai-compatible is active (defaults to openai-compatible). |
OPENAI_MAX_DIFF_CHARS / LLM_MAX_DIFF_CHARS |
Max size of unified diff text sent to the model (default ~120k characters). |
OPENAI_MAX_TOKENS / LLM_MAX_TOKENS |
Max completion tokens (default 4000). |
LLM_TEMPERATURE |
Sampling temperature, clamped to 0–2 (default 0.2). Lower = more deterministic; higher = more varied prose. |
LLM_MAX_RETRIES |
Retry count for transient LLM call failures — rate limits, 5xx, network errors (default 2, matching the Vercel AI SDK's own default). Set to 0 to disable retries. |
$env:OPENAI_API_KEY = "sk-..."
# Optional: $env:LLM_MODEL = "gpt-4o"$env:ANTHROPIC_API_KEY = "sk-ant-..."
$env:LLM_MODEL = "claude-3-5-sonnet-latest" # optional override$env:OPENAI_BASE_URL = "https://llm-gateway.example.com"
$env:OPENAI_DEFAULT_HEADERS = '{"x-company-rbac":"your-rbac-token-here","Authorization":"Bearer sk-your-api-key-here"}'
# LLM_PROVIDER is auto-detected as "openai-compatible" because LLM_BASE_URL/OPENAI_BASE_URL is set.$env:GOOGLE_GENERATIVE_AI_API_KEY = "..."
$env:LLM_MODEL = "gemini-2.0-flash"Use smart-diff as a library, or as the smart-diff CLI binary that ships with the package via the bin field — no separate install. Provider configuration is identical either way; see Provider configuration. Every option below has a library form (a key in the summarizeGitDiff object) and, where applicable, an equivalent kebab-case CLI flag.
import { summarizeGitDiff } from '@mcarvin/smart-diff';
const { summary, usage } = await summarizeGitDiff({
from: 'origin/main',
to: 'HEAD',
cwd: '/path/to/repo', // optional; default process.cwd()
includeFolders: ['src'],
excludeFolders: ['node_modules', 'dist'],
commitMessageExcludeRegexes: ['^\\[bot\\]'],
commitMessageIncludeRegexes: ['^feat:'], // optional; OR across patterns
teamName: 'Platform',
systemPrompt: undefined, // optional; overrides DEFAULT_GIT_DIFF_SYSTEM_PROMPT
provider: 'anthropic', // optional; overrides LLM_PROVIDER env + auto-detection
model: 'claude-3-5-sonnet-latest', // optional
maxDiffChars: 120_000, // optional; also see LLM_MAX_DIFF_CHARS
});
// summary: Markdown string
// usage: LlmUsageReport aggregated across every LLM call — see Token usage reporting belownpm install @mcarvin/smart-diff
npx smart-diff origin/main HEAD --team Platform --max-diff-chars 20000
npx smart-diff --help<from>(required) and[to](defaultHEAD) can be passed positionally or via--from/--to.- Pass
--merge-base/-balongside<from>/--fromto resolve it as the merge base oftoandfrominstead of using it directly — e.g.--to develop --from main --merge-baseis the tsgit-native equivalent of--to develop --from $(git merge-base develop main), with no local git binary required. - Repeatable options (
--include,--exclude,--commit-include,--commit-exclude) accept multiple flags. - The Markdown summary is printed to stdout; errors go to stderr and exit with code 1.
- Run
smart-diff --helpfor the full flag reference, orsmart-diff --versionfor the installed version.
| Option | CLI flag | Description |
|---|---|---|
from / to |
<from> [to] / --from / --to |
Git refs for the range; to defaults to HEAD. |
mergeBase |
--merge-base, -b |
Resolve from as the merge base of to and from, in-process via tsgit — no local git binary needed. |
cwd / git |
--cwd <path> |
Working directory path, or inject your own GitClient instance (library only; see Lower-level API). |
includeFolders |
--include <path> |
Limit diff to these paths relative to repo root (omit for full repo minus excludes). |
excludeFolders |
--exclude <path> |
Excluded paths, applied client-side to the changed-path list (directory-prefix match), e.g. node_modules. |
commitMessageIncludeRegexes |
--commit-include <regex> |
If any pattern is non-empty, only commits whose full message matches at least one pattern are kept (after excludes). Case-insensitive. |
commitMessageExcludeRegexes |
--commit-exclude <regex> |
Drop commits whose message matches any of these patterns. |
teamName |
--team <name> |
Adds a Team: line to the user payload for the model. |
systemPrompt |
--system-prompt <text> |
Replaces the default system prompt. |
provider |
--provider <id> |
LlmProviderId — wins over LLM_PROVIDER env and auto-detection. |
model |
--model <id> |
Chat model id; overrides LLM_MODEL and the provider default. |
llmModelProvider |
— (library only) | () => Promise<LanguageModel> — bypass env-based resolution entirely; hand-wire a Vercel AI SDK LanguageModel (required in tests or custom setups). |
Token reduction — see Handling large diffs
| Option | CLI flag | Description |
|---|---|---|
maxDiffChars |
--max-diff-chars <n> |
Caps unified diff size for the request; see LLM_MAX_DIFF_CHARS. |
mapReduce |
--map-reduce |
Split an oversized diff into per-file batches instead of hard-truncating it. |
contextLines |
--context-lines <n> |
Number of context lines around each change (git diff -U<n>). Lower values (1 or 0) are the single biggest token saver on modification-heavy diffs. |
ignoreWhitespace |
--ignore-whitespace |
Passes -w / --ignore-all-space to git diff so pure-whitespace hunks don't consume tokens. Also applies to --numstat / --name-status so counts stay consistent. |
stripDiffPreamble |
--strip-diff-preamble |
Removes low-value lines from the unified diff (diff --git, index, mode changes, similarity/rename/copy metadata). --- a/…, +++ b/…, and @@ hunk headers are kept. |
maxHunkLines |
--max-hunk-lines <n> |
Caps the body of each hunk; anything past the limit is replaced with a single elision marker. The @@ header and DiffSummary totals are preserved. |
excludeDefaultNoise |
--exclude-default-noise |
Merges the built-in DEFAULT_NOISE_EXCLUDES list (lockfiles, dist, build, out, coverage, node_modules, __snapshots__) into excludeFolders. |
| Option | CLI flag | Description |
|---|---|---|
maxRetries |
--max-retries <n> |
Retry count for transient LLM call failures; see LLM_MAX_RETRIES. |
| Option | CLI flag | Description |
|---|---|---|
redactSecrets |
--redact-secrets |
Masks likely secrets/credentials in the diff text before it's sent to the LLM — cloud provider keys, VCS/chat tokens, PEM private key blocks, JWTs, Bearer headers, basic-auth URL passwords, and generic KEY=value assignments. Uses DEFAULT_SECRET_PATTERNS unless secretPatterns overrides them. |
secretPatterns |
— (library only) | Overrides the built-in secret-detection rules used when redactSecrets is true. |
| Option | CLI flag | Description |
|---|---|---|
— (summarizeGitDiff always returns usage) |
--usage |
Print the same LlmUsageReport as Token usage reporting to stderr, after the Markdown summary on stdout. |
| — | -h, --help |
Show CLI help. |
| — | -v, --version |
Print the installed version. |
For most repos, the cheapest wins are:
await summarizeGitDiff({
from: 'origin/main',
contextLines: 1, // -U1 cuts 30-60% of tokens on typical diffs
ignoreWhitespace: true, // drop pure-whitespace hunks entirely
stripDiffPreamble: true, // kill `index`/`mode`/`similarity` lines
maxHunkLines: 400, // truncate monster hunks but keep the @@ header
excludeDefaultNoise: true // skip lockfiles, dist/, coverage/, node_modules/
});These options only reshape the unified diff text — the structured DiffSummary still reports true file counts and line totals, so the model always sees the full change inventory.
By default, a diff over maxDiffChars is hard-truncated — only the first N characters are sent, and a notice is prepended to the summary. Set mapReduce: true to instead split the diff into per-file batches, summarize each batch independently, then synthesize one final summary from the batch summaries:
await summarizeGitDiff({
from: 'origin/main',
maxDiffChars: 20_000,
mapReduce: true, // no-op when the diff already fits within maxDiffChars
});This costs one extra LLM call per batch plus one reduce call, so it's slower and more expensive than a single request — use it when losing coverage of the tail of a large diff matters more than latency/cost. systemPrompt (if set) applies to the reduce phase only; the map phase always uses DEFAULT_MAP_REDUCE_MAP_SYSTEM_PROMPT.
summarizeGitDiff returns the Markdown summary plus token usage aggregated across every LLM call made to produce it — one call by default, or every map-reduce batch plus the reduce call when mapReduce is used:
import { summarizeGitDiff } from '@mcarvin/smart-diff';
const { summary, usage } = await summarizeGitDiff({ from: 'origin/main' });
// usage: { requestCount, inputTokens, outputTokens, totalTokens, cachedInputTokens }This reports token counts only — no dollar cost estimate, since per-token pricing varies by provider/model and changes over time; a hardcoded price table would drift out of date. Fields are 0 when a provider doesn't report a given figure. generateSummaryWithUsage is the lower-level equivalent of generateSummary.
If you want full control — for example, to configure retries, middlewares, or hit an in-process mock — pass llmModelProvider:
import { summarizeGitDiff } from '@mcarvin/smart-diff';
import { createAnthropic } from '@ai-sdk/anthropic';
const { summary } = await summarizeGitDiff({
from: 'origin/main',
llmModelProvider: async () =>
createAnthropic({ apiKey: process.env.MY_ANTHROPIC_KEY })(
'claude-3-5-sonnet-latest',
),
});- Single unified diff for
from..towhen no commit-message filters apply and the filtered commit list matches the full log for that range. - Concatenated per-commit patches (each commit diffed against its first parent, or the empty tree for a root commit) when you use include/exclude regexes or when the filtered commit list differs in length from the full range (so the diff reflects only the commits that remain).
The package also exports helpers for building a custom pipeline on top of the same git and LLM behavior:
- Git:
createGitClient(cwd?, timeout?)(async — returnsPromise<GitClient>, a live@scolladon/tsgitrepository handle; callgit.dispose()when done),getRepoRoot,getCommits,getDiff,getDiffSummary,getChangedFiles,filterCommitsByMessageRegexes,buildPathFilterPredicate,matchesAnyPath,renderFileDiff,renderUnifiedDiff,buildFileSummary,mergeFileSummariesByPath,summarizeFiles,shapeUnifiedDiff,redactSecrets,DEFAULT_NOISE_EXCLUDES,DEFAULT_SECRET_PATTERNS—timeoutis in milliseconds; omit for no timeout - AI:
generateSummary,generateSummaryWithUsage,resolveLlmMaxDiffChars,resolveLlmMaxRetries,truncateUnifiedDiffForLlm,splitUnifiedDiffIntoFileChunks,groupDiffChunksByBudget - Provider resolution:
resolveLanguageModel,detectLlmProvider,isLlmProviderConfigured,defaultModelForProvider,resolveLlmBaseUrl,parseLlmDefaultHeadersFromEnv - Constants:
DEFAULT_GIT_DIFF_SYSTEM_PROMPT,DEFAULT_MAP_REDUCE_MAP_SYSTEM_PROMPT,DEFAULT_MAP_REDUCE_REDUCE_SYSTEM_PROMPT,LLM_GATEWAY_REQUIRED_MESSAGE - Types:
LlmProviderId,LlmModelProvider,ResolveLanguageModelOptions,GenerateSummaryInput,SummarizeFlags,LlmUsageReport,DiffFileSummary,DiffSummary,CommitInfo,GitClient,GitDiffRangeQuery,DiffPathFilter,DiffShapingOptions,PathFilterPredicate,RenderedFileDiff,SecretRedactionRuleDiffFileSummary.binary?: booleanistruewhen either blob sniffs as binary; absent for text files
This package is used downstream by:
- sf-git-ai-meta-insights — Salesforce metadata wrapper compatible with Salesforce DX projects
- gitlab-llm-kit — TypeScript toolkit for GitLab REST API access and LLM-powered insights (merge requests, issues, pipelines, security, releases, wiki, and more)