Skip to content

[Feature]: out-of-tree graph store keyed by repository, shared across worktrees #3889

Description

@lexfrei

Problem or use case

I'd like graphify to keep its graphs outside the repositories it indexes, in a per-user store. I want the graphs. What I don't want is graphify-out/ inside other people's repos, or a set of workarounds I have to maintain to keep it out.

On a typical day I clone about ten repositories and open around three git worktrees in each, one per agent task. Since graphify writes graphify-out/ into the checkout, every clone and every worktree needs its own handling:

  • The output has to be ignored somewhere, and the repo's .gitignore isn't mine to edit. So it goes into .git/info/exclude or a global excludes file, without a trailing slash in case it's a symlink.
  • A fresh worktree has no graph, and rebuilding one per worktree costs time and LLM calls. Copying the primary checkout's graph in doesn't help: incremental extraction then treats the whole corpus as changed (Incremental semantic extraction reprocesses the corpus and rewrites topology in a Git worktree #2654).
  • A symlink from each worktree to the primary checkout's graphify-out needs a post-checkout hook in init.templateDir. That hook doesn't fire in repos that set their own core.hooksPath. An update from the worktree also writes straight through the link into the primary graph.
  • An absolute GRAPHIFY_OUT is one global value with no per-repo key. The skill ignores it anyway (graphify skill doesnt respect GRAPHIFY_OUT #2571).

Each of these works on its own terms. Together they are a pile of glue per clone, just to read someone else's code with a graph.

Proposed solution

My proposal is an opt-in store mode. Graphs live in a per-user directory keyed by repository, nothing is written to the working tree, and one global setting turns it on. After that a new clone or worktree needs no preparation.

I read through v8 to check how much this touches. Most of it already works with an absolute GRAPHIFY_OUT, so the design is mostly about computing that value per repo:

  1. Resolution happens in graphify/paths.py at import time, and the first match wins. If GRAPHIFY_OUT is set, nothing changes. If an in-tree graphify-out/ already exists at the repo top, nothing changes either, so teams that commit their graph are not affected. If store mode is on, the result is <store>/<key>. Otherwise it stays graphify-out. The resolved value is absolute and goes into GRAPHIFY_OUT, so the modules that copy the constant at import keep working as they are.
  2. Store mode is enabled with GRAPHIFY_STORE=1 or with {"store": true} in ~/.graphify/config.json. The config file is there because git hooks run from GUI clients don't see the shell environment. The default location is ~/.graphify/store/, next to the per-user state graphify already keeps, and GRAPHIFY_STORE_DIR overrides it.
  3. The key is the repo directory name plus a short hash of git rev-parse --path-format=absolute --git-common-dir. That gives every linked worktree of a clone the same key. Outside git, the hash comes from the absolute cwd. Each entry also writes an origin.json with the common dir, the primary worktree path and the remote URL, for listing and pruning later.
  4. Worktrees share the clone's graph read-only. Read commands, hook-guard and MCP work from any worktree. Write commands (extract, update, watch, cluster-only) exit with a message pointing at the primary checkout. The existing worktree guard in the hooks already skips rebuilds there. Since source_file paths are relative to the scan root, a graph built in the primary checkout resolves fine from a worktree. In store mode .graphify_root is always written absolute. This would also cover the read-side fallback from fix(cli): resolve the graph from the primary checkout in linked worktrees (#2008) #2064.
  5. A new graphify where command prints the resolved output dir exactly as the process sees it. It is the only new CLI surface, and the skill uses it to find the graph.
  6. The skill and the always-on text stop hardcoding the path. Each bash block starts with GFY_OUT="$(graphify where 2>/dev/null || echo graphify-out)", and every literal graphify-out/ becomes "$GFY_OUT"/. Inline Python gets the path from an env var, never through string interpolation. In default mode graphify where prints graphify-out, so current behaviour doesn't change. This overlaps with fix: respect GRAPHIFY_OUT env var in skill files and install templates #1054 and Honor GRAPHIFY_OUT in installed docs and IDE plugins #2720, and I'd rather work with those authors than open a competing change.
  7. graphify store list shows each entry with its origin, size and last write, and flags orphans whose clone is gone. graphify store prune removes orphans and is a dry run by default.
  8. In store mode hook install doesn't write the .gitattributes merge-driver line, since there is no in-tree graph to merge. Nothing else in the hooks changes.

I'd split the work into three PRs, each useful on its own:

  1. Resolution, the key, origin.json, the write guard in linked worktrees, and graphify where. Tests would use real git worktree add fixtures, like the existing worktree-guard test in test_hooks.py. CLI, hooks and MCP users get the benefit from this one alone.
  2. Skill fragments, always-on blocks and the agent text in install.py switch to graphify where, including the --bless run and the sanctioned always-on edits.
  3. graphify store list and graphify store prune.

A few things I'm not sure about:

  • Should store mode stay opt-in, or become the default in some later major version?
  • Is a shared read-only graph the right policy for worktrees? The alternative is an opt-in per-worktree graph keyed by --git-dir, seeded from the clone's graph once Incremental semantic extraction reprocesses the corpus and rewrites topology in a Git worktree #2654 is fixed.
  • Is ~/.graphify/store/ a good default, or would you prefer $XDG_DATA_HOME/graphify/ on Linux? I wouldn't put it under a cache dir, because the semantic pass costs money to rebuild.

Alternatives considered

Area

CLI or installation

Are you willing to submit a PR?

Yes, I can work on this. If the design looks right to you, I'll start with the first PR, and I'm happy to change any part of it before that.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions