Gigagraph
БесплатноНе проверенSemantic code-graph MCP server for coding agents: call graphs, structural similarity, API endpoint maps, blast radius
Описание
Semantic code-graph MCP server for coding agents: call graphs, structural similarity, API endpoint maps, blast radius
README
Semantic code-graph MCP server for coding agents. Indexes what code means structurally — every function, where it's called, what it calls, which packages it depends on — rather than embedding source text. Written in Rust with tree-sitter parsing and rayon multithreading; built to chew through large multi-language monorepos.
Languages: C, C++, C#, Bash (incl. .bats), Go, Java, JavaScript, Kotlin,
Objective-C, PHP, Python, Ruby, Rust, Swift, TypeScript (+TSX) — plus schema
languages:
SQL (tables, views, functions, and their dependency edges), Prisma (models +
relations), GraphQL SDL (types + type dependencies), and shallow YAML.
What it builds
For every function in the tree:
Definition: name, qualified name (
scope::Type::name), file:lines, signature, containing type, param countCall graph: every call site, resolved heuristically to the target function (same file → import-directed → same package → same directory → global name match, with confidence + ambiguity reporting), plus reverse (caller) edges
Package edges: calls that leave the project are attributed to external packages via import analysis (
express,java.util,stdio.h,Foundation,builtin:console, …)Semantic vector: two always-on signals per function, blended at query time (0.6/0.4, fixed — no configuration):
- structural (256 dims): feature-hashing the function's build — callee
names, identifier bag, AST node-type histogram, control-flow shape
(loops/branches/nesting), arity, size, external packages — enriched with
verb-synonym buckets (
fetch/load/get→ one READ feature), camelCase/snake_case subwords, typed-local types, and depth-weighted transitive effect features (the helpers a function reaches and the external packages it ultimately touches, 3 calls deep) — IDF-weighted, L2-normalized; - semantic (64 dims): the function's identifier/callee/type/doc words
embedded with a distilled static embedding model compiled into the binary
(int8-quantized potion-base-2M, ~2.2 MB, MIT — see
src/embed/NOTICE); static table lookup + mean-pool, microseconds per function, no network, no runtime downloads.
Similarity search is brute-force blended cosine over the in-memory matrices, parallelized and fully deterministic: two functions are similar when they're built the same way and named/documented with the same meaning.
- structural (256 dims): feature-hashing the function's build — callee
names, identifier bag, AST node-type histogram, control-flow shape
(loops/branches/nesting), arity, size, external packages — enriched with
verb-synonym buckets (
- API endpoint map: routes the code publishes (Express/Koa/Fastify/Hono/
restify, NestJS, Flask/FastAPI/Django, Laravel/Symfony/Slim, Spring,
ASP.NET, gin/echo/chi/gorilla/net-http, Sinatra/Rails, axum/actix, Ktor)
correlated with the outbound HTTP calls that hit them (fetch, XHR, jQuery,
axios/got/ky/superagent, requests/httpx/aiohttp, Guzzle, net/http,
HttpClient, HTTParty, Retrofit) — matched on normalized path templates
(
/users/:id≡/users/{id}≡/users/{*}) + method, with mount/group/ controller prefix joining and the same high/heuristic confidence labeling as call resolution. - React Native bridge map: JS
NativeModulescall sites correlated with native@ReactMethod(Java/Kotlin) andRCT_EXPORT_METHOD(ObjC) implementations — the cross-language edge static resolution can't see. - Dead-code review queue: functions whose name is referenced nowhere (no call site even unresolved, no identifier reference, no import binding), with framework/entry-point/export suppression and honest caveats.
- Touch memory: a ring of recent edits with agent-supplied rationale
(
record_touch/recent_touches; the handshake instructs agents to record every substantive edit with its rationale — no mechanical hook, docs/HOOKS.md). - 3D code map:
gigagraph visualizerenders an offline WebGL map of the codebase, PCA-projected from the similarity vectors so structurally similar code clusters together.
Index and extraction cache persist under <root>/.gigagraph/ (auto-gitignored).
Re-indexing is incremental: unchanged files (by content hash) skip parsing.
Every MCP tool call runs a stat-only staleness probe first and transparently
re-indexes changed files — answers always reflect the current tree, and an
unchanged tree costs zero file reads.
Install
Prebuilt binaries (recommended)
Grab the latest release for your platform from Releases. One-liners:
# macOS (Apple silicon)
curl -fsSL https://github.com/Raptosaur/gigagraph/releases/latest/download/gigagraph-aarch64-apple-darwin.tar.gz \
| tar -xz && sudo mv gigagraph /usr/local/bin/
# macOS (Intel)
curl -fsSL https://github.com/Raptosaur/gigagraph/releases/latest/download/gigagraph-x86_64-apple-darwin.tar.gz \
| tar -xz && sudo mv gigagraph /usr/local/bin/
# Linux (x86_64)
curl -fsSL https://github.com/Raptosaur/gigagraph/releases/latest/download/gigagraph-x86_64-unknown-linux-gnu.tar.gz \
| tar -xz && sudo mv gigagraph /usr/local/bin/
# Linux (arm64)
curl -fsSL https://github.com/Raptosaur/gigagraph/releases/latest/download/gigagraph-aarch64-unknown-linux-gnu.tar.gz \
| tar -xz && sudo mv gigagraph /usr/local/bin/
Windows: download gigagraph-x86_64-pc-windows-msvc.zip from
Releases, unzip, and put
gigagraph.exe somewhere on your PATH.
Each artifact ships with a .sha256 checksum file.
Cargo
cargo install gigagraph # from crates.io
cargo install --git https://github.com/Raptosaur/gigagraph # from git
From source
git clone https://github.com/Raptosaur/gigagraph && cd gigagraph
cargo build --release # binary at target/release/gigagraph
MCP registration
Claude Code
With gigagraph on your PATH, register it once for all projects:
claude mcp add --scope user gigagraph -- gigagraph serve --root .
Or per project (--scope project writes a shareable .mcp.json). Manual
config equivalent:
{
"mcpServers": {
"gigagraph": {
"command": "gigagraph",
"args": ["serve", "--root", "."]
}
}
}
--root . resolves against the directory the client launches the server in
(your project root for Claude Code), so one user-level registration serves
every project with its own index.
Any MCP client speaking stdio works the same way. The initialize handshake
kicks off a background index sync, so by the time the agent issues its first
structural query the cache is warm; every later tool call re-checks a
stat-only tree fingerprint and refreshes incrementally when files changed.
index_project remains available for an explicit force re-parse.
Optionally, a Claude Code SessionStart hook can pre-warm the index even
before the MCP handshake (useful when sessions start with file reads, not
tool calls) — see docs/HOOKS.md.
Tools
| Tool | What it answers |
|---|---|
index_project |
(Re)index the tree. Incremental; force re-parses all. |
index_stats |
Files/functions/calls, resolution rates, per-language counts, endpoint handler links, and skipped files. |
search_functions |
Find functions by name (exact/prefix/substring/fuzzy). |
get_function |
Full card: location, signature, calls out (resolved), packages used, callers. |
get_callers |
Who calls this? Call sites with file:line, ambiguity flagged. |
get_callees |
Everything this calls: internal, external package, unresolved. |
find_similar |
Structurally similar functions — by indexed function or raw snippet. |
call_path |
Shortest call chain between two functions (BFS). |
file_overview |
One file's imports (classified) + functions — or a whole directory via dir. |
extract_file |
Raw parser output for one file (functions, annotations, literal call args) — why something isn't showing up. |
supported_languages |
Languages + extensions; path answers "would this file be indexed?". |
list_packages |
External packages ranked by call-site count. |
list_endpoints |
Published API surface: REST routes, SOAP/XML-RPC/JSON-RPC operations, gRPC services, GraphQL/AppSync resolvers, and IaC-declared routes (kind filter). |
find_endpoint_callers |
Who calls POST /api/users/:id? Endpoints + their in-repo HTTP callers. |
get_endpoint |
One endpoint's full card: handler, matched clients, confidence. |
list_client_calls |
Outbound HTTP/RPC calls, with unmatched: true filter. |
unreferenced_endpoints |
Endpoints no in-repo client hits (external callers invisible — not proof of dead code). |
unreferenced_functions |
Dead-code review queue; framework/entry-point conventions auto-excluded. |
blast_radius |
Pre-emptive change impact: transitive callers by depth, implicated endpoints, cross-service consumers via correlated HTTP/RPC calls, RN bridge sites, affected-test count. |
affected_tests |
Which tests can a change to this function/file dirty, grouped by file (the re-run unit). |
list_tests |
The standing test inventory: every named case with its runner and suite (file/name/framework/language filters). |
test_command |
The shell command that runs a given test file/case, using the project's own build tooling. |
bridge_map |
React Native bridge: native methods ↔ JS NativeModules call sites. |
visualize |
Self-contained 3D HTML map of the codebase. |
record_touch / recent_touches |
Shared editing memory: what was changed and why, across agents. |
Function references accept fn:<id>, a qualified name, or an unambiguous
simple name.
CLI (debugging)
gigagraph index . # build index, print stats
gigagraph query search_functions '{"query":"parse"}' --root .
gigagraph query find_similar '{"function":"fn:42"}' --root .
cargo run --example dump -- path/to/file.kt # inspect a file's AST
Architecture
src/
lang/ one tree-sitter query + metadata per language (docs/QUERY_CONTRACT.md)
extract.rs generic query-driven extraction: functions, calls, imports, features
graph.rs id assignment, import classification, heuristic call resolution
tests.rs test discovery: annotations, runner naming conventions, BDD blocks
vector.rs feature-hashed structural vectors + parallel blended cosine top-k
verbs.rs identifier word-splitting + verb-synonym bucketing
embed.rs compiled-in distilled static embeddings (src/embed/, ~2.2 MB)
indexer.rs parallel walk (gitignore-aware) -> cached extract -> graph -> vectors
lsp.rs optional LSP enrichment of uncertain edges (tsserver, pyright)
mcp.rs stdio JSON-RPC MCP server
api.rs tool implementations
Parsing is per-file parallel (rayon); resolution is per-function parallel; similarity search is chunk-parallel. Adding a language = one file with a tree-sitter query following the capture contract, plus fixtures.
Testing
Three layers, all under tests/:
- Per-language extraction tests (
python_test.rs,swift_test.rs, …) — spot-checks that a construct is handled. - App fixtures (
tests/fixtures/apps/, asserted byapps_test.rsandlist_tests_test.rs) — eight small but realistic applications spanning all twenty languages, each with its own idiomatic test suite.apps_test.rsasserts set equality between the functions a file defines and the functions extracted, so a regression that silently stops extracting a construct fails even when every spot-check still passes. Constructs the extractor genuinely cannot see are asserted as absent, so the day one starts working is a deliberate change rather than a surprise. - Real-world corpus (
tests/corpus.json,corpus_test.rs) — thirteen open-source applications (Spring PetClinic, axum, googletest, bats-core, Vapor, AFNetworking, the RealWorld reference apps, …) cloned at pinned commits. Fixtures prove a construct is handled; the corpus proves the handling survives vendored trees, generated files and scale.
cargo test # fixtures + unit tests (offline)
scripts/fetch-corpus.sh # clone the pinned corpus
scripts/fetch-corpus.sh --report # measured counts + suggested floors
cargo test --test corpus_test -- --ignored # validate against real apps
Corpus expectations are floors, not equalities — regression detectors rather than quality bars. Where an application's architecture defeats static resolution (Rails autoloading, MediatR indirection), the floor is honestly low and the reason is recorded in the manifest.
Benchmarks
Apple M1 (8 cores), release build (v0.5.x + similarity/LSP layers).
Cold = empty cache, full parse; warm = extraction cache hot, graph +
endpoints + vectors + embeddings rebuilt (gigagraph index CLI, process
start included). An MCP server with an unchanged tree skips rebuilding
entirely — the stat-only fingerprint probe is the whole cost.
| Repo | Files | Functions | Call sites | Cold | Warm |
|---|---|---|---|---|---|
| vuejs/core (TS) | 529 | 5,328 | 56,476 | 1.40 s | 1.14 s |
| square/okhttp (Kotlin/Java) | 677 | 7,693 | 56,927 | 1.37 s | 1.18 s |
| BurntSushi/ripgrep (Rust) | 118 | 3,026 | 17,132 | 1.08 s | 1.01 s |
| pallets/flask (Python) | 85 | 1,515 | 3,897 | 0.99 s | 0.97 s |
| cdk-patterns/serverless (CDK/IaC) | 502 | 846 | 5,439 | 1.08 s | 1.04 s |
Times grew over the early structural-only prototype: the same pass now also detects endpoints (source + IaC), captures DI type information, computes transitive effect features, and embeds every function with the built-in static model — the embedding cost itself is negligible.
Queries against a live server (serve mode, okhttp index in memory,
staleness probe included): search_functions 26 ms, get_callers
22 ms, find_similar 24 ms, blast_radius 18 ms. One-shot CLI queries
pay index load on top (~1 s total). Binary is ~45 MB including the
embedding table.
Honesty about resolution
Cross-file resolution is heuristic (no full type inference), but it is
dependency-injection-aware: declared field/property types, typed
parameters, x = new T() locals, and implements/extends hierarchies are
captured per language, so this.userService.getUser() narrows to
UserService's methods, and calls through an interface expand to its
implementations (single implementor resolves cleanly; several are honest
heuristic with ambiguous_with rivals; implementations outrank abstract
signatures). Explicit container registrations (Laravel $app->bind(A::class, B::class)) name THE implementation and pre-empt the hierarchy fan-out.
Remaining method calls through untyped values (obj.save()) resolve by
method name + receiver hints and are labeled confidence: "heuristic".
Agents should treat high as trustworthy and heuristic as a strong lead.
LSP enrichment
When a real language server is available, gigagraph asks it to settle exactly the call sites the static resolver was unsure about. Providers, detected automatically and run side by side in one pass (each within the shared budget):
- TypeScript — via the
tsserverthat ships inside the project's ownnode_modules(its native protocol, not LSP): needstsconfig.json(orjsconfig.json) at the root,node_modules/typescript/lib/tsserver.js, andnodeon PATH. - Python — via pyright, spoken over actual LSP (JSON-RPC 2.0,
Content-Length framed both ways): needs any one of a project-local npm
install (
node_modules/pyright/langserver.index.js, plusnode), apyright-langserverinside a local.venv/venv(pip install pyrightships one), orpyright-langserveron PATH.
No configuration; any missing piece means that provider silently never runs and the static graph is served as-is.
After each index build, a background pass (never blocking indexing or
queries, ≤30 s, ≤2000 sites, ambiguous sites first) sends definition
requests at each uncertain call name. Where the server answers, the edge is
rewritten: confirmed or re-pointed callee, ambiguous_with cleared, and
confidence: "lsp" in tool output — strictly stronger than the static
high. Definitions the server places in node_modules expose false internal
edges (e.g. joi's .validate() credited to an in-repo validate) and are
rewritten to external package calls. Answers are cached per file content hash
(.gigagraph/lsp.bin), so re-enrichment after edits only re-queries changed
files; the enriched index is persisted and stamped (lsp_enriched) so a
restart does not redo the work.
Honest limits: the server only answers where the type system does —
untyped JS (e.g. mongoose models without typings) mostly yields no
definition, and those edges keep their static heuristic label; pyright
answers on unannotated Python only as far as its inference reaches, and
definitions it places out of tree (site-packages, typeshed stubs) settle
nothing. String-based correlation (endpoints, DI containers, CDK, RN bridge)
stays fully static. The LspProvider trait is designed for more servers;
TypeScript 7's native compiler no longer ships tsserver.js and is currently
skipped by detection.
Установка Gigagraph
У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.
▸ github.com/Raptosaur/gigagraphFAQ
Gigagraph MCP бесплатный?
Да, Gigagraph MCP бесплатный — установка в пару кликов через Unyly без оплаты.
Нужен ли API-ключ для Gigagraph?
Нет, Gigagraph работает без API-ключей и переменных окружения.
Gigagraph — hosted или self-hosted?
Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.
Как установить Gigagraph в Claude Desktop, Claude Code или Cursor?
Открой Gigagraph на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.
Похожие MCP
GitHub
PRs, issues, code search, CI status
автор: GitHubFilesystem
Secure file operations with configurable access controls.
Memory
Knowledge graph-based persistent memory system.
Template MCP Server
A CLI tool to create a new Model Context Protocol server project with TypeScript support, dual transport options, and an extensible structure
автор: mcpdotdirectAmap Maps Mcp Server
MCP server for using the AMap Maps API
автор: duxiaohuiSupabase
Database, auth and storage
автор: SupabaseEverything
Reference / test server with prompts, resources, and tools.
Git
Tools to read, search, and manipulate Git repositories.
Sequential Thinking
Dynamic and reflective problem-solving through thought sequences.
Time
Time and timezone conversion capabilities.
Compare Gigagraph with
Не уверен что выбрать?
Найди свой стек за 60 секунд
Автор?
Embed-бейдж для README
Похожее
Все в категории development
