Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Context Cortex

БесплатноНе проверен

MCP server with 26 tools and 3 resources for knowledge management, research, and scheduling via Claude Desktop.

GitHubEmbed

Описание

MCP server with 26 tools and 3 resources for knowledge management, research, and scheduling via Claude Desktop.

README

Context Cortex is a domain-agnostic personal research operating system: continuously ingest documents and web findings, index them in LanceDB (hybrid vector + full-text search with Voyage AI reranking), maintain a git-backed markdown memory, and work from Claude Desktop via the Model Context Protocol (MCP) or from the bundled Express server (RSS, scheduled research, HTTP ingest, legacy UI).

It evolved from a ResearchBrain-style codebase (transfer-pricing–focused upstream) with practice-specific defaults replaced by files under memory/config/.


Table of contents

  1. Features at a glance
  2. Quick start
  3. Architecture
  4. RAG pipeline
  5. Memory layer
  6. Knowledge base layer
  7. MCP server (mcp-server.js)
  8. Express server (server.js)
  9. Scheduled jobs (crons)
  10. Configuration
  11. Environment variables
  12. npm scripts
  13. OpenClaw and external ingest
  14. Observability
  15. Deployment and ops
  16. Further reading

Features at a glance

Area What you get
Retrieval Hybrid vector + FTS on LanceDB, optional multi-query expansion (Claude), Voyage rerank, session context, filters (topic, jurisdiction, concept_path, technique tags, tier, doc status, recency).
Ingestion Raw text, PDF (pdf-parse or optional Docling), DOCX / MD / TXT, RSS with paywall scraping (Puppeteer), Perplexity (manual + scheduled tiers), OpenClaw Gmail/Telegram PDF endpoints.
Graph metadata At ingest, Claude can fill entity_refs, regulation_refs, tp_method_tags (legacy column name; use as technique/facet tags), concept_path validated against your taxonomy.json.
Memory Markdown under memory/ (topics, briefs, articles, sources, synthesis, config); git sync to GitHub on writes; pull on MCP/Express startup and every 30 minutes while Express runs.
Curated KB Concept articles in knowledge-base/concepts/, outputs in knowledge-base/outputs/, registry (MASTER_INDEX, ACTIVITY_LOG, LINT_REPORT), compile, lint, file_output loop.
MCP 26 tools + 3 resources for Claude Desktop (search, ingest, Perplexity, agenda, briefs, compiler, linter, citations, DOCX, weekly topic scoring, observability, chunk deletion).
Express REST API for chat (orchestrator), clip/research, briefs, upload, RAG retrieve/verify/search/dashboard, authenticated ingest for OpenClaw; node-cron jobs (RSS, briefs, synthesis, maintenance, scheduled research).

Quick start

  1. cp .env.example .env and set API keys. Use CONTEXT_CORTEX_API_KEY for Bearer-protected ingest routes; RESEARCHBRAIN_API_KEY is still accepted as an alias.
  2. Edit memory/config/taxonomy.json, research-agenda.md, and corpus-profile.md for your domain.
  3. Optionally fill perplexity-queries.json so scheduled tiers (current_awareness, thematic, horizon) run non-empty queries.
  4. npm install, then npm run validate-env (requires live keys) or npm run dev to start Express on port 3000 (or PORT).
  5. Point Claude Desktop MCP at mcp-server.js (see MCP server).

Embeddings: .env.example defaults to voyage-finance-2. For non-finance corpora, set EMBEDDING_MODEL (and matching EMBEDDING_DIMENSIONS) to a general-purpose Voyage model per Voyage AI docs.


Architecture

┌─────────────────────────────────────────────────────────────┐
│              Claude Desktop (or other MCP client)            │
│                     MCP stdio transport                      │
└──────────────────────────┬──────────────────────────────────┘
                           │
              ┌────────────▼────────────┐
              │      mcp-server.js       │
              │   26 tools · 3 resources │
              └────────────┬────────────┘
                           │
   ┌───────────┬───────────┼───────────┬───────────┐
   │           │           │           │           │
┌──▼───┐ ┌────▼────┐ ┌────▼────┐ ┌────▼────┐ ┌────▼─────┐
│ RAG  │ │ Memory  │ │Perplexity│ │   KB    │ │ Output   │
│Lance │ │  + git  │ │  Sonar   │ │compiler │ │  loop    │
└──┬───┘ └─────────┘ └──────────┘ └─────────┘ └──────────┘
   │
┌──▼──────────────────────────────────────────────────────┐
│              Express server.js (:3000)                   │
│  RSS · crons · OpenClaw ingest · RAG HTTP · dashboard     │
└──────────────────────────────────────────────────────────┘

RAG pipeline

Primary code lives under src/rag/.

Component Role
retriever.js Main retrieval entry: hybrid search + rerank; optional query expansion (QUERY_EXPANSION_ENABLED, retrievalPipeline.js). Logs retrieval events for observability.
vectorStore.js LanceDB Cloud: upserts, filters, Phase 2 schema toggles (LANCEDB_PHASE2_SCHEMA), graph column probing.
embedder.js Voyage embeddings; disk cache (EMBEDDING_CACHE_PATH).
reranker.js Voyage rerank-2.5 (or VOYAGE_RERANK_MODEL) second-stage ranking.
chunker.js Sentence/chunk boundaries (compromise NLP), size limits and overlap for non-Docling text.
hierarchicalChunker.js Docling path: section breadcrumbs, section_path, merge/split without overlap.
ingestionPipeline.js Chunk → embed → graph metadata → upsert → markdown sidecar → git sync / registry hooks. Entry points: ingest(), ingestPages(), ingestFromParsedDoc().
graph-metadata-extractor.js Claude extracts structured fields; concept_path validated against memory/config/taxonomy.json.
concept-taxonomy.js Loads taxonomy from disk; reloadTaxonomy() if you change JSON at runtime.
validation.js Safe upsert validation; failed tagging can mark status.
maintenance.js TTL expiry (deleteExpired), knowledge-gap analysis → scheduler hints.

PDFs: Set DOCLING_ENABLED=true and install Python deps from ingestion/python/requirements.txt for layout-aware parsing via scripts/docling_parse.py; otherwise pdf-parse. Metadata inference in src/ingest/pdfProcessor.js.

Citations: Drafts can use [CHUNK_ID: claim]; src/output/citationVerifier.js checks against LanceDB (and optionally Claude) — exposed as MCP verify_citations and POST /api/rag/verify-citations.


Memory layer

Implemented in src/memory.js.

Path Purpose
memory/topics/ Topic notes (append via MCP save_memory_note).
memory/articles/ Per-ingestion markdown (RSS, uploads, Perplexity, etc.).
memory/briefs/ Daily intelligence briefs (YYYY-MM-DD.md).
memory/sources/ Source logs (ingested lines per feed).
memory/synthesis/ master-thesis.md, weekly-synthesis.md.
memory/config/ research-agenda.md, taxonomy.json, corpus-profile.md, perplexity-queries.json, etc.
memory/conversations/ Optional conversation logs when using Express chat.

Writes trigger syncToGitHub when GITHUB_TOKEN and GITHUB_REPO are set. MEMORY_PATH can point to an external volume.


Knowledge base layer

Path / module Purpose
knowledge-base/concepts/ Compiled concept articles (markdown + frontmatter).
knowledge-base/outputs/ Filed white papers, alerts, query responses (file_output / output loop).
knowledge-base/registry/ MASTER_INDEX.md, ACTIVITY_LOG.md, LINT_REPORT.md, SCHEMA_CHANGELOG.md.
src/knowledgeCompiler.js discoverConcepts, compileArticle, batch compileAll, staleness, schema suggestions. Uses corpus-profile.md then SOUL.md for tone.
src/knowledgeLinter.js Consistency, staleness, coverage, orphans, citations, cross connections, research questions.
src/knowledgeRegistry.js Registry queries and updates.
src/outputLoop.js Files outputs, summary chunk, re-ingest, stale concept flags.

MCP server (mcp-server.js)

  • Server name: context-cortex
  • Transport: stdio (stdout must remain JSON-RPC only; logging goes to stderr)
  • Startup: optional git pull via src/git-sync.js; optional LanceDB Phase 2 preflight when LANCEDB_PHASE2_SCHEMA=true

Resources (3)

URI Description
contextcortex://system-prompt Session system prompt: corpus profile / SOUL, master thesis, agenda, topic files, concept index, recent ingested lines.
contextcortex://master-thesis Contents of memory/synthesis/master-thesis.md.
contextcortex://agenda Raw memory/config/research-agenda.md.

Tools (26)

Tool Purpose
search_knowledge_base Primary RAG search; filters: topic, jurisdiction, content_type, concept_path_prefix, technique tag (tp_method param maps to tp_method_tags), tier_max, doc_status, days_ago; mode: standard | deep_dive; optional session_id for context + rate_last_retrieval.
search_by_ref Filter-only retrieval by entity_ref, regulation_ref, technique tag, or concept_path_prefix (no embedding/rerank).
rate_last_retrieval Save good/bad feedback for the last retrieval in a session.
ingest_content Ingest arbitrary text into LanceDB + pipeline (metadata, TTL, etc.).
ingest_file Ingest local PDF, DOCX, MD, TXT by path; Docling path when enabled; updates memory topic sidecars.
search_perplexity Live web research via Perplexity Sonar Pro.
get_research_agenda / update_research_agenda Read/update memory/config/research-agenda.md (core topics + search terms).
get_articles_digest List ingested RSS-style articles (optional date/source filters).
get_knowledge_stats Chunk totals, topic distribution, date range.
get_ingestion_dashboard Rich dashboard: periods (24h, 48h, local “today”, week, month, year), distributions (topic, content type, channel, source, tier, doc status).
verify_citations Verify [CHUNK_ID: claim] citations in a draft.
get_daily_brief Load brief by date or generate: true for today (Perplexity + Claude; requires agenda core topics).
save_memory_note Append structured note to memory/topics/<file>.md.
get_master_thesis Load master thesis markdown.
compile_concept Build/rebuild a concept article from chunks (force_recompile, scope: full | incremental).
discover_concepts Suggest new concept slugs from chunk density (optional domain_filter on taxonomy L1).
lint_knowledge_base KB health report; preset: quick | full; optional checks array; output_format: summary | full_report | actionable_only.
file_output File a finished artifact into knowledge-base/outputs/ and re-ingest summary (output_type enum includes white_paper, client_alert, query_response, etc.).
knowledge_registry query_type: summary, domain_coverage, recent_additions, stale_articles, source_list, activity_log.
convert_to_docx Markdown → Word via Python python-docx (src/weekly-topics/convert_to_docx.py).
score_weekly_topics Deterministic selection from scored candidates (recency, impact, rigor, novelty, whitespace + diversity / repeat rules).
get_observability_report JSON summary over last N days (retrieval, ingestion, MCP usage, tokens when logged).
get_low_score_queries Recent queries with low reranker scores.
get_recent_alerts Threshold alerts from cc-alerts.jsonl.
delete_chunks LanceDB SQL predicate deletion; confirm: true required after dry-run.

Example MCP config (paths and keys are yours to fill):

{
  "mcpServers": {
    "context-cortex": {
      "command": "node",
      "args": ["/absolute/path/to/context-cortex/mcp-server.js"],
      "env": {
        "ANTHROPIC_API_KEY": "...",
        "VOYAGE_API_KEY": "...",
        "LANCEDB_URI": "...",
        "LANCEDB_API_KEY": "...",
        "PERPLEXITY_API_KEY": "...",
        "GITHUB_TOKEN": "...",
        "CONTEXT_CORTEX_API_KEY": "...",
        "MEMORY_PATH": "./memory"
      }
    }
  }
}

Note: npm run mcp does not run RSS or Express crons — use npm start / npm run dev for background automation, or run npm run ingest-rss-once manually.


Express server (server.js)

Default port 3000 (PORT). Selected routes:

Health and status

Method Path Notes
GET /api/health Liveness.
GET /api/status Broader status payload.
GET /api/rag/health RAG / LanceDB oriented check.

Chat (orchestrator)

Method Path Notes
GET /api/chat/start Start session.
POST /api/chat User message → Claude with memory/RAG context.
POST /api/chat/end End session.
GET /api/session Session snapshot.

Agenda, sources, clip, research, briefs

Method Path Notes
GET/PUT /api/agenda Read/update research agenda.
GET /api/sources/status, /api/sources/verify Feed / auth status.
POST /api/sources/authenticate Cookie / premium flow helpers.
POST /api/clip, GET /api/clip/detect, POST /api/clip/paste URL clipping pipeline.
POST /api/research/search, /api/research/save Research helpers.
GET /api/briefs, /api/briefs/:date List / load briefs.
POST /api/briefs/generate Trigger daily brief.

Ingestion and uploads

Method Path Notes
POST /api/ingest Generic ingest hook.
GET /api/ingest/stats Ingest stats.
GET /api/reprocess/candidates, POST /api/reprocess Reprocess flow.
GET /api/articles Article listing.
POST /api/upload Raw upload (size limit UPLOAD_MAX_SIZE_MB).
GET /api/upload/formats Supported MIME/extensions.
GET /api/debug/uploads Debug helper.

RAG HTTP API

Method Path Auth Notes
POST /api/rag/ingest Bearer CONTEXT_CORTEX_API_KEY or RESEARCHBRAIN_API_KEY JSON ingest for agents (e.g. OpenClaw writer).
POST /api/rag/ingest/pdf Bearer Base64 PDF + metadata.
POST /api/rag/ingest/gmail-pdf Bearer OpenClaw Gmail skill shape.
POST /api/rag/ingest/telegram-pdf Bearer OpenClaw Telegram skill; optional async queue (RAG_INGEST_QUEUE_ENABLED).
GET /api/rag/stats Vector stats.
GET /api/rag/dashboard Ingestion dashboard JSON (timezone DASHBOARD_TIMEZONE, scan caps in .env.example).
GET /api/rag/search Query params for hybrid search.
POST /api/rag/retrieve JSON body { query, filters, topK, topN } for tools like OpenClaw.
POST /api/rag/verify-citations JSON { draft_text }; optional Telegram alert on flags.

Static UI

  • public/index.html — legacy web UI.
  • public/dashboard.html — ingestion dashboard.

Scheduled jobs (crons)

All schedules run only while the Express process is running (npm start / npm run dev).

Schedule Job
RSS_POLL_INTERVAL 6h0 */6 * * *; any other value → hourly ingestFeeds() — RSS + gates (agenda, dedupe, age, optional SOURCES_ENABLED).
Daily 7:00 generateDailyBrief() — Perplexity sweep from agenda topics → Claude brief → memory/briefs/.
Sunday 8:00 Weekly synthesis appended to memory/synthesis/weekly-synthesis.md (prompt is in server.jscustomize for your domain).
1st of month 4:00 Full knowledge base lintknowledge-base/registry/LINT_REPORT.md.
Sunday 9:00 compileAll() — recompile stale concept articles.
SCHEDULE.current_awareness (default 0 */4 * * *) Run next current_awareness query from perplexity-queries.json (skips if empty).
SCHEDULE.thematic (default Monday 2:00) Thematic tier.
SCHEDULE.horizon (default 1st of month 3:00) Horizon tier.
Daily 3:00 RAG maintenance: TTL cleanup + gap analysis (maintenance.js).
Monday 6:00 Court monitor — only if COURT_MONITOR_ENABLED=true (ingest path still uses tax-case defaults in code; adjust if you reuse this job).
Monday 7:00 OECD monitor — only if OECD_MONITOR_ENABLED=true (same note).
Every 30 minutes git pull to sync memory from GitHub.

Weekly CLI pipeline (not Express): scripts/weekly-trigger.sh + skill skills/weekly-topics-whitepaper/SKILL.md.


Configuration

File Role
memory/config/README.md Index of config files.
memory/config/taxonomy.json Allowed concept_path L1/L2 prefixes for graph metadata + discovery.
memory/config/research-agenda.md Core topics + search terms for RSS gating and daily brief Perplexity queries.
memory/config/corpus-profile.md Voice, audience, citation norms for compiler and linter (preferred over SOUL.md).
memory/config/perplexity-queries.json Arrays: current_awareness, thematic, horizon — each item { id, query, metadata }.
memory/config/perplexity-queries.example.json Copy/paste starter shape.
src/config/sources.js RSS feed rows + env-built premium feeds.
src/config/premiumRssUrls.js Default URL lists (empty by default; set env vars to enable).

LanceDB / taxonomy: If you point at an existing table built under another taxonomy, old concept_path values will not automatically match a new taxonomy.json. Prefer a new table/project, or run scripts/backfill-graph-metadata.js after changing taxonomy.


Environment variables

Copy .env.example to .env. Highlights:

Category Variables (non-exhaustive)
LLM / search ANTHROPIC_API_KEY, PERPLEXITY_API_KEY, optional Perplexity Agent flags.
Embeddings / rerank VOYAGE_API_KEY, EMBEDDING_MODEL, EMBEDDING_DIMENSIONS, VOYAGE_RERANK_MODEL, batching / delay.
LanceDB LANCEDB_URI, LANCEDB_API_KEY, LANCEDB_PHASE2_SCHEMA, LANCEDB_UPSERT_BATCH_SIZE, PHASE2_PREFLIGHT_STRICT.
Auth CONTEXT_CORTEX_API_KEY (preferred), RESEARCHBRAIN_API_KEY (legacy).
Git GITHUB_TOKEN, GITHUB_REPO, GITHUB_BRANCH.
Paths MEMORY_PATH, COOKIE_DIR, EMBEDDING_CACHE_PATH.
RSS RSS_POLL_INTERVAL, SOURCES_ENABLED, RSS_MAX_PER_FEED, RSS_MAX_AGE_DAYS, publisher URL env vars (see .env.example).
PDF DOCLING_ENABLED, UPLOAD_MAX_SIZE_MB.
Retrieval QUERY_EXPANSION_ENABLED, HYDE_ENABLED, DEEP_DIVE_MAX_CHUNKS, graph flags.
OpenClaw / skills GMAIL_PDF_SKILL_ENABLED, TELEGRAM_PDF_SKILL_ENABLED, MY_EMAIL, TRUSTED_GMAIL_SENDERS, MY_TELEGRAM_USER_ID, queue flag.
Dashboard DASHBOARD_TIMEZONE, DASHBOARD_MAX_SCAN_ROWS, etc.
Optional monitors COURT_MONITOR_ENABLED, OECD_MONITOR_ENABLED.

npm scripts

Script Command Purpose
start node server.js Production Express.
dev node --watch server.js Dev server with reload.
mcp node mcp-server.js MCP stdio only.
validate-env node scripts/validateEnv.js Keys + API smoke tests.
validate-sources node scripts/validate-sources.js HTTP check all RSS URLs.
validate-premium-feeds node scripts/validate-premium-feeds.js Economist/WSJ URL lists.
ingest-rss-once node scripts/ingest-rss-once.js One-shot RSS pass.
migrate-memory node scripts/migrateMemory.js Bulk index memory/articles → LanceDB.
test-rag node scripts/testRag.js RAG regression tests (needs LanceDB + Voyage env).
test-mcp node scripts/testMcp.js MCP registry + tool smoke (needs full env).
schema-guard compound Phase 2 columns + graph preflight + Lance check + test-rag.
import-cookies node scripts/import-cookies.js Premium feed cookies.

OpenClaw and external ingest

  • Skill YAMLs: openclaw/skills/ (gmail-pdf-ingest, telegram-pdf-ingest, context-cortex-writer, retrieval-feedback).
  • Install symlinks: ./scripts/install-openclaw-skills.sh
  • Set CONTEXT_CORTEX_API_KEY (or legacy key) and your Express base URL in OpenClaw.

Observability

  • JSONL logs (gitignored): logs/cc-observability.jsonl, logs/cc-alerts.jsonl
  • MCP tools: get_observability_report, get_low_score_queries, get_recent_alerts
  • Optional: scripts/dead-chunk-detector.js (compare Lance doc_ids to retrieval logs)

Deployment and ops

Full Mac mini / volume notes: README-deployment.md.


Further reading

Document Contents
README-workflow.md Day-to-day usage, MCP patterns, weekly topics pipeline.
README-deployment.md launchd, logs, OpenClaw on a server, Railway env checklist.
CLAUDE.md Repo map for Claude Code / contributors.
memory/config/README.md First-edit checklist for new domains.

License

See Polyform Noncommerical 1.0.0. (LICENSE).

from github.com/gdavidreeves/context-cortex

Установка Context Cortex

У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.

▸ github.com/gdavidreeves/context-cortex

FAQ

Context Cortex MCP бесплатный?

Да, Context Cortex MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Context Cortex?

Нет, Context Cortex работает без API-ключей и переменных окружения.

Context Cortex — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Context Cortex в Claude Desktop, Claude Code или Cursor?

Открой Context Cortex на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Context Cortex with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории ai