Unity Catalog
FreeNot checkedSemantic search over Databricks Unity Catalog metadata using BGE-large embeddings and pgvector.
About
Semantic search over Databricks Unity Catalog metadata using BGE-large embeddings and pgvector.
README
Semantic catalog discovery MCP server for Databricks Unity Catalog.
Deployed as a Databricks App, backed by Lakebase (pgvector) indexed from UC system tables. Fronted by uc-mcp-proxy for MCP client connectivity.
No SQL execution. Agents use spark-connect-mcp for that.
Architecture
UC system tables Lakebase (pgvector)
system.information_schema.tables → catalog_metadata
system.information_schema.columns (full_name PK, comment, columns JSONB,
content_hash TEXT, embedding vector(1024),
synced_at TIMESTAMPTZ)
Sync Job (Databricks Job, every 6h)
1. Read system tables for allowed catalogs/schemas
2. Compute content_hash = SHA-256(table_comment + all column names/types/comments)
3. Compare vs stored hashes in Lakebase
4. Call Databricks FM API (BGE-large) ONLY for new/changed tables
5. Upsert changed rows, delete removed tables
FastAPI App (Databricks App)
/mcp ← uc-mcp-proxy routes here
Tools: search (pgvector ANN), describe (Lakebase SELECT), list (Lakebase SELECT)
lineage (direct Databricks API passthrough)
MCP Tools
| Tool | Source | Description |
|---|---|---|
search_tables(query) |
Lakebase pgvector | Semantic search over table+column descriptions |
describe_table(full_name) |
Lakebase | Full schema: columns, types, comments |
list_catalogs() |
Lakebase | All indexed catalogs |
list_schemas(catalog) |
Lakebase | Schemas within a catalog |
get_table_lineage(full_name) |
Databricks API | Upstream/downstream tables |
get_column_lineage(full_name, column) |
Databricks API | Column-level provenance |
Requirements
- Databricks workspace with
system.information_schema.*enabled - Lakebase (provisioned via
make deploy) - Databricks App service principal with UC metastore access
- uc-mcp-proxy for MCP client routing
Deploy
# Configure allowlist in databricks.yml, then:
make deploy
Single target provisions the App, Lakebase, runs migrations, and triggers the initial sync job.
Configuration
Operator specifies which catalogs (or catalog+schema combinations) to index in databricks.yml:
variables:
catalog_allowlist:
default: |
- catalog: main
- catalog: analytics
schema_pattern: "prod_*"
Only namespaces in the allowlist are indexed. No "index everything" default.
Embedding Strategy
- Model: Databricks Foundation Models API (BGE-large, 1024 dimensions)
- Content:
{full_name}: {table_comment}. Columns: {col} ({type}): {col_comment}, ... - Granularity: one vector per table (column context included, not per-column)
- Index: HNSW in pgvector
Hash-based incremental ETL — stable workspaces skip 90%+ of embedding API calls.
Installing Unity Catalog
This server has no published package — it is built from source. Open the repository and follow its README.
▸ github.com/icerhymers/uc-catalog-mcpFAQ
Is Unity Catalog MCP free?
Yes, Unity Catalog MCP is free — one-click install via Unyly at no cost.
Does Unity Catalog need an API key?
No, Unity Catalog runs without API keys or environment variables.
Is Unity Catalog hosted or self-hosted?
Self-hosted: the server runs locally on your machine via the install command above.
How do I install Unity Catalog in Claude Desktop, Claude Code or Cursor?
Open Unity Catalog on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.
Related MCPs
GitHub
PRs, issues, code search, CI status
by GitHubFilesystem
Secure file operations with configurable access controls.
Memory
Knowledge graph-based persistent memory system.
Template MCP Server
A CLI tool to create a new Model Context Protocol server project with TypeScript support, dual transport options, and an extensible structure
by mcpdotdirectCompare Unity Catalog with
Not sure what to pick?
Find your stack in 60 seconds
Author?
Embed badge for your README
Browse similar
All development MCPs
