Test Bench
БесплатноНе проверенOpen-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctn
Описание
Open-source evaluation framework for MCP servers powered by Claude. Auto-discovers tools, generates test scenarios, runs LLM-as-judge scoring to assess correctness and safety, audits for security issues, and visualizes traces in a web dashboard.
README
Video walkthrough: https://youtu.be/rGaS_9dsOao 60-second overview: https://youtu.be/eqE1Nj6BDIk
Open-source evaluation harness for MCP servers: auto-discovers tools, runs LLM-judged scenarios, flags security issues, and shows traces in a web dashboard.

What it is
MCP Test Bench is a local evaluation harness for Model Context Protocol servers. Point it at any MCP server — a stdio command or an SSE URL — and it discovers every tool, resource, and prompt the server exposes, then auto-generates realistic test scenarios from the schemas using Claude. It drives those scenarios through an agent loop, records every tool call and response, scores each run with an LLM-as-judge against configurable rubrics (correctness, safety, efficiency, hallucination), and audits the server for security issues such as prompt-injection patterns in tool descriptions, unbounded outputs, and PII leakage.
Everything is stored in a local SQLite file and visible in a Next.js dashboard: run history, score timelines, side-by-side server comparisons, and full trace views. A mcpbench CLI makes the whole pipeline scriptable for CI.
Quickstart
git clone https://github.com/RitikPatill/mcp-test-bench.git
cd mcp-test-bench
pnpm i
export ANTHROPIC_API_KEY=sk-ant-...
pnpm dev
# open http://localhost:3000
Requires Node >= 20 and pnpm 9.
Usage
Add a server through the dashboard by pasting a stdio command (e.g. npx -y @modelcontextprotocol/server-filesystem /tmp). Test Bench discovers its tools, auto-generates scenarios, and lets you kick off an eval with one click. The trace timeline streams live as Claude calls tools and the judge scores each turn. When it finishes, the report shows an overall score, per-rubric breakdowns, failed scenarios with reasoning, and any security findings.
For CI, build the CLI and run a config file directly:
pnpm --filter cli build
node apps/cli/dist/index.js run examples/demo/filesystem-server.yaml
Config files are plain YAML — see examples/demo/filesystem-server.yaml for the minimal shape.
Architecture
flowchart LR
UI[Next.js Dashboard] -->|REST/SSE| API[API Routes]
API --> Runner[Eval Runner]
API --> DB[(SQLite)]
Runner --> MCPClient[MCP Client\nstdio + SSE]
Runner --> Agent[Claude Agent Loop]
Runner --> Judge[LLM-as-Judge]
Runner --> Scanner[Security Scanner]
MCPClient -.spawns.-> Target[Target MCP Server]
Agent --> Anthropic[Anthropic API]
Judge --> Anthropic
Project structure
mcp-test-bench/
├── apps/
│ ├── cli/ # mcpbench binary (tsup build → dist/index.js)
│ └── web/ # Next.js 15 App Router dashboard
├── packages/
│ └── core/ # MCP client, eval runner, judge, scanner, DB schema
├── examples/
│ ├── ci/ # GitHub Actions workflow + example config
│ └── demo/ # Ready-to-run YAML configs for public MCP servers
├── docs/ # Architecture doc, roadmap, demo gif
└── scripts/ # DB seed and demo helper scripts
Roadmap
- Custom judge models — swap the judge to any OpenAI-compatible endpoint via a
judge.modelconfig key. - Plugin scanners — a
ScannerPlugininterface so community security checks can ship as npm packages. - Hosted mode — optional Turso/libsql backend so teams can share results across machines.
- Replay mode — re-score already-saved turns without re-running the agent when rubrics change.
- Scenario library — shareable community YAML packs, one per common server type, importable via
mcpbench import.
License
MIT — see LICENSE.
Built autonomously by autodev, a multi-agent orchestrator I designed. Each commit in this repo was authored by me; the implementation work was performed by Sonnet under the orchestrator's control. Read the orchestrator's README to see how.
Установка Test Bench
У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.
▸ github.com/RitikPatill/mcp-test-benchFAQ
Test Bench MCP бесплатный?
Да, Test Bench MCP бесплатный — установка в пару кликов через Unyly без оплаты.
Нужен ли API-ключ для Test Bench?
Нет, Test Bench работает без API-ключей и переменных окружения.
Test Bench — hosted или self-hosted?
Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.
Как установить Test Bench в Claude Desktop, Claude Code или Cursor?
Открой Test Bench на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.
Похожие MCP
Fetch
Web content fetching and conversion for efficient LLM usage.
AWS KB Retrieval
Retrieval from AWS Knowledge Base using Bedrock Agent Runtime.
автор: modelcontextprotocolSpring AI MCP Server
Provides auto-configuration for setting up an MCP server in Spring Boot applications.
llm-analysis-assistant
A very streamlined mcp client that supports calling and monitoring stdio/sse/streamableHttp, and can also view request responses through the /logs page. It also
автор: xuzexin-hzMCP-Agent
A simple, composable framework to build agents using Model Context Protocol by [LastMile AI](https://www.lastmileai.dev)
автор: lastmile-aiSpring AI MCP Client
Provides auto-configuration for MCP client functionality in Spring Boot applications.
mcp.natoma.ai
A Hosted MCP Platform to discover, install, manage and deploy MCP servers by [Natoma Labs](https://www.natoma.ai)
MCPHub
Website to list high quality MCP servers and reviews by real users. Also provide online chatbot for popular LLM models with MCP server support.
MCP Servers Rating and User Reviews
Website to rate MCP servers, write authentic user reviews, and [search engine for agent & mcp](http://www.deepnlp.org/search/agent)
mkinf
An Open Source registry of hosted MCP Servers to accelerate AI agent workflows.
Compare Test Bench with
Не уверен что выбрать?
Найди свой стек за 60 секунд
Автор?
Embed-бейдж для README
Похожее
Все в категории ai
