Evalbench
БесплатноНе проверенOffline LLM / agent eval harness with regression gates
Описание
Offline LLM / agent eval harness with regression gates
README
EVALBENCH
Offline LLM / agent eval harness with regression gates
PyPI CI License: COCL 1.0 Suite
AI Agents & LLMOps — build, route, evaluate, and secure agents.
pip install cognis-evalbench
evalbench scan . # → prioritized findings in seconds
🔎 Example output
Real, reproducible output from the tool — runs offline:
$ evalbench-emit --version
evalbench 2.0.0
$ evalbench-emit --help
usage: evalbench [-h] [--version] {run,demo,gate,types} ...
evalbench — offline eval harness with a regression gate.
positional arguments:
{run,demo,gate,types}
run evaluate a suite (JSON)
demo run the bundled golden-set suite
gate compare candidate vs baseline run
types list supported assertion types
options:
-h, --help show this help message and exit
--version show program's version number and exit
$ evalbench-emit demo
evalbench v2.0.0 — eval run: support-bot-golden-set
CASE SCORE RESULT
------------------------------------------------
refund-policy 1.000 PASS
[ok] icontains found '30 days'
[ok] regex /support@\S+\.\w+/ matched
[ok] not-contains "I don't know" absent
[ok] word-count words=20 in [8,60]
[ok] latency latency=240ms (<= 800.0)
json-extraction 1.000 PASS
[ok] json-valid valid JSON
[ok] json-schema schema valid
[ok] json-path status='shipped'
summary-similarity 0.905 PASS
[ok] similarity cosine=0.714 (>= 0.55)
[ok] contains found 'Cancel Subscription'
[ok] length len=85 in [20,200]
tone-guardrail 1.000 PASS
[ok] all-of all-of(regex:P, icontains:P)
[ok] not-regex /(?i)stupid|idiot|whatever/ absent
------------------------------------------------
cases: 4/4 passed pass_rate=1.000 mean_score=0.976
Blocks above are real
evalbenchoutput — reproduce them from a clone.
Usage — step by step
Install the harness:
pip install cognis-evalbenchTry the bundled golden set (no files needed), or run your own suite JSON and save the result for later gating:
evalbench demo evalbench run suite.json --save baseline.jsonGate a candidate against a baseline.
gateaccepts saved run results or raw suites and flags score drops beyond--tolerance(add--strictfor pass-rate / mean-score regressions):evalbench gate baseline.json candidate.json --tolerance 0.02 --format jsonRead the result.
evalbench typeslists the assertion types (contains, regex, json-schema, similarity, latency, cost, and more). Exit0= success,1= case failure / regression,2= usage/IO error.Automate in CI. Evaluate then gate so regressions fail the build:
evalbench run suite.json --save run.json && evalbench gate baseline.json run.json
Contents
- Why evalbench? · Features · Quick start · Example · Architecture · AI stack · How it compares · Integrations · Install anywhere · Related · Contributing
Why evalbench?
CI for agents
evalbench is single-purpose, scriptable, and self-hostable: point it at a target, get prioritized results in the format your workflow already speaks (table · JSON · SARIF), gate CI on it, and let agents drive it over MCP.
Features
- ✅ Load Suite
- ✅ Run Suite
- ✅ Compare Baseline
- ✅ Runs on Linux/macOS/Windows · Docker · devcontainer
- ✅ Ports in Python, JavaScript, Go, and Rust (
ports/)
Quick start
pip install cognis-evalbench
evalbench --version
evalbench scan . # scan current project
evalbench scan . --format json # machine-readable
evalbench scan . --fail-on high # CI gate (non-zero exit)
Example
$ evalbench scan .
[HIGH ] EVA-001 example finding (./src/app.py)
[MEDIUM ] EVA-002 another signal (./config.yaml)
2 findings · risk score 5 · 38ms
Architecture
flowchart LR
IN[agent / A2A traffic] --> P[evalbench<br/>map + analyze]
P --> OUT[graph + flags]
Use it from any AI stack
evalbench is interoperable with every popular way of using AI:
- MCP server —
evalbench mcp(Claude Desktop, Cursor, Cognis.Studio, uncensored-fleet) - OpenAI-compatible / JSON — pipe
evalbench scan . --format jsoninto any agent or LLM - LangChain · CrewAI · AutoGen · LlamaIndex — wrap the CLI/JSON as a tool in one line
- CI / scripts — exit codes + SARIF for non-AI pipelines
How it compares
| Cognis evalbench | promptfoo | |
|---|---|---|
| Self-hostable, no account | ✅ | varies |
| Single command, zero config | ✅ | ⚠️ |
| JSON + SARIF for CI | ✅ | varies |
| MCP-native (AI agents) | ✅ | ❌ |
| Polyglot ports (JS/Go/Rust) | ✅ | ❌ |
| Open license | ✅ COCL | varies |
Built in the spirit of promptfoo / deepeval, re-framed the Cognis way. Missing a credit? Open a PR.
Integrations
Pipes into your stack: SARIF for code-scanning, JSON for anything, an MCP server (evalbench mcp) for AI agents, and a webhook forwarder for SIEM/Slack/Jira. See docs/INTEGRATIONS.md.
Install — every way, every platform
pip install "git+https://github.com/cognis-digital/evalbench.git" # pip (works today)
pipx install "git+https://github.com/cognis-digital/evalbench.git" # isolated CLI
uv tool install "git+https://github.com/cognis-digital/evalbench.git" # uv
pip install cognis-evalbench # PyPI (when published)
docker run --rm ghcr.io/cognis-digital/evalbench:latest --help # Docker
brew install cognis-digital/tap/evalbench # Homebrew tap
curl -fsSL https://raw.githubusercontent.com/cognis-digital/evalbench/main/install.sh | sh
| Linux | macOS | Windows | Docker | Cloud |
|---|---|---|---|---|
scripts/setup-linux.sh |
scripts/setup-macos.sh |
scripts/setup-windows.ps1 |
docker run ghcr.io/cognis-digital/evalbench |
DEPLOY.md (AWS/Azure/GCP/k8s) |
Related Cognis tools
- agentsmith — Config-first scaffolding and orchestration for multi-agent workflows
- skillhub — Local skill registry and installer for AI agents
- toolguard — Runtime allowlist and policy for agent tool-calls
- ragkit — Batteries-included local RAG pipeline — ingest, index, serve
- memorybank — Portable long-term memory store for agents, exposed over MCP
- promptpack — Versioned prompt / template registry with A/B and rollbacks
Explore the suite → 🗂️ all 170+ tools · ⭐ awesome-cognis · 🔗 cognis-sources · 🤖 uncensored-fleet · 🧠 engram
Contributing
PRs, new rules, and demo scenarios are welcome under the collaboration-pull model — see CONTRIBUTING.md and SECURITY.md.
⭐ If
evalbenchsaved you time, star it — it genuinely helps others find it.
Interoperability
{} composes with the 300+ tool Cognis suite — JSON in/out and a shared
OpenAI-compatible /v1 backbone. See INTEROP.md for the
suite map, composition patterns, and reference stacks.
License
Source-available under the Cognis Open Collaboration License (COCL) v1.0 — free for personal, internal-evaluation, research, and educational use; commercial / production use requires a license ([email protected]). See LICENSE.
Установить Evalbench в Claude Desktop, Claude Code, Cursor
unyly install evalbenchСтавит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.
Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh
Или настроить вручную
Выполни в терминале:
claude mcp add evalbench -- uvx --from git+https://github.com/cognis-digital/evalbench cognis-evalbenchПошаговые гайды: как установить Evalbench
FAQ
Evalbench MCP бесплатный?
Да, Evalbench MCP бесплатный — установка в пару кликов через Unyly без оплаты.
Нужен ли API-ключ для Evalbench?
Нет, Evalbench работает без API-ключей и переменных окружения.
Evalbench — hosted или self-hosted?
Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.
Как установить Evalbench в Claude Desktop, Claude Code или Cursor?
Открой Evalbench на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.
Похожие MCP
Fetch
Web content fetching and conversion for efficient LLM usage.
AWS KB Retrieval
Retrieval from AWS Knowledge Base using Bedrock Agent Runtime.
автор: modelcontextprotocolSpring AI MCP Server
Provides auto-configuration for setting up an MCP server in Spring Boot applications.
llm-analysis-assistant
A very streamlined mcp client that supports calling and monitoring stdio/sse/streamableHttp, and can also view request responses through the /logs page. It also
автор: xuzexin-hzMCP-Agent
A simple, composable framework to build agents using Model Context Protocol by [LastMile AI](https://www.lastmileai.dev)
автор: lastmile-aiSpring AI MCP Client
Provides auto-configuration for MCP client functionality in Spring Boot applications.
mcp.natoma.ai
A Hosted MCP Platform to discover, install, manage and deploy MCP servers by [Natoma Labs](https://www.natoma.ai)
MCPHub
Website to list high quality MCP servers and reviews by real users. Also provide online chatbot for popular LLM models with MCP server support.
MCP Servers Rating and User Reviews
Website to rate MCP servers, write authentic user reviews, and [search engine for agent & mcp](http://www.deepnlp.org/search/agent)
mkinf
An Open Source registry of hosted MCP Servers to accelerate AI agent workflows.
Compare Evalbench with
Не уверен что выбрать?
Найди свой стек за 60 секунд
Автор?
Embed-бейдж для README
Похожее
Все в категории ai
