Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Evalbench

БесплатноНе проверен

Offline LLM / agent eval harness with regression gates

GitHubEmbed

Описание

Offline LLM / agent eval harness with regression gates

README

EVALBENCH

EVALBENCH

Offline LLM / agent eval harness with regression gates

PyPI CI License: COCL 1.0 Suite

AI Agents & LLMOps — build, route, evaluate, and secure agents.

pip install cognis-evalbench
evalbench scan .            # → prioritized findings in seconds

🔎 Example output

Real, reproducible output from the tool — runs offline:

$ evalbench-emit --version
evalbench 2.0.0
$ evalbench-emit --help
usage: evalbench [-h] [--version] {run,demo,gate,types} ...

evalbench — offline eval harness with a regression gate.

positional arguments:
  {run,demo,gate,types}
    run                 evaluate a suite (JSON)
    demo                run the bundled golden-set suite
    gate                compare candidate vs baseline run
    types               list supported assertion types

options:
  -h, --help            show this help message and exit
  --version             show program's version number and exit
$ evalbench-emit demo
evalbench v2.0.0 — eval run: support-bot-golden-set

CASE                      SCORE  RESULT
------------------------------------------------
refund-policy             1.000  PASS
    [ok] icontains     found '30 days'
    [ok] regex         /support@\S+\.\w+/ matched
    [ok] not-contains  "I don't know" absent
    [ok] word-count    words=20 in [8,60]
    [ok] latency       latency=240ms (<= 800.0)
json-extraction           1.000  PASS
    [ok] json-valid    valid JSON
    [ok] json-schema   schema valid
    [ok] json-path     status='shipped'
summary-similarity        0.905  PASS
    [ok] similarity    cosine=0.714 (>= 0.55)
    [ok] contains      found 'Cancel Subscription'
    [ok] length        len=85 in [20,200]
tone-guardrail            1.000  PASS
    [ok] all-of        all-of(regex:P, icontains:P)
    [ok] not-regex     /(?i)stupid|idiot|whatever/ absent
------------------------------------------------
cases: 4/4 passed   pass_rate=1.000   mean_score=0.976

Blocks above are real evalbench output — reproduce them from a clone.

Usage — step by step

  1. Install the harness:

    pip install cognis-evalbench
    
  2. Try the bundled golden set (no files needed), or run your own suite JSON and save the result for later gating:

    evalbench demo
    evalbench run suite.json --save baseline.json
    
  3. Gate a candidate against a baseline. gate accepts saved run results or raw suites and flags score drops beyond --tolerance (add --strict for pass-rate / mean-score regressions):

    evalbench gate baseline.json candidate.json --tolerance 0.02 --format json
    
  4. Read the result. evalbench types lists the assertion types (contains, regex, json-schema, similarity, latency, cost, and more). Exit 0 = success, 1 = case failure / regression, 2 = usage/IO error.

  5. Automate in CI. Evaluate then gate so regressions fail the build:

    evalbench run suite.json --save run.json && evalbench gate baseline.json run.json
    

Contents

Why evalbench?

CI for agents

evalbench is single-purpose, scriptable, and self-hostable: point it at a target, get prioritized results in the format your workflow already speaks (table · JSON · SARIF), gate CI on it, and let agents drive it over MCP.

Features

  • ✅ Load Suite
  • ✅ Run Suite
  • ✅ Compare Baseline
  • ✅ Runs on Linux/macOS/Windows · Docker · devcontainer
  • ✅ Ports in Python, JavaScript, Go, and Rust (ports/)

Quick start

pip install cognis-evalbench
evalbench --version
evalbench scan .                       # scan current project
evalbench scan . --format json         # machine-readable
evalbench scan . --fail-on high        # CI gate (non-zero exit)

Example

$ evalbench scan .
  [HIGH    ] EVA-001  example finding             (./src/app.py)
  [MEDIUM  ] EVA-002  another signal              (./config.yaml)

  2 findings · risk score 5 · 38ms

Architecture

flowchart LR
  IN[agent / A2A traffic] --> P[evalbench<br/>map + analyze]
  P --> OUT[graph + flags]

Use it from any AI stack

evalbench is interoperable with every popular way of using AI:

  • MCP serverevalbench mcp (Claude Desktop, Cursor, Cognis.Studio, uncensored-fleet)
  • OpenAI-compatible / JSON — pipe evalbench scan . --format json into any agent or LLM
  • LangChain · CrewAI · AutoGen · LlamaIndex — wrap the CLI/JSON as a tool in one line
  • CI / scripts — exit codes + SARIF for non-AI pipelines

How it compares

Cognis evalbench promptfoo
Self-hostable, no account varies
Single command, zero config ⚠️
JSON + SARIF for CI varies
MCP-native (AI agents)
Polyglot ports (JS/Go/Rust)
Open license ✅ COCL varies

Built in the spirit of promptfoo / deepeval, re-framed the Cognis way. Missing a credit? Open a PR.

Integrations

Pipes into your stack: SARIF for code-scanning, JSON for anything, an MCP server (evalbench mcp) for AI agents, and a webhook forwarder for SIEM/Slack/Jira. See docs/INTEGRATIONS.md.

Install — every way, every platform

pip install "git+https://github.com/cognis-digital/evalbench.git"    # pip (works today)
pipx install "git+https://github.com/cognis-digital/evalbench.git"   # isolated CLI
uv tool install "git+https://github.com/cognis-digital/evalbench.git" # uv
pip install cognis-evalbench                                          # PyPI (when published)
docker run --rm ghcr.io/cognis-digital/evalbench:latest --help        # Docker
brew install cognis-digital/tap/evalbench                             # Homebrew tap
curl -fsSL https://raw.githubusercontent.com/cognis-digital/evalbench/main/install.sh | sh
Linux macOS Windows Docker Cloud
scripts/setup-linux.sh scripts/setup-macos.sh scripts/setup-windows.ps1 docker run ghcr.io/cognis-digital/evalbench DEPLOY.md (AWS/Azure/GCP/k8s)

Related Cognis tools

  • agentsmith — Config-first scaffolding and orchestration for multi-agent workflows
  • skillhub — Local skill registry and installer for AI agents
  • toolguard — Runtime allowlist and policy for agent tool-calls
  • ragkit — Batteries-included local RAG pipeline — ingest, index, serve
  • memorybank — Portable long-term memory store for agents, exposed over MCP
  • promptpack — Versioned prompt / template registry with A/B and rollbacks

Explore the suite → 🗂️ all 170+ tools · ⭐ awesome-cognis · 🔗 cognis-sources · 🤖 uncensored-fleet · 🧠 engram

Contributing

PRs, new rules, and demo scenarios are welcome under the collaboration-pull model — see CONTRIBUTING.md and SECURITY.md.

⭐ If evalbench saved you time, star it — it genuinely helps others find it.

Interoperability

{} composes with the 300+ tool Cognis suite — JSON in/out and a shared OpenAI-compatible /v1 backbone. See INTEROP.md for the suite map, composition patterns, and reference stacks.

License

Source-available under the Cognis Open Collaboration License (COCL) v1.0 — free for personal, internal-evaluation, research, and educational use; commercial / production use requires a license ([email protected]). See LICENSE.


Cognis Digital · one of 170+ tools in the Cognis Neural Suite · Making Tomorrow Better Today

from github.com/cognis-digital/evalbench

Установить Evalbench в Claude Desktop, Claude Code, Cursor

Рекомендуется · одна команда, все IDE
unyly install evalbench

Ставит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.

Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh

Или настроить вручную

Выполни в терминале:

claude mcp add evalbench -- uvx --from git+https://github.com/cognis-digital/evalbench cognis-evalbench

Пошаговые гайды: как установить Evalbench

FAQ

Evalbench MCP бесплатный?

Да, Evalbench MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Evalbench?

Нет, Evalbench работает без API-ключей и переменных окружения.

Evalbench — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Evalbench в Claude Desktop, Claude Code или Cursor?

Открой Evalbench на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Evalbench with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории ai