Harness Rs
БесплатноНе проверенMCP (Model Context Protocol) stdio server exposing the harness-rs tool registry to Claude Code, Cursor, and Codex.
Описание
MCP (Model Context Protocol) stdio server exposing the harness-rs tool registry to Claude Code, Cursor, and Codex.
README
Agent = Model + Harness. This is the Harness — the scaffolding that turns an LLM into an autonomous agent. Any domain: research, ops, assistants, data work, coding.
A Rust framework for production agents. A governed agent loop that earns autonomy in stages, an OS-enforced sandbox that's honest about what it isolates, and a three-layer local memory that makes the agent know more every session. Compile-time type-safe, deterministic-first, observable, governance built in. Built on the harness engineering discipline (Böckeler/Thoughtworks, Lopopolo/OpenAI, 2026). Full rationale in DESIGN.md.
What makes it
- A loop, not a call. ReAct with tool dispatch, sensor feedback, and auto-fix — running under a token budget until the task is verifiably done.
- Autonomy earned in stages. L1 report → L2 assisted (human gate) → L3 unattended (allowlist only). Graduate the agent as you build trust.
- Honest isolation. OS-enforced sandboxing (macOS Seatbelt / Linux Bubblewrap) that reports what it actually enforces.
- Memory that compounds. Three local layers — procedural (skills), semantic (facts), episodic (searchable history) — so it knows more each run.
- Code does what code can. lint · format · git run as Sensors and Hooks, not model turns. Measured, not asserted.
What you get
| Layer | What | Crate |
|---|---|---|
| Models | 3 protocol families (OpenAI-compat · Anthropic · Gemini), one ApiKind::build(url, model, key) |
harness-models |
| Tools | fs · shell (risk-gated) · web search/fetch | harness-tools-* |
| Loop | ReAct + tool dispatch + sensor feedback + auto-fix | harness-loop |
| Loop engineering | recurring loops: maturity levels L1/L2/L3, human gates, action executors, token budgets | harness-loop-engine |
| Orchestration | async Run = concurrent Job DAG + retry/backoff + dynamic replanning + resumable state | harness-orchestrator |
| Learning | record episodes (situation → tools used → outcome) + semantic recall · CortexDB-backed Memory |
harness-experience, harness-cortexdb |
| Skills · Guides · Hooks · Sensors | proc-macro registered, agentskills.io-compliant | harness-macros, harness-skills |
| Memory · Recall | Memory trait + JSONL store · cross-session search (FTS5 / CJK) |
harness-core, harness-recall-sqlite |
| Privacy | PII detect + redact (label/mask/hash/block, Luhn-checked cards) · redact-on-write Memory decorators |
harness-redact, harness-context |
| Documents | read_document — PDF/Word/Excel/PPT locally (pure Rust) · offline OCR for scanned PDFs (pdftoppm + tesseract) · LLM/vision fallback |
harness-tools-docs |
| Observability | TelemetryHook (structured tracing spans → OTLP) · JSONL session record + deterministic offline replay |
harness-loop |
| Scheduler · MCP · Sandbox · CLI | recurring jobs · MCP server+client · OS-native sandbox (macOS Seatbelt) · git-worktree / Docker · harness code / run / replay / trace / sched / new / mcp serve |
— |
Quick start
use harness_loop::AgentLoop;
use harness_models::ApiKind;
use harness_tools_fs::{ListDir, ReadFile};
use harness_context::default_world;
use harness_core::Task;
use std::sync::Arc;
#[tokio::main]
async fn main() -> Result<(), harness::HarnessError> {
// One model API: protocol family + base_url + model + key. No hardcoded URLs.
let model = ApiKind::OpenAI.build("https://api.deepseek.com", "deepseek-chat",
std::env::var("DEEPSEEK_API_KEY").unwrap());
let mut world = default_world(".");
let outcome = AgentLoop::new(model)
.with_tool(Arc::new(ReadFile))
.with_tool(Arc::new(ListDir))
.run(Task { description: "What is the workspace name?".into(),
source: None, deadline: None }, &mut world)
.await?;
println!("{outcome:?}");
Ok(())
}
Register tools/skills/guides/sensors/hooks with #[harness::tool] / #[skill]
/ #[guide] / #[sensor] / #[hook] — they auto-register via inventory.
Scaffold a new project with harness new.
Composable layers
Start with one agent; compose upward as the work grows.
harness-loopruns one agent (ReAct: think → call tools → observe).harness-loop-enginegoverns a recurring loop: it earns autonomy in stages — L1 report → L2 assisted (human gates every change) → L3 unattended (allowlisted actions only) — under a token budget, with anActionExecutorfor the side effect after a verified approval.harness-orchestratorfans one goal across many concurrent, dependent Jobs (a DAG) with retry/backoff, a run budget, crash-resumable state, and dynamic replanning (aPlannermutates the DAG mid-run from results).harness-experiencemakes an agent learn: it records each run as an episode (situation → tools used → outcome) and recalls similar past episodes to guide the next run. Pair withharness-cortexdb(a CortexDB-backedMemory) for semantic recall over a brain shared with Claude Code / Codex.
use harness_orchestrator::{Dag, Job, Orchestrator, Run, SubagentJobRunner};
// notion/airtable/coda run concurrently; `compare` waits for all three.
let dag = Dag::from_jobs([
Job::new("notion", "what is Notion best at?"),
Job::new("airtable", "what is Airtable best at?"),
Job::new("coda", "what is Coda best at?"),
Job::new("compare", "compare them").with_deps(["notion", "airtable", "coda"]),
]);
let report = Orchestrator::new(Arc::new(SubagentJobRunner::new(model, ".")))
.run(Run::new("run-1", "compare tools", dag)).await;
Design principles
- Don't burn tokens on what code can do — lint/format/git run via Sensors and Hooks, not the model. The Compactor manages scarce context.
- Isolate, don't interrupt — permissions are decided at sandbox spawn, not
prompted per call. Backends are honest about what they enforce
(
Isolation::{None, Changes, Process}):SeatbeltSandbox(macOS, kernel-level viasandbox-exec) andContainerSandbox(Docker) enforce;WorktreeSandboxisolates git changes, not capability. Today a sandbox wraps shell exec; in-process fs tools are jailed separately. - Earn autonomy in stages — start at L1, set a budget, graduate only as you build trust. Unattended loops make unattended mistakes; verification is on you.
Coding agent
harness code is an interactive, opencode-style coding REPL built entirely on
the framework above — multi-turn, streaming, with read/write/edit/list/grep/glob
and shell tools. It runs in NORMAL mode (every write, edit, and shell command
waits for a y/N you approve) or --yolo (unattended). A single Rust binary,
any OpenAI-compatible model:
HARNESS_API_KEY=… HARNESS_BASE_URL=… HARNESS_MODEL=… harness code # NORMAL
harness code --yolo --workspace . # YOLO
Examples
See examples/ — memory, recall, the scheduler, MCP,
redaction-demo (PII redaction: the engine + GuardedMemory / RedactingMemory),
experience-cortexdb (the learning layer over a CortexDB brain),
cap (a coding agent reimplementing oh-my-pi's
hashline editing — content-hash line anchors instead of line numbers), and
two end-to-end agents over a live PostgreSQL database: ecommerce-analyst
(concurrent analysis DAG) and ecommerce-ops-agent (the full stack —
dynamic replanning, L1/L2/L3 governed DB writes, cross-run memory).
Privacy & documents
Redact PII before it persists. harness-redact detects card numbers
(Luhn-checked, so order numbers survive), emails, phones, and money, then
rewrites them — Label (<EMAIL>), Mask (************1111), Hash (a
stable pseudonym), or Block. It's redact-not-drop: the surrounding fact is
kept. Two Memory decorators wire it to the two write boundaries that leak:
use harness_context::{GuardedMemory, RedactingMemory};
// (1) agent long-term memory — redact on write; money/secrets drop
let memory: Arc<dyn Memory> = Arc::new(
GuardedMemory::new(file_mem).with_blocked_substring("password"),
);
// (2) transcript / experience → CortexDB — redact-only, never drops a turn
let safe: Arc<dyn Memory> = Arc::new(RedactingMemory::new(cortex_mem));
spawn_transcript_writer(rx, safe); // every captured turn is scrubbed
Or use the engine directly on any text: Redactor::new().scrub(text).text.
Read documents, including scanned PDFs. read_document extracts text from
PDF/Word/Excel/PowerPoint locally in pure Rust (zero tokens). A scanned,
image-only PDF has no text layer — enable the ocr-tesseract feature for
offline OCR (rasterise with pdftoppm, recognise with tesseract; still zero
tokens, still deterministic):
harness-rs-tools-docs = { version = "0.0.25", features = ["ocr-tesseract"] }
let tool = ReadDocument::new(); // local + OCR
let tool = ReadDocument::with_llm_fallback(model); // + vision fallback for images
// model call: {"path": "scan.pdf", "ocr_lang": "eng+chi_sim"}
// result carries source = "local" | "ocr" | "llm"
Observability
Two seams, one instrumentation. TelemetryHook projects the agent's
lifecycle onto structured tracing spans (agent_run → model.complete /
tool.call / budget.warning, with token, latency, and ok/err fields).
Attach tracing-opentelemetry and they export to Jaeger / Tempo / any OTLP
backend unchanged; attach tracing_subscriber::fmt().json() for a log pipeline.
Record + replay makes a run reproducible for free:
harness run "…" --workspace ./ws --write --record run.jsonl # capture live
harness replay run.jsonl --workspace ./ws # re-run offline, no LLM
harness trace run.jsonl --verbose # inspect the timeline
replay drives the loop from the recorded model outputs (a MockModel) and
re-executes the tool calls, so it reproduces the exact Outcome with zero API
cost — record once, regression-test in CI forever.
Benchmarks
Completion rate. bench-suite runs a task set where each task carries a
machine verifier (a shell assertion the harness runs outside the agent), so
"resolved" is objective — not the model grading itself. qwen3.7-plus via
Aliyun MaaS, 2026-07-09:
| task | status | iters | tools | in tok | out tok |
|---|---|---|---|---|---|
| sum-file | resolved | 3 | 2 | 2793 | 214 |
| rename-key | resolved | 3 | 2 | 2805 | 261 |
| count-lines | resolved | 3 | 2 | 2730 | 156 |
| fix-typo | resolved | 3 | 2 | 2782 | 219 |
| create-readme | resolved | 2 | 1 | 1733 | 144 |
pass@1 = 5/5 (100%) on the Rust-native set. Reproduce with
HARNESS_API_KEY=… HARNESS_BASE_URL=… HARNESS_MODEL=… cargo run -p eval-bench --bin bench-suite
(exits non-zero on any failure, so CI can gate on it).
Cost. Measured token cost on a fixed task set — deepseek-v4-flash via
Aliyun MaaS, 2026-07-04. Every task finished (Done) with side effects verified
(sum.txt = 42, etc.). Reproduce any row with harness run "<task>" --json:
| task | iters | tool calls | in tok | out tok |
|---|---|---|---|---|
| list a directory | 2 | 1 | 975 | 103 |
| read a file, then answer | 2 | 1 | 992 | 130 |
| create a file | 2 | 1 | 1350 | 107 |
| read → sum numbers → write result | 3 | 2 | 2336 | 260 |
File writes go through the write_file tool (small structured args), not the
model re-emitting whole file bodies each turn — "don't burn tokens on what code
can do", measured rather than asserted. cargo run -p eval-bench emits the same
per-task cost fields for cross-framework comparison.
Status
Latest: v0.0.25 — verifier-driven completion-rate benchmark
(eval-bench --bin bench-suite, objective pass@1 via machine verifiers) and
stuck detection (StuckPolicy / Outcome::Stuck: nudge on repeated identical
tool calls, then terminate cleanly instead of burning the budget).
Full history in CHANGELOG.md.
License
MIT OR Apache-2.0.
Установка Harness Rs
У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.
▸ github.com/liliang-cn/harness-rsFAQ
Harness Rs MCP бесплатный?
Да, Harness Rs MCP бесплатный — установка в пару кликов через Unyly без оплаты.
Нужен ли API-ключ для Harness Rs?
Нет, Harness Rs работает без API-ключей и переменных окружения.
Harness Rs — hosted или self-hosted?
Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.
Как установить Harness Rs в Claude Desktop, Claude Code или Cursor?
Открой Harness Rs на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.
Похожие MCP
GitHub
PRs, issues, code search, CI status
автор: GitHubFilesystem
Secure file operations with configurable access controls.
Memory
Knowledge graph-based persistent memory system.
Template MCP Server
A CLI tool to create a new Model Context Protocol server project with TypeScript support, dual transport options, and an extensible structure
автор: mcpdotdirectCompare Harness Rs with
Не уверен что выбрать?
Найди свой стек за 60 секунд
Автор?
Embed-бейдж для README
Похожее
Все в категории development
