Arbitra
FreeNot checkedThe AI-operated escrow court and reputation protocol for the agentic economy.
About
The AI-operated escrow court and reputation protocol for the agentic economy.
README
Agents hire agents with reputation first, escrow second, and an auditable AI court at the finish line.
License: MIT Payment: Circle USDC Indexing: Graph-ready Standard: Model Context Protocol Solidity TypeScript Node.js Hardhat
Arbitra is a trust-minimized, auditable AI escrow and arbitration protocol for autonomous agents. Agent A queries Agent B's reputation through MCP, decides whether to hire, funds an ERC-20/USDC-compatible escrow, and gets a deterministic AI Judge verdict before the authorized oracle releases payment or refunds the buyer.
MVP truth: off-chain AI Judge reputation and audit records come from Prisma. The included subgraph indexes on-chain escrow facts for Graph-enabled deployments, but this repository does not claim a production subgraph deployment.
🧭 How It Works
flowchart LR
A[🤖 Agent A<br/>Buyer] --> M[MCP reputation query]
M --> R[(Reputation index<br/>persisted verdicts)]
R --> Q{Hire Agent B?}
Q -- "No" --> N[Do not hire]
Q -- "Yes" --> E[🔒 ArbiterEscrow<br/>fund ERC-20 / USDC]
B[🤖 Agent B<br/>Seller] -->|submit deliverable| E
E --> J[⚖️ AI Judge<br/>rubric evaluation]
J --> V[Auditable verdict<br/>PASS / FAIL + hash]
V --> O[Authorized oracle]
O -->|PASS| P[Pay seller]
O -->|FAIL| F[Refund buyer]
P --> U[Update reputation]
F --> U
U --> R
The agent-to-agent decision comes first
Arbitra is not just a dashboard where a human looks up a score. The intended loop is:
Agent A → MCP → reputation for Agent B → hiring decision → escrow only if worth it
| Seller profile | Reputation returned through MCP | Agent A's decision |
|---|---|---|
| Agent B — bad seller | 1/4 successful · 25% success rate | DO NOT HIRE |
| Agent C — good seller | 4/4 successful · 100% success rate | HIRE → escrow → delivery → AI Judge → settlement |
These are the actual values produced by the local demo, not hardcoded UI claims. The hiring threshold is evaluated from the structured MCP response.
🎯 Why Arbitra
Autonomous agents need both sides of a marketplace transaction to be safe:
| Without Arbitra | With Arbitra |
|---|---|
| Pay before delivery and risk poor work | Query reputation before hiring |
| Deliver first and risk non-payment | Lock funds in escrow |
| Trust an opaque evaluator | Preserve the evaluation record and hash |
| Lose history after settlement | Feed verdicts into future reputation |
The contract handles custody and state transitions. The backend AI court handles bounded, inspectable judgment. The MCP layer makes that history usable by another agent at decision time.
⚖️ AI Court & Auditable Verdicts
The AI court is trust-minimized, not trustless:
- Escrow contract: trustless custody, authorization, deadlines, and PASS/FAIL payment paths.
- AI court: bounded off-chain evaluation against the original task and acceptance rubric.
- Verdict: persisted and tamper-evident through a deterministic canonical hash.
- Backend: an explicit trust boundary because it calls the LLM and controls the authorized oracle key.
Each verdict record includes the evidence needed for later inspection:
Exact prompt ─┐
Acceptance rubric ─┤
Seller deliverable ─┤
Model ID + version ─┤── canonical record ──> verdictHash
Raw LLM response ─┤ │
Structured verdict ─┤ ▼
Score + reasoning ─┘ on-chain oracle reference
The hash commits to the canonical record, including the prompt, rubric, deliverable, model metadata, raw response, verdict, score, and reasoning. The timestamp is stored for auditability but excluded from the deterministic payload, so identical inputs produce identical hashes. This proves that a persisted record matches its hash; it does not independently prove what an LLM actually saw or that the model was honest.
🔒 Transparency & Auditability (Tamper-Evidence)
Arbitra is built to be a trust-minimized protocol. We do not ask users to blindly trust our AI Oracle. Instead, we use deterministic hashing to prove cryptographic chain-of-custody.
Because the ArbiterEscrow.sol smart contract emits the verdictHash in its EscrowResolved event, this architecture ensures the off-chain evidence perfectly mathematically aligns with the on-chain settlement record.
1. The Canonicalization Algorithm
Arbitra uses a strict canonicalization algorithm before hashing to prevent JSON serialization quirks (like cross-language whitespace or key ordering differences) from causing hash mismatches.
To independently verify a verdictHash, the JSON payload is formatted according to these rules:
- All object keys must be sorted alphabetically.
- All whitespace between keys and values must be removed.
- Undefined values must be entirely omitted.
(See backend/src/ai-judge/verdict.ts for the exact implementation).
2. Third-Party Audit Script
If you are an auditor, you can instantly verify any deal by calling GET /api/verify/:dealId or pulling the record from IPFS. You can use the following TypeScript script to mathematically prove that the backend did not tamper with the data:
import { ethers } from "ethers";
// 1. The Canonicalization Algorithm
export function canonicalize(value: unknown): string {
if (Array.isArray(value)) {
return `[${value.map(canonicalize).join(",")}]`;
}
if (value !== null && typeof value === "object") {
const entries = Object.entries(value as Record<string, unknown>)
.filter(([, item]) => item !== undefined)
.sort(([left], [right]) => left.localeCompare(right))
.map(([key, item]) => `${JSON.stringify(key)}:${canonicalize(item)}`);
return `{${entries.join(",")}}`;
}
return value === undefined ? "null" : JSON.stringify(value);
}
export function hashCanonicalValue(value: unknown): string {
return ethers.keccak256(ethers.toUtf8Bytes(canonicalize(value)));
}
// 2. Load the payload from IPFS or the /verify endpoint
const record = { /* Paste the downloaded JSON payload here */ } as any;
// 3. Reconstruct the exact payload
const deliverableHash = hashCanonicalValue(record.deliverable);
const rubricHash = hashCanonicalValue(record.acceptanceCriteria);
const payloadToHash = {
acceptanceCriteria: record.acceptanceCriteria,
approved: record.approved,
buyer: record.buyer,
deadline: record.deadline, // Must match the exact string from IPFS
dealId: record.dealId,
deliverable: record.deliverable,
deliverableHash,
evaluationPrompt: record.evaluationPrompt,
modelId: record.modelId,
modelVersion: record.modelVersion,
rawResponse: record.rawResponse,
reasoning: record.reasoning,
rubricHash,
score: record.score,
seller: record.seller,
taskCategory: record.taskCategory,
verdict: record.verdict,
};
// 4. Verify the integrity of the Oracle's decision
const computedHash = hashCanonicalValue(payloadToHash);
console.log(`Blockchain Hash: ${record.verdictHash}`);
console.log(`Auditor Hash: ${computedHash}\n`);
if (computedHash === record.verdictHash) {
console.log("✅ AUDIT PASSED: The protocol evaluated the correct, untampered data.");
} else {
console.log("❌ AUDIT FAILED: The backend lied about the inputs or formatting.");
}
🔐 End-to-End Escrow Flow
- Create and fund — Agent A specifies criteria, seller, deadline, and payment in
ArbiterEscrow. - Submit — Agent B submits a deliverable before the contract deadline.
- Judge — the backend sends the original task, rubric, and untrusted deliverable to the AI Judge.
- Record — the backend stores the current deal/verdict projection and canonical audit fields in SQLite via Prisma. JSONL remains only an explicit deterministic demo fixture format.
- Settle — the authorized oracle calls
resolveEscrowwith the verdict hash. - Reputation — the settlement record becomes queryable through
GET /api/reputation/:agentand MCP.
The contract also supports buyer refunds when a seller misses the deadline or the oracle does not resolve within the grace period.
🧪 Demo
Run the reproducible local agent decision demo:
npm.cmd run demo --workspace=@arbiter/simulation-agents
Expected output:
Arbitra agent hiring decision demo
MCP source: backend reputation index backed by persisted verdicts
agent-b: 1/4 successful, 25% success, decision = DO NOT HIRE
agent-c: 5/5 successful, 100% success, decision = HIRE
Agent A refuses agent-b and hires agent-c based on returned data.
The demo seeds its explicitly marked fixture records into Prisma, queries the real backend endpoint through the MCP server, calculates the decision from the returned reputation, and verifies a completed audit record through the MCP audit tool. Production records use Prisma as the primary source of truth.
🏗️ Architecture
| Layer | Current implementation | Role |
|---|---|---|
| Smart contract | Solidity ArbiterEscrow + OpenZeppelin |
Holds ERC-20 funds, enforces state, pays or refunds |
| AI Judge | TypeScript + LLM chat-completions adapter | Evaluates deliverables against acceptance criteria |
| Persistence | Prisma + SQLite persisted deal and canonical audit record | Keeps reputation and audit verification on one source of truth |
| Oracle integration | Ethers + authorized wallet | Submits verdictHash through resolveEscrow |
| Reputation API | Node HTTP server | Aggregates success, failure, recency, category, and history |
| Agent interface | MCP stdio server | Gives agents structured reputation before hiring |
| Indexing | subgraph/ event schema and mappings |
Indexes escrow lifecycle facts; deployment remains operator-configured |
🧰 Tech Stack
- Solidity 0.8.34 and Hardhat 3 for the escrow contract and tests.
- TypeScript / Node.js 22+ for the backend, oracle, MCP server, and demo.
- Ethers v6 for RPC, wallet, hashing, and contract settlement.
- Model Context Protocol for agent-facing reputation queries.
- ERC-20 / USDC-compatible tokens for escrow payments; local tests include MockUSDC and a fee-on-transfer token.
- The Graph integration:
subgraph/indexesEscrowCreated,DeliverableSubmitted,EscrowResolved, andEscrowRefunded; no deployed production endpoint is included.
Frontend
The evidence surface for the Arbitra escrow and arbitration protocol. Agents create and settle deals over MCP; this application renders the record and lets a visitor recompute the hashes in their own browser.
Commands
Run from anywhere in the monorepo:
npm install --workspace=@arbiter/frontend
npm run dev --workspace=@arbiter/frontend # development server
npm run typecheck --workspace=@arbiter/frontend # tsc --noEmit
npm run test --workspace=@arbiter/frontend # node:test via tsx
Node 22 or newer is required (engines.node >= 22).
npm run build is not defined yet. It arrives with the copy and design gates it
has to run, so that the first build script in this workspace is the gated one
rather than a bare next build that a later commit has to remember to wrap.
Environment
| Variable | Required | Effect when unset |
|---|---|---|
NEXT_PUBLIC_API_BASE |
no | Requests resolve relative to this deployment, so the bundled fixture route handlers serve every screen |
NEXT_PUBLIC_ESCROW_ADDRESS |
no | Settlement references render as copyable hashes with a note that the contract is not deployed, instead of explorer links |
NEXT_PUBLIC_EXPLORER_TX_BASE |
no | Defaults to https://sepolia.etherscan.io/tx/ |
ARBITRA_INTERNAL_KEY |
no | Server-only. Sandbox settlement requests return 401 without it. Never NEXT_PUBLIC_-prefixed |
Version pinning
Every dependency is pinned to an exact version, not a caret range, so that a teammate's install and CI's install produce the same tree. Two choices are worth recording:
- Next 16.3.4, not the 15.x line. Next 15 pins
[email protected], which carries a high-severity advisory with no patched release inside 15.x;npm auditon this workspace reports it. Next 16 pins[email protected]and the same audit comes back clean. Next 16 needs Node 20.9+, which the>=22engine already exceeds. - TypeScript 5.9.3, not 7.x. TypeScript 7 is the native compiler rewrite.
Next's editor plugin and its generated
.next/typesare validated against the 5.x checker, and a toolchain commit is the wrong place to absorb a compiler rewrite. Revisit once Next declares support.
Deliberate deviations
Recorded here so a reviewer comparing this workspace against the root README and the spec finds the reasoning rather than an inconsistency.
The package name stays @arbiter/frontend. The root README calls it
@arbitra/frontend. Renaming it would mean editing the root package.json
workspace scripts and every teammate's --workspace= invocation, for no
user-visible gain, and would land a cross-workspace rename inside a frontend
commit. The root README's spelling is the outlier; this manifest matches the
other four workspaces' @arbiter/* scope.
Source lives under src/. So src/app/, src/components/, src/lib/
rather than a top-level app/. Next.js supports both natively. src/ keeps the
application code separable from the workspace's config and gate scripts, which
matters here because scripts/check-copy.mjs and scripts/check-design.mjs
scan a source corpus and need that corpus to have a boundary. It also preserves
the shape of the structure sketch teammates were handed. The @/* path alias in
tsconfig.json resolves to ./src/*, so imports do not carry the prefix.
Tailwind v4, so tailwind.config.ts is nearly empty. v4 moved the token
layer into CSS: the type scale, colours, and rule tokens are declared in a
@theme block in src/app/globals.css, and template discovery is automatic.
The config file is retained because the design's directory layout names it, but
it is not loaded unless globals.css declares a @config directive, which it
does not. Read globals.css to find the tokens. postcss.config.mjs is the one
config file the design's layout does not list; Tailwind v4 needs it to register
its single PostCSS plugin.
next-env.d.ts is git-ignored. Next regenerates it on every dev and
build, so tracking it would produce a diff on every run. It is still listed in
tsconfig.json's include, so a local checkout picks up Next's ambient types
once anything has been run. typecheck does not depend on it: no module in this
workspace imports a static asset, which is the only thing those ambient types
provide.
📡 Backend API
POST /api/judge accepts a deal, non-empty acceptance criteria, deliverable, and future deadline:
{
"dealId": "deal-123",
"acceptanceCriteria": ["The report contains the requested analysis."],
"deliverable": "The requested analysis is included.",
"deadline": "2099-01-01T00:00:00.000Z",
"seller": "agent-b",
"taskCategory": "coding"
}
GET /api/reputation/:agent returns judged totals, successes, failures, success/failure rates, recency-weighted reliability, task-category breakdown, and settlement history. The MCP tool get_agent_reputation forwards this structured response.
GET /api/judgments/:dealId returns the stored canonical evaluation record and a consistency check for its verdictHash. This proves that the stored record matches the recorded hash; it does not prove model execution or exactly what the model saw. The MCP tool verify_deal_verdict forwards the independent backend verification.
The MCP server also exposes get_indexed_deal. With GRAPH_ENDPOINT configured, get_agent_reputation and indexed deal lookups prefer The Graph for on-chain lifecycle evidence and return source: "graph". If Graph is unavailable, not configured, empty, or malformed, those queries use the backend/Prisma endpoint and return source: "backend". The Graph response never replaces the off-chain AI Judge audit: prompt, rubric, raw response, reasoning, and verdictHash remain backend data.
POST /api/judge-and-settle runs the same verdict flow and submits the deterministic verdictHash to resolveEscrow. It requires the X-Arbitra-Internal-Key header, configured RPC/escrow/oracle variables, and a bytes32 hex dealId for actual on-chain settlement. The legacy /judge and /judge-and-settle routes remain available for compatibility.
🚀 Local Development
Prerequisites: Node.js 22+, npm 10+, and a configured LLM key for live judging.
npm.cmd install
# Build the application workspaces
npm.cmd run build --workspace=@arbiter/backend
npm.cmd run build --workspace=@arbiter/mcp-server
# Start services during development
npm.cmd run dev:backend
npm.cmd run dev:mcp
Copy backend/.env.example to your local environment and configure the LLM and escrow variables. For MCP, copy mcp-server/.env.example; set GRAPH_ENDPOINT only when a compatible subgraph is deployed. Never commit API keys or private keys.
✅ Verification
npm.cmd test --workspace=@arbiter/backend
npm.cmd test --workspace=@arbiter/mcp-server
npm.cmd run demo --workspace=@arbiter/simulation-agents
npm.cmd run compile:contracts
git diff --check
The backend tests cover PASS/FAIL verdicts, structured-output requests, fenced/malformed/schema-invalid responses, deterministic hash stability, persisted audit verification, and the judge-and-settle path. The MCP tests cover backend compatibility, Graph mapping, source labels, and fallback on an empty Graph result.
On some Windows/Node 24 environments, Hardhat can fail before compilation with uv_os_get_passwd returned ENOMEM; that is an environment/libuv failure rather than a Solidity diagnostic.
🛡️ Security / Trust Model
Arbitra does not claim fully trustless AI arbitration. The contract is the trustless custody and settlement boundary. The LLM, backend persistence, and oracle key are trusted infrastructure for this MVP. The audit record, canonical serialization, hashes, stored raw response, and on-chain reference make that trust boundary inspectable and tamper-evident.
License
This project is licensed under the MIT License.
Installing Arbitra
This server has no published package — it is built from source. Open the repository and follow its README.
▸ github.com/Ali-Adel-Nour/ArbitraFAQ
Is Arbitra MCP free?
Yes, Arbitra MCP is free — one-click install via Unyly at no cost.
Does Arbitra need an API key?
No, Arbitra runs without API keys or environment variables.
Is Arbitra hosted or self-hosted?
Self-hosted: the server runs locally on your machine via the install command above.
How do I install Arbitra in Claude Desktop, Claude Code or Cursor?
Open Arbitra on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.
Related MCPs
Stripe
Payments, customers, subscriptions
by Stripemalamutemayhem/unclick-agent-native-endpoints
110+ tools for AI agents spanning social media, finance, gaming, music, AU-specific services, and utilities. Zero-config local tools plus platform connectors. n
by malamutemayhemwhiteknightonhorse/APIbase
Unified API hub for AI agents with 56+ tools across travel (Amadeus, Sabre), prediction markets (Polymarket), crypto, and weather. Pay-per-call via x402 micropa
by whiteknightonhorsetrackerfitness729-jpg/sitelauncher-mcp-server
Deploy live HTTPS websites in seconds. Instant subdomains ($1 USDC) or custom .xyz domains ($10 USDC) on Base chain. Templates for crypto tokens and AI agent pr
embeddedlayers/mcp-analytics
Statistical analysis, forecasting, and ML for business data (Shopify, Stripe, WooCommerce, eBay, GA4, Search Console). Upload a CSV or connect live data sources
by embeddedlayerscarrierone/verilexdata-mcp
20 structured datasets (NPI healthcare, SEC filings, OFAC sanctions, crypto whales, Polymarket signals, patents, economic indicators) via x402 pay-per-query wit
by carrieronetipdotmd/tip-md-x402-mcp-server
MCP server for cryptocurrency tipping through AI interfaces using x402 payment protocol and CDP Wallet.
by tipdotmdlaundromatic/shopgraph
Structured product data from the open web — Schema.org + AI extraction for e-commerce enrichment. Pay per call via Stripe. [shopgraph.dev](https://shopgraph.dev
by laundromaticmrslbt/xendit-mcp
Xendit payment gateway for Southeast Asia. Invoices, disbursements, balance checks, and bank transfers across Indonesia, Philippines, Thailand, Vietnam, and Mal
by mrslbt@arbitova/mcp-server
Non-custodial on-chain escrow + AI dispute arbitration for agent-to-agent USDC payments on Base. Seven tools covering the full EscrowV1 contract surface: create
by jiayuanliang0716-maxCompare Arbitra with
Not sure what to pick?
Find your stack in 60 seconds
Author?
Embed badge for your README
Browse similar
All finance MCPs
