Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Semantix Verify

БесплатноНе проверен

Semantix-Verify is an MCP server for semantic validation of AI/LLM outputs. It exposes a single tool, verify_text_intent(text, intent_description, threshold), w

GitHubEmbed

Описание

Semantix-Verify is an MCP server for semantic validation of AI/LLM outputs. It exposes a single tool, verify_text_intent(text, intent_description, threshold), which uses a local quantized NLI cross-encoder (INT8 ONNX) to return a 0.0–1.0 probability that the text satisfies the given intent — and, when it doesn't, a structured correction prompt for agent retry loops. Useful for building com

README

semantix-ai

Regulator-clause-trained NLI judges, datasets, and audit trails — for compliance work in jurisdictions large vendors skip.

PyPI version Python versions License Downloads Docs semantix-ai MCP server


The SA AI Compliance Stack

semantix-ai is the Python entry point to a coherent, Apache-licensed compliance stack built around South Africa's Protection of Personal Information Act (POPIA), the first publicly-distributed regulator-clause-fine-tuned NLI judge for any jurisdiction.

Artifact What it is Where
semantix-ai Decorator + library that wraps a judge around every LLM call, with hash-chained audit certificates PyPI
nli-popia-v2 10-clause POPIA-grounded NLI judge (consent, minimality, security, breach, cross-border, data-subject-rights, children, special PI, automated decision-making, general processing) HuggingFace
sa-compliance-embeddings-v1 384-dim embeddings fine-tuned on POPIA Act text + grounded scenarios — POPIA-section retrieval (recall@1 0.211 → 0.477 over bge-small-en-v1.5) HuggingFace
popia-instruct-v0 QLoRA adapter on Phi-3-mini for grounded POPIA Q&A. v0 — narrow but real: clause routing + section text recitation, not free-form legal reasoning HuggingFace
POPIA-Bench v1 197-pair public benchmark for clause-level POPIA NLI, with pinned eval hashes and a community leaderboard bench/popia-v1/
POPIAJudge preprint arXiv cs.CL paper documenting the recipe, results, and limitations papers/popiajudge-arxiv/

No competitor publicly distributes a regulator-clause-fine-tuned NLI judge. Llama Guard is trained on hazard taxonomies, Patronus Lynx on RAG faithfulness, Guardrails Hub on PII patterns + classifiers. The clause-pinned compliance niche is empty.

The library below makes the judge usable in a Python program. The stack above makes it defensible in a regulatory review.


Quick start

Validate every LLM output against an explicit intent — a score and a verdict, locally, in ~15–70 ms (varies by CPU), without an API key. Pair it with the audit engine for a hash-chained, tamper-evident receipt.

pip install semantix-ai[turbo]     # local quantized NLI judge; no API key
from semantix import Intent, QuantizedNLIJudge, validate_intent
from semantix.audit.engine import AuditEngine

class ResolutionPolite(Intent):
    """The response must acknowledge the customer's issue and propose a concrete next step, in a polite tone."""

judge = QuantizedNLIJudge()

# Validation: the Intent is the function's return-type annotation.
@validate_intent(judge=judge)
def handle_complaint(message: str) -> ResolutionPolite:
    return call_my_llm(message)

reply = handle_complaint(incoming)   # validated reply, or raises SemanticIntentError

# Audit trail: score, record a certificate, verify the chain, persist it.
# The decorator validates; it does NOT auto-write a certificate — you record it.
engine = AuditEngine()
verdict = judge.evaluate(premise=str(reply),
                         hypothesis="the reply is polite and proposes a next step")
engine.record(intent="ResolutionPolite", output=str(reply),
              score=verdict.score, passed=verdict.passed, judge_id="QuantizedNLIJudge")
assert engine.verify_chain()         # True while the chain is intact
engine.flush("audit.jsonl")          # hash-chained receipts on disk

Why this exists

LLM applications quietly skip the step where you prove the output was fit for purpose. The common fix — calling a bigger LLM as a judge — has three problems:

  1. It drifts. Same input, different score on different runs. A regulator asking "rerun this validation" gets a different answer, which is indistinguishable from evidence the system is broken.
  2. It ships personal information out of your network. Every judge call sends the output to a third-party API. Under POPIA §72 (or GDPR Art. 44, or the EU AI Act's high-risk-system obligations) that's a problem to document, not a default.
  3. It produces no receipt. The validation happened, a score came back, nothing was recorded in a form that survives an audit.

semantix replaces that reflex with a local, deterministic validator and a tamper-evident log. Every validation produces a JSON-LD certificate hash-chained to the previous one. Modify any entry mid-chain and every subsequent hash breaks. The regulator doesn't need to trust your database — the math proves the chain is internally consistent.


What you get

1. Validation as a decorator

from semantix import Intent, validate_intent

class MedicalAdvice(Intent):
    """The text provides a medical diagnosis or treatment recommendation."""

@validate_intent(~MedicalAdvice)  # Must NOT give medical advice
def chatbot(msg: str) -> str:
    return call_my_llm(msg)

Compose with & (all must pass) and | (any must pass):

SafeAndPolite = Polite & ~MedicalAdvice & ~LegalAdvice

2. Tamper-evident audit trail

from semantix.audit.engine import AuditEngine
engine = AuditEngine()

# Bind each certificate to WHAT was judged, BY WHICH judge, ABOUT WHOM.
engine.record(
    intent="POPIA cross-border transfers",
    output=policy_text,                                  # the premise (hashed, never stored raw)
    hypothesis="Personal information is transferred outside South Africa",
    judge_id="POPIAJudge/[email protected]",
    subject="user:5b4c9d12",
    metadata={"destination": "Ashby", "country": "US"},
    score=0.53, passed=False, reason="No consent basis on record.",
)

engine.verify_chain()   # True if no tampering
engine.chain_report()   # integrity AND variety — flags a chain that verifies
                        # perfectly while certifying one repeated result

Each certificate records the hash of the validated text (output_hash) and of the judged claim (claim_hash), the intent and hypothesis, the judge identity and configuration (judge_id, metadata), the subject, the verdict, the timestamp, and the hash of the previous certificate. New certificates use the …/v2 schema; existing …/v1 certificates still verify unchanged, so a chain that upgrades mid-life stays one intact chain. Compatible with JSON-LD tooling and standard audit pipelines.

3. Self-healing retries

On failure, semantix injects structured feedback so the LLM knows what went wrong:

from typing import Optional

@validate_intent(ResolutionPolite, retries=2)
def reply(msg: str, semantix_feedback: Optional[str] = None) -> str:
    prompt = f"Reply to: {msg}"
    if semantix_feedback:
        prompt += f"\n\n{semantix_feedback}"
    return call_llm(prompt)

First call: semantix_feedback is None. On retry: it receives a Markdown report with the score, reason, and rejected output. Measured reliability improves from 21% to 70% across three intent categories.

4. Forensic token-level attribution

from semantix import ForensicJudge, QuantizedNLIJudge
judge = ForensicJudge(QuantizedNLIJudge())
# Verdict.reason: "Suspect tokens: [indemnify, forfeit, waive]"

5. pytest integration

from semantix.testing import assert_semantic

def test_chatbot_is_polite():
    response = my_chatbot("handle angry customer")
    assert_semantic(response, "polite and professional")

On failure:

AssertionError: Semantic check failed (score=0.12)
  Intent:  polite and professional
  Output:  "You're an idiot for asking that."
  Reason:  Text contains aggressive language

First-class pytest plugin with fixtures, markers, and CI reporting: pytest-semantix.


Framework integrations

Drop into your existing stack — retries are handled natively by each framework.

DSPy

import dspy
from semantix.integrations.dspy import semantic_reward

qa = dspy.ChainOfThought("question -> answer")
refined = dspy.Refine(module=qa, N=3, reward_fn=semantic_reward(Polite))

semantic_reward / semantic_metric also plug into dspy.BestOfN, dspy.Evaluate, and MIPROv2 — local, no API calls, ~15 ms per eval. See benchmarks/ for reproducible comparisons against LLM-judge reward functions.

LangChain
from semantix.integrations.langchain import SemanticValidator
validator = SemanticValidator(Polite)
chain = prompt | llm | StrOutputParser() | validator
Pydantic AI
from pydantic_ai import Agent
from semantix.integrations.pydantic_ai import semantix_validator
agent = Agent("openai:gpt-4o", output_type=str)
agent.output_validator(semantix_validator(Polite))
Guardrails AI
from guardrails import Guard
from semantix.integrations.guardrails import SemanticIntent
guard = Guard().use(SemanticIntent("must be polite and professional"))
Instructor
from semantix.integrations.instructor import SemanticStr
from pydantic import BaseModel
class Response(BaseModel):
    reply: SemanticStr["must be polite and professional", 0.85]
MCP
pip install "semantix-ai[mcp,nli]"
mcp run semantix/mcp/server.py

Any MCP-capable agent (Claude Desktop, Cursor, etc.) can validate intents as a tool.

GitHub Actions
- uses: labrat-akhona/semantic-test-action@v1
  with:
    test-path: tests/

Posts a semantic test report as a PR comment.

Install extras: pip install "semantix-ai[dspy]", "[langchain]", "[pydantic-ai]", "[guardrails]", "[instructor]", "[mcp]", "[all]".


Pluggable judges

Choose the speed / accuracy / reasoning trade-off:

from semantix import NLIJudge, EmbeddingJudge, LLMJudge, CachingJudge

@validate_intent(judge=NLIJudge())                           # local, ~15 ms, deterministic
@validate_intent(judge=EmbeddingJudge())                     # local, ~5 ms, similarity-based
@validate_intent(judge=LLMJudge(model="gpt-4o-mini"))        # reasoning, ~500 ms, API
@validate_intent(judge=CachingJudge(NLIJudge(), maxsize=256))  # LRU-wrapped

Quantized mode (INT8 ONNX, ~79 MB, no PyTorch):

pip install "semantix-ai[turbo]"

When this is the right tool

  • You're running an LLM-backed system that processes personal information and need an auditable validation step.
  • You're optimising a DSPy program and the LLM-judge reward loop is too slow, too expensive, or too non-deterministic.
  • You need semantic test assertions in pytest / CI that don't call a paid API.
  • You're in a regulated industry (financial services, insurance, healthcare) and "the model said it was fine" isn't a defensible answer.

When it isn't

  • Your validation intent requires multi-hop reasoning or world knowledge ("is this compliant with section 4(b) of the 2026 tax code"). NLI can't do this; reasoning LLMs can.
  • You need the judge to explain why in prose, not just give a score.
  • You're evaluating fewer than 100 outputs per month and the latency / cost of LLM-as-judge doesn't matter.

See Where semantix fits for a comparison against TruLens, DeepEval, Vectara HHEM, Guardrails, RAGAS, and NeMo.


Key properties

  • Local inference — NLI model runs on CPU, no data leaves your machine.
  • Deterministic per CPU architecture — same input, same score, every time on a given machine (single-threaded ONNX inference). A different pre-quantized INT8 variant loads per architecture (AVX2 / AVX-512 / ARM64), so scores and latency vary across hardware.
  • Fast — ~15–70 ms per check with the quantized judge, depending on CPU.
  • Zero API cost — no tokens burned for validation.
  • Auditable — hash-chained JSON-LD certificates per check.
  • Well-tested — 274 tests, MIT licensed (model weights and datasets ship under Apache-2.0 / CC-BY-4.0).

Installation

pip install semantix-ai                    # Core (default NLI judge)
pip install "semantix-ai[turbo]"           # Quantized ONNX (smallest footprint)
pip install "semantix-ai[openai]"          # LLM judge (GPT-4o-mini)
pip install "semantix-ai[all]"             # Everything

Package name on PyPI is semantix-ai. Import is from semantix import ....


Contributing

See CONTRIBUTING.md for dev setup, testing, and submission guidelines.

License

MIT — see LICENSE.


Built by Akhona Eland in South Africa

from github.com/labrat-akhona/semantix-ai

Установить Semantix Verify в Claude Desktop, Claude Code, Cursor

Рекомендуется · одна команда, все IDE
unyly install semantix-verify

Ставит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.

Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh

Или настроить вручную

Выполни в терминале:

claude mcp add semantix-verify -- uvx semantix-ai

Пошаговые гайды: как установить Semantix Verify

FAQ

Semantix Verify MCP бесплатный?

Да, Semantix Verify MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Semantix Verify?

Нет, Semantix Verify работает без API-ключей и переменных окружения.

Semantix Verify — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Semantix Verify в Claude Desktop, Claude Code или Cursor?

Открой Semantix Verify на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Semantix Verify with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории ai