Tallowfield Technologies Data Quality
БесплатноНе проверенData quality mcp-server for Tallowfield Technologies — CPU-only, no GPU available in the target environment
Описание
Data quality mcp-server for Tallowfield Technologies — CPU-only, no GPU available in the target environment
README
Record-level lineage with replayable transformations, delivered as an MCP server, for data-quality work at Tallowfield Technologies.
The problem
Freight and logistics records arrive dirty: whitespace from CSV exports, inconsistent casing on port and carrier codes, dates in half a dozen formats, missing fields. Cleaning them is easy. Proving what was changed, why, and that the cleaned record genuinely follows from the raw one is the hard part — and it is what an auditor or a downstream dispute actually needs.
tidewaterdata makes every change go through a named, parameterised, deterministic transformation and records each application with content hashes on both sides. Given the original record and the recorded lineage you can replay the whole history and prove the stored result follows from the stored steps — or find the exact step where something was altered.
How the CPU-only constraint shapes the design
The target environment is CPU-only, with no GPU, and in practice keeps a small, often-isolated runtime footprint. That drives concrete choices:
- Correctness rests on canonical JSON + stdlib
blake2bhashing and pure-Python transforms. Nothing wants an accelerator. (ADR 0002) - The runtime core has no third-party dependencies, including the MCP server itself, which implements the JSON-RPC stdio protocol directly. (ADR 0001)
- Any model/LLM behaviour goes through a provider interface. The default
StubProvideris deterministic and offline, so the whole test suite runs with no API key and no network. If the networkedRealProvideris selected but unconfigured — the normal case on an isolated box — the service degrades to the stub with a logged warning instead of failing. (ADR 0004)
tests/test_offline_cpu.py holds this line: it runs the full flow with the
socket layer stubbed out and asserts no GPU/model libraries were imported.
Architecture
MCP client (stdio JSON-RPC) shell
| |
tidewaterdata/server.py tidewaterdata/cli.py
\_____________ _____________/
\/
tidewaterdata/service.py
validate -> run core -> time -> structured log (stderr)
|
validation.py config.py lineage.py (Pipeline / replay / verify_chain)
|
transforms.py (registry) hashing.py (canonical JSON + blake2b)
|
providers/ base | stub (offline) | real (networked)
The core (lineage, transforms, hashing, types) knows nothing about
transport. service.py is the one place that validates input, times work, and
logs. server.py and cli.py are thin adapters over it. Transformations are
identified by (name, params) and resolved through a registry, so a lineage
serialises to pure JSON and replays anywhere without shipping code. (ADR 0003)
Decision records for the contested calls are in docs/adr/.
Install
make install # creates .venv and installs the package with dev extras
make uses PY ?= .venv/bin/python; override with make test PY=/path/to/python.
The core imports only the standard library, so you can also run it straight from
a checkout without installing anything:
python3.12 -m tidewaterdata list-transforms
Quickstart
Every command below runs against a checkout with python3.12 -m tidewaterdata;
after make install, the console script tidewater is equivalent.
List the available transforms:
python3.12 -m tidewaterdata list-transforms
Clean a record and get back the result plus a replayable lineage:
echo '{"record": {"id": "1", "name": " acme freight "},
"pipeline": [{"name": "trim_whitespace", "params": {"field": "name"}},
{"name": "collapse_spaces", "params": {"field": "name"}}]}' \
| python3.12 -m tidewaterdata clean -i -
The result.name is "acme freight", and lineage carries the per-step hashes
and diffs. Save the whole {record, lineage} and verify it later:
echo '{"record": {"id": "1", "name": " acme freight "},
"lineage": { ...the lineage printed above... }}' \
| python3.12 -m tidewaterdata replay -i -
replay prints {"result": ..., "verified": true} on success, or exits 2
with a replay mismatch at step N message if the record or lineage was altered.
Run the MCP server (reads newline-delimited JSON-RPC on stdin):
python3.12 -m tidewaterdata serve
Library use:
from tidewaterdata import Pipeline, Transformation, replay
rec = {"id": "shp-1", "name": " acme freight ", "port": "lax"}
pipe = Pipeline([
Transformation("trim_whitespace", {"field": "name"}),
Transformation("collapse_spaces", {"field": "name"}),
Transformation("to_upper", {"field": "port"}),
])
lineage, cleaned = pipe.run(rec)
assert replay(rec, lineage) == cleaned # raises ReplayMismatch if tampered
MCP tools
The server exposes four tools via tools/list / tools/call:
| Tool | Arguments | Returns |
|---|---|---|
list_transforms |
none | available transform names |
clean_record |
record, pipeline |
{result, lineage} |
replay_lineage |
record, lineage |
{result, verified} |
suggest_transforms |
record, required_fields? |
{suggestions} |
Validation failures and replay mismatches come back as a normal tools/call
result with isError: true and a message; malformed JSON or an unknown method
comes back as a JSON-RPC error.
Transforms
trim_whitespace, collapse_spaces, to_upper, to_lower, default_missing,
drop_field, rename_field, coerce_int, standardize_date. Register your own
with @tidewaterdata.register("name") — a transform is a pure dict -> dict
function that must not mutate its input.
Configuration
All optional, read from the environment at startup and validated there (a bad
value fails immediately with a configuration error):
| Variable | Default | Meaning |
|---|---|---|
TIDEWATER_PROVIDER |
stub |
stub (offline) or real (networked) |
TIDEWATER_MAX_RECORD_BYTES |
1000000 |
reject records larger than this |
TIDEWATER_MAX_STEPS |
100 |
reject pipelines longer than this |
TIDEWATER_MAX_BATCH |
1000 |
batch-size cap (reserved; see limitations) |
TIDEWATER_LOG_LEVEL |
INFO |
DEBUG..CRITICAL |
Config.from_file(path) overlays these defaults with keys from a JSON object.
Logs are one JSON object per line on stderr (stdout is the protocol/result stream). The size and length limits are the defence against resource exhaustion: oversized input is rejected before any work is done.
The real provider additionally reads TIDEWATER_PROVIDER_ENDPOINT and
TIDEWATER_PROVIDER_API_KEY, and needs the real extra
(pip install -e .[real], which pulls in httpx).
Test
make test # or: python3.12 -m pytest
CLI exit codes: 0 success, 2 bad input / validation / missing file / replay
mismatch, 1 unexpected.
Known limitations
- No batch tool yet.
TIDEWATER_MAX_BATCHis wired into config but there is noclean_batchMCP tool; records are cleaned one at a time. The cap exists so the limit is enforced the moment batching lands. - Replay depends on transform behaviour staying stable across versions.
Changing what a registered transform does will break replay of older lineages —
surfaced as a
ReplayMismatch, not a silent wrong answer, but it is a real compatibility obligation. (ADR 0003) - 16-byte hashes detect tampering and drift, not a determined forger who recomputes the whole chain. Signing the final lineage is a future option. (ADR 0002)
- The
realprovider's response contract is provisional. It assumes a{"suggestions": [...]}shape; see the TODO inproviders/real.pypending a frozen endpoint spec. - No deployment wrapper. systemd/container packaging is out of scope for this delivery.
Tallowfield Technologies is an illustrative client; this repository is a self-directed reference implementation built to work end to end.
from github.com/J-X0/tallowfield-technologies-data-quality-mcp
Установить Tallowfield Technologies Data Quality в Claude Desktop, Claude Code, Cursor
unyly install tallowfield-technologies-data-qualityСтавит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.
Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh
Или настроить вручную
Выполни в терминале:
claude mcp add tallowfield-technologies-data-quality -- uvx --from git+https://github.com/J-X0/tallowfield-technologies-data-quality-mcp tidewaterdataПошаговые гайды: как установить Tallowfield Technologies Data Quality
FAQ
Tallowfield Technologies Data Quality MCP бесплатный?
Да, Tallowfield Technologies Data Quality MCP бесплатный — установка в пару кликов через Unyly без оплаты.
Нужен ли API-ключ для Tallowfield Technologies Data Quality?
Нет, Tallowfield Technologies Data Quality работает без API-ключей и переменных окружения.
Tallowfield Technologies Data Quality — hosted или self-hosted?
Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.
Как установить Tallowfield Technologies Data Quality в Claude Desktop, Claude Code или Cursor?
Открой Tallowfield Technologies Data Quality на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.
Похожие MCP
GitHub
PRs, issues, code search, CI status
автор: GitHubFilesystem
Secure file operations with configurable access controls.
Memory
Knowledge graph-based persistent memory system.
Template MCP Server
A CLI tool to create a new Model Context Protocol server project with TypeScript support, dual transport options, and an extensible structure
автор: mcpdotdirectAmap Maps Mcp Server
MCP server for using the AMap Maps API
автор: duxiaohuiSupabase
Database, auth and storage
автор: SupabaseEverything
Reference / test server with prompts, resources, and tools.
Git
Tools to read, search, and manipulate Git repositories.
Sequential Thinking
Dynamic and reflective problem-solving through thought sequences.
Time
Time and timezone conversion capabilities.
Compare Tallowfield Technologies Data Quality with
Не уверен что выбрать?
Найди свой стек за 60 секунд
Автор?
Embed-бейдж для README
Похожее
Все в категории development
