Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Tallowfield Technologies Data Quality

БесплатноНе проверен

Data quality mcp-server for Tallowfield Technologies — CPU-only, no GPU available in the target environment

GitHubEmbed

Описание

Data quality mcp-server for Tallowfield Technologies — CPU-only, no GPU available in the target environment

README

Record-level lineage with replayable transformations, delivered as an MCP server, for data-quality work at Tallowfield Technologies.

The problem

Freight and logistics records arrive dirty: whitespace from CSV exports, inconsistent casing on port and carrier codes, dates in half a dozen formats, missing fields. Cleaning them is easy. Proving what was changed, why, and that the cleaned record genuinely follows from the raw one is the hard part — and it is what an auditor or a downstream dispute actually needs.

tidewaterdata makes every change go through a named, parameterised, deterministic transformation and records each application with content hashes on both sides. Given the original record and the recorded lineage you can replay the whole history and prove the stored result follows from the stored steps — or find the exact step where something was altered.

How the CPU-only constraint shapes the design

The target environment is CPU-only, with no GPU, and in practice keeps a small, often-isolated runtime footprint. That drives concrete choices:

  • Correctness rests on canonical JSON + stdlib blake2b hashing and pure-Python transforms. Nothing wants an accelerator. (ADR 0002)
  • The runtime core has no third-party dependencies, including the MCP server itself, which implements the JSON-RPC stdio protocol directly. (ADR 0001)
  • Any model/LLM behaviour goes through a provider interface. The default StubProvider is deterministic and offline, so the whole test suite runs with no API key and no network. If the networked RealProvider is selected but unconfigured — the normal case on an isolated box — the service degrades to the stub with a logged warning instead of failing. (ADR 0004)

tests/test_offline_cpu.py holds this line: it runs the full flow with the socket layer stubbed out and asserts no GPU/model libraries were imported.

Architecture

                MCP client (stdio JSON-RPC)          shell
                          |                             |
                   tidewaterdata/server.py       tidewaterdata/cli.py
                          \_____________  _____________/
                                        \/
                            tidewaterdata/service.py
              validate -> run core -> time -> structured log (stderr)
                                        |
        validation.py   config.py   lineage.py (Pipeline / replay / verify_chain)
                                        |
              transforms.py (registry)   hashing.py (canonical JSON + blake2b)
                                        |
                        providers/ base | stub (offline) | real (networked)

The core (lineage, transforms, hashing, types) knows nothing about transport. service.py is the one place that validates input, times work, and logs. server.py and cli.py are thin adapters over it. Transformations are identified by (name, params) and resolved through a registry, so a lineage serialises to pure JSON and replays anywhere without shipping code. (ADR 0003)

Decision records for the contested calls are in docs/adr/.

Install

make install          # creates .venv and installs the package with dev extras

make uses PY ?= .venv/bin/python; override with make test PY=/path/to/python. The core imports only the standard library, so you can also run it straight from a checkout without installing anything:

python3.12 -m tidewaterdata list-transforms

Quickstart

Every command below runs against a checkout with python3.12 -m tidewaterdata; after make install, the console script tidewater is equivalent.

List the available transforms:

python3.12 -m tidewaterdata list-transforms

Clean a record and get back the result plus a replayable lineage:

echo '{"record": {"id": "1", "name": "  acme   freight "},
       "pipeline": [{"name": "trim_whitespace", "params": {"field": "name"}},
                    {"name": "collapse_spaces", "params": {"field": "name"}}]}' \
  | python3.12 -m tidewaterdata clean -i -

The result.name is "acme freight", and lineage carries the per-step hashes and diffs. Save the whole {record, lineage} and verify it later:

echo '{"record": {"id": "1", "name": "  acme   freight "},
       "lineage": { ...the lineage printed above... }}' \
  | python3.12 -m tidewaterdata replay -i -

replay prints {"result": ..., "verified": true} on success, or exits 2 with a replay mismatch at step N message if the record or lineage was altered.

Run the MCP server (reads newline-delimited JSON-RPC on stdin):

python3.12 -m tidewaterdata serve

Library use:

from tidewaterdata import Pipeline, Transformation, replay

rec = {"id": "shp-1", "name": "  acme   freight ", "port": "lax"}
pipe = Pipeline([
    Transformation("trim_whitespace", {"field": "name"}),
    Transformation("collapse_spaces", {"field": "name"}),
    Transformation("to_upper", {"field": "port"}),
])
lineage, cleaned = pipe.run(rec)
assert replay(rec, lineage) == cleaned   # raises ReplayMismatch if tampered

MCP tools

The server exposes four tools via tools/list / tools/call:

Tool Arguments Returns
list_transforms none available transform names
clean_record record, pipeline {result, lineage}
replay_lineage record, lineage {result, verified}
suggest_transforms record, required_fields? {suggestions}

Validation failures and replay mismatches come back as a normal tools/call result with isError: true and a message; malformed JSON or an unknown method comes back as a JSON-RPC error.

Transforms

trim_whitespace, collapse_spaces, to_upper, to_lower, default_missing, drop_field, rename_field, coerce_int, standardize_date. Register your own with @tidewaterdata.register("name") — a transform is a pure dict -> dict function that must not mutate its input.

Configuration

All optional, read from the environment at startup and validated there (a bad value fails immediately with a configuration error):

Variable Default Meaning
TIDEWATER_PROVIDER stub stub (offline) or real (networked)
TIDEWATER_MAX_RECORD_BYTES 1000000 reject records larger than this
TIDEWATER_MAX_STEPS 100 reject pipelines longer than this
TIDEWATER_MAX_BATCH 1000 batch-size cap (reserved; see limitations)
TIDEWATER_LOG_LEVEL INFO DEBUG..CRITICAL

Config.from_file(path) overlays these defaults with keys from a JSON object.

Logs are one JSON object per line on stderr (stdout is the protocol/result stream). The size and length limits are the defence against resource exhaustion: oversized input is rejected before any work is done.

The real provider additionally reads TIDEWATER_PROVIDER_ENDPOINT and TIDEWATER_PROVIDER_API_KEY, and needs the real extra (pip install -e .[real], which pulls in httpx).

Test

make test             # or: python3.12 -m pytest

CLI exit codes: 0 success, 2 bad input / validation / missing file / replay mismatch, 1 unexpected.

Known limitations

  • No batch tool yet. TIDEWATER_MAX_BATCH is wired into config but there is no clean_batch MCP tool; records are cleaned one at a time. The cap exists so the limit is enforced the moment batching lands.
  • Replay depends on transform behaviour staying stable across versions. Changing what a registered transform does will break replay of older lineages — surfaced as a ReplayMismatch, not a silent wrong answer, but it is a real compatibility obligation. (ADR 0003)
  • 16-byte hashes detect tampering and drift, not a determined forger who recomputes the whole chain. Signing the final lineage is a future option. (ADR 0002)
  • The real provider's response contract is provisional. It assumes a {"suggestions": [...]} shape; see the TODO in providers/real.py pending a frozen endpoint spec.
  • No deployment wrapper. systemd/container packaging is out of scope for this delivery.

Tallowfield Technologies is an illustrative client; this repository is a self-directed reference implementation built to work end to end.

from github.com/J-X0/tallowfield-technologies-data-quality-mcp

Установить Tallowfield Technologies Data Quality в Claude Desktop, Claude Code, Cursor

Рекомендуется · одна команда, все IDE
unyly install tallowfield-technologies-data-quality

Ставит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.

Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh

Или настроить вручную

Выполни в терминале:

claude mcp add tallowfield-technologies-data-quality -- uvx --from git+https://github.com/J-X0/tallowfield-technologies-data-quality-mcp tidewaterdata

Пошаговые гайды: как установить Tallowfield Technologies Data Quality

FAQ

Tallowfield Technologies Data Quality MCP бесплатный?

Да, Tallowfield Technologies Data Quality MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Tallowfield Technologies Data Quality?

Нет, Tallowfield Technologies Data Quality работает без API-ключей и переменных окружения.

Tallowfield Technologies Data Quality — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Tallowfield Technologies Data Quality в Claude Desktop, Claude Code или Cursor?

Открой Tallowfield Technologies Data Quality на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Tallowfield Technologies Data Quality with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории development