Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Analyst Toolkit

БесплатноНе проверен

Modular data QA and preprocessing toolkit — run as a Jupyter notebook pipeline, CLI, or MCP server with Docker and GCS support.

GitHubEmbed

Описание

Modular data QA and preprocessing toolkit — run as a Jupyter notebook pipeline, CLI, or MCP server with Docker and GCS support.

README

Analyst Toolkit Logo
Self-Healing Data Audit  ·  Data QA + Cleaning Engine  ·  MCP Server

MIT License Status Version CI GHCR

🧪 Analyst Toolkit

Modular data QA and preprocessing toolkit — run as a Jupyter notebook pipeline, CLI, or MCP server with Docker and GCS support.

🆕 Version 0.5.0: MCP Platform Upgrade

This release turns the MCP server into a fuller product surface, not just a dashboard add-on.

  1. Ingest First: register_input and upload_input establish canonical input_ids, deterministic error handling, and retry-aware session reuse.
  2. Durable Sessions: manage_session plus the optional SQLite backend make session lifecycle explicit: list, inspect, fork, rebind, clear, and survive restarts when operators opt in.
  3. Artifact-First UX: read_artifact, stable local artifact URLs, and the localhost artifact server make dashboards and reports usable across stdio, HTTP, and containerized clients.
  4. Safer Contracts: config normalization, certification alignment, and retry/idempotency hardening now keep MCP behavior deterministic under reruns and partial failures.

👀 MCP Ecosystem

Ship the toolkit as an MCP server and plug it into Claude Desktop, FridAI, or any JSON-RPC 2.0 client.

  • ⛓️ Pipeline Mode: Chain multiple tools in memory using session_id — no intermediate saves.
  • 🕹️ Executive Cockpit: Get a 0-100 Data Health Score, artifact links, and a detailed Healing Ledger.
  • 📀 Golden Templates: Example templates tuned for typical fraud/migration/compliance patterns.
  • 📚 Template Resources: MCP resources/list + resources/read expose standard and golden YAML templates directly to clients/agents.
  • 🤖 Auto-Heal: One-click inference and repair — from raw data to certified output in a single tool call.
  • 🧭 Artifact Server: Optional local dashboard serving lets cockpit and module outputs open as stable browser links instead of raw file paths.
  • 📡 MCP Server Guide — full setup, tool reference, and host integrations

TL;DR

  • Modular execution by stage (diagnostics, validation, normalization, etc.)
  • Inline dashboards and exportable HTML + Excel reports
  • Full pipeline execution (notebook or CLI)
  • YAML-configurable logic per module
  • Checkpointing and joblib persistence
  • MCP server — expose all toolkit modules as tools to any MCP-compatible host
  • Cockpit, diagnostics, final audit, and data dictionary dashboards as standalone HTML artifacts
  • 🐧 Built using synthetic data from the dirty_birds_data_generator
  • 📂 Sample output (HTML dashboards, reports, plots, cleaned dataset)

📎 Resource Hub (Start Here)


📚 Quick Start Notebooks

Modular Demo    Pipeline Demo


📸 Dashboard Snapshots

Cockpit dashboard Diagnostics dashboard
Final audit dashboard

Cockpit hub, diagnostics, and final audit — three of the ten standalone HTML dashboards produced per run.

For full HTML examples instead of screenshots:


🧰 Installation

🔧 Local Development

git clone https://github.com/G-Schumacher44/analyst_toolkit.git
cd analyst_toolkit
make install-dev       # editable install + dev tooling + notebook extras

With MCP server deps

pip install -e ".[mcp]"

The mcp extra installs analyst_toolkit_deploy from a pinned GitHub Release wheel, so MCP builds do not rely on git checkouts during dependency resolution.

With notebook extras

pip install "analyst_toolkit[notebook] @ git+https://github.com/G-Schumacher44/analyst_toolkit.git"

Install from GitHub (bare)

pip install git+https://github.com/G-Schumacher44/analyst_toolkit.git

🤖 MCP Server

The toolkit ships with a built-in MCP server that exposes every module as a tool callable by any MCP-compatible host — Claude Desktop, FridAI, VS Code, or any JSON-RPC 2.0 client.

Recent MCP-facing additions include:

  • real MCP resources for quickstart, agent playbook, and capability catalog
  • explicit module/workflow/runtime template exposure
  • standalone data dictionary generation
  • optional local artifact serving for cockpit and dashboard links

Pull from GHCR:

docker pull ghcr.io/g-schumacher44/analyst-toolkit-mcp:latest

Or build and start locally:

make mcp-up        # docker compose up --build -d
make mcp-health    # curl /health and pretty-print response
make mcp-logs      # tail logs
make mcp-down      # stop
# extra runtime checks:
curl http://localhost:8001/ready | python3 -m json.tool
curl http://localhost:8001/metrics | python3 -m json.tool

Call a tool:

curl -X POST http://localhost:8001/rpc \
  -H "Content-Type: application/json" \
  -d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"outliers","arguments":{"gcs_path":"gs://my-bucket/data/"}}}'

Tools accept a gcs_path (GCS URI, local .parquet, or local .csv) and an optional config dict matching the module's YAML structure.

Setting Purpose
export_html: true in config (or ANALYST_REPORT_BUCKET) Generate HTML dashboard reports
ANALYST_MCP_ENABLE_ARTIFACT_SERVER=true Serve dashboards as browser-openable links
ANALYST_MCP_STRUCTURED_LOGS=true Structured request lifecycle logging
ANALYST_MCP_AUTH_TOKEN Bearer token auth for networked deployments
ANALYST_MCP_RESOURCE_TIMEOUT_SEC Tune template/resource read timeouts
ANALYST_MCP_SESSION_BACKEND Keep session state in memory by default or opt into durable local SQLite persistence

Deployment Profiles

Use one of these operating modes intentionally:

Profile Bind/Auth Posture Intended Use
local-dev loopback bind, auth token optional local testing, Claude Desktop, local FridAI integration
internal-trusted explicit non-loopback bind, bearer token strongly recommended team/internal network use behind normal network controls
public-or-prod explicit non-loopback bind, bearer token strongly recommended, docs+ops review completed managed or internet-reachable deployment

Notes:

  • Default HTTP posture is localhost-first. Do not treat docker-compose port publishing as a reason to skip auth.
  • If you set a non-loopback host, set ANALYST_MCP_AUTH_TOKEN as an operator policy. Current runtime behavior warns when the token is unset; it does not hard-fail startup.
  • The artifact server is also localhost-first by default and should only be widened deliberately.
  • ANALYST_MCP_SESSION_BACKEND=sqlite writes durable session state to the local filesystem. By default the database lives in a private user-local state directory, not under exports/. Treat that as an explicit trust-boundary expansion: keep filesystem permissions narrow and prefer the default memory backend unless operators intentionally want restart-persistent sessions.
  • In stdio/local mode, run one analyst_toolkit MCP server instance per client. Repeated reconnects can orphan extra stdio workers, which can make session visibility and input bindings look inconsistent during QA. If behavior looks split-brain, restart the client and terminate stale python -m analyst_toolkit.mcp_server.server --stdio processes before retesting.

See 📡 MCP Server Guide for full setup, tool reference, FridAI integration, Claude Desktop wiring, and environment variable reference.

MCP Retry Semantics

  • register_input and upload_input are input-idempotent when you provide a stable idempotency_key.
  • If load_into_session=true and you do not provide an explicit session_id, anonymous retries now reuse the existing live bound session for that canonical input instead of minting a new session each time.
  • Repeating the same-run module and dashboard calls now reuses the primary remote artifact object when it already exists. Versioned fallback object names are reserved for true first-write collisions or permission failures, not normal retries.

🧾 Configuration

Each module is controlled by a YAML file stored in config/.

Example:

validation:
  input_path: "data/raw/synthetic_penguins_v3.5.csv"
  schema_validation:
    run: true
    rules:
      expected_columns: [...]

For full structure and explanation, 📘 Read the Full Configuration Guide


🧪 Usage

📓 Notebook Use (Modular)

Run each module interactively inside a Jupyter notebook.

Example

from analyst_toolkit.m02_validation.run_validation_pipeline import run_validation_pipeline
from analyst_toolkit.m00_utils.config_loader import load_config
from analyst_toolkit.m00_utils.load_data import load_csv

# --- Load config and data ---
config = load_config("config/validation_config_template.yaml")
df = load_csv("path/to/your/data.csv")

# --- Extract global settings ---
notebook_mode = config.get("notebook", True)
run_id = config.get("run_id", "demo_run")

# --- Run Validation Module ---
df_validated = run_validation_pipeline(
    config=config,
    df=df,
    notebook=notebook_mode,
    run_id=run_id
)

Modules render dashboards inline if notebook: true is set in the YAML config.

See 📗 Notebook Usage Guide for a full breakdown

📓 Notebook Use (Full Pipeline)

Run the full pipeline interactively inside a Jupyter notebook.

Example

from analyst_toolkit.run_toolkit_pipeline import run_full_pipeline

final_df = run_full_pipeline(config_path="config/run_toolkit_config.yaml")

Each module reads its own YAML config file, with optional global overrides in config/run_toolkit_config.yaml. Example:

# --- Global Run Settings ---
run_id: "CLI_2_QA"
notebook: false

# --- Pipeline Entry Point ---
pipeline_entry_path: "data/raw/synthetic_penguins_v3.5.csv"

modules:
  diagnostics:
    run: true
    config_path: "config/diag_config_template.yaml"

  validation:
    run: true
    config_path: "config/validation_config_template.yaml"

See 📗 Notebook Usage Guide for a full breakdown

🔁 Full Pipeline (CLI)
make pipeline                              # uses config/run_toolkit_config.yaml
make pipeline CONFIG=config/my_config.yaml # custom config
# or directly:
python -m analyst_toolkit.run_toolkit_pipeline --config config/run_toolkit_config.yaml

For full structure and explanation, 📘 Read the Full Usage Guide


📝 Notes from the Dev

Why build a toolkit for analysts?

I built the Analyst Toolkit to eliminate the most frustrating part of the analytics workflow — wasting hours on boilerplate cleaning when we should be exploring, validating, and learning. This system gives you:

  • A one-stop first-pass QA and cleaning run, fully executable in a single notebook
  • Total modularity — run stage by stage or all at once
  • YAML-driven control over everything from null handling to audit thresholds

Every step leaves behind artifacts: dashboards, exports, warnings, checkpoints. You don't just run the pipeline — you see it working. You know what changed, where it changed, and what the implications are downstream. This is auditable automation — and the current dashboard set is meant to be reviewed, not just generated.

It is overbuilt in the ways that matter: transparency, reproducibility, trust. It's designed for team collaboration, for portfolio projects, for production QA. It's for your current self — and your future self — when you need to revisit a workflow six months from now.

The system is human readable and YAML-driven — for your team, your stakeholders, and yourself.

🐧 Dirty Birds: Palmer Penguins Synthetic Dataset v3.5

This toolkit is developed and tested using the Dirty Birds v3.5 dataset — a fully synthetic recreation of the Palmer Penguins dataset, purposefully enriched with ambiguity, anomalies, and missing data. The dataset is generated using penguin_synthetic_data_generator.py, a synthetic data generator that simulates viable research data and injects realistic biological variance and field collection noise for robust QA testing.

🐧 Features include:

  • Categorical anomalies (typos, whitespace, & swaps)
  • Numeric outliers and skew (both in error and in biological boundaries)
  • Nullable fields in both wide and narrow formats
  • Simulated noise to match real-world field data collection
🫆 Version Release Notes

v0.5.0 — MCP Platform Upgrade

  • Input Ingest + Retry Semantics: Added register_input and upload_input as first-class MCP flows with canonical input_ids, stable idempotency keys, and conflict-safe retry handling.
  • Session Lifecycle Management: Added manage_session with retention policy visibility, on-demand config inspection, fork/rebind support, and an optional SQLite-backed durable session store.
  • Artifact Access + Delivery: Added read_artifact, hardened local artifact server routing, and stabilized same-run remote artifact identities on retry.
  • Certification + Contract Alignment: infer_configs, validation, and final_audit now align inferred rules to transformed session state instead of failing on obvious runtime drift.
  • Operator Hardening: Improved trust boundaries around artifact serving, SQLite state paths, input registry/session behavior, and stdio troubleshooting guidance.

v0.4.4 — Full Dashboard Rollout

  • Complete Dashboard Surface: All ten pipeline modules now produce standalone exportable HTML dashboards — cockpit, diagnostics, validation, normalization, duplicates, outlier detection, outlier handling, imputation, auto-heal, and data dictionary.
  • Cockpit Hub: A unified landing page links every module dashboard for a run into a single navigable session view with run metadata, artifact links, and health score.
  • Dashboard Renderer Modularization: dashboard_html.py refactored from a monolith into a composable renderer stack (dashboard_core, dashboard_shared, per-module renderers) — easier to review, extend, and test.
  • Local Artifact Server: Optional server turns cockpit and module artifact file references into browser-openable links for local development and review.
  • Data Dictionary Artifacts: Standalone data dictionary generation, surfaced as an MCP tool and as an HTML artifact with column-level schema, type, and sample coverage.
  • MCP Template + Resource Inventory: resources/list and resources/read now expose the full template and resource inventory — quickstart, playbook, capability catalog, module YAML templates — directly to MCP clients and agents.
  • Export Destination Routing: Configurable artifact routing with explicit GCS, local, and (stubbed) Google Drive destinations.
  • Observability + Auth Hardening: Structured request lifecycle logging (ANALYST_MCP_STRUCTURED_LOGS), /metrics and /ready endpoints, and optional bearer token auth (ANALYST_MCP_AUTH_TOKEN).
  • CI Hardening: Tightened lint, mypy, and test gates; CodeRabbit review workflow added to PR process.

v0.4.0 — The Cockpit Upgrade

  • State Management: Introduced StateStore for in-memory DataFrame persistence between tool calls via session_id.
  • Data Health Score: Every run now generates a weighted 0-100 score (Completeness, Validity, Uniqueness, Consistency).
  • Healing Ledger: Persistent JSON/GCS history tracking every transformation made during a run.
  • Golden Templates: Example templates tuned for typical fraud/migration/compliance patterns (bundled in the image under config/golden_templates/).
  • Autonomous Tools: Added auto_heal (one-click cleaning) and drift_detection (schema/statistical comparison).
  • Configuration Intelligence: Added get_config_schema to return JSON Schemas for every module.

v0.3.0

  • MCP Server: New analyst_toolkit/mcp_server/ package exposes all toolkit modules as MCP tools over JSON-RPC 2.0 (HTTP /rpc) and stdio transport.
  • HTML Reports: All modules can emit self-contained single-page HTML reports.
  • Docker / GHCR: Image published to ghcr.io/g-schumacher44/analyst-toolkit-mcp on every push to main.
  • CI + Quality: GitHub Actions: ruff lint, mypy, pytest, Docker build + GHCR push on main.

v0.2.1

  • Normalization · Datetime parsing: Multi-format support, strict mode, dayfirst/yearfirst/utc options.
  • Exports · Excel date stability: Explicit date formats for cross-platform rendering.
  • Duplicates · Subset-focused clusters: Dashboard now focuses on chosen subset_columns for clarity.

v0.2.0

  • Standardized Configuration Handling: All modules now intelligently parse their own configuration blocks.
  • Simplified Module API: Runners accept the full config object — no manual unpacking needed.

v0.1.3

  • Refactored Duplicates Module (M04) with correct flag/remove modes and decoupled detection logic.

v0.1.2

  • Core module scaffolding complete (M01–M10), full pipeline execution, inline dashboards, joblib checkpointing.
📂 Project Structure
📦 src/                                    # Source root
│
├── analyst_toolkit/                       # 🔧 Main toolkit package
│   ├── run_toolkit_pipeline.py            # CLI + notebook entrypoint
│   │
│   ├── m00_utils/                         # Shared utilities
│   │   ├── config_loader.py               # YAML config loading and merging
│   │   ├── load_data.py                   # CSV/parquet ingestion
│   │   ├── export_utils.py                # Excel + HTML export helpers
│   │   ├── report_generator.py            # Self-contained HTML report builder
│   │   ├── scoring.py                     # Data health scoring (0-100)
│   │   ├── rendering_utils.py             # Shared display/rendering helpers
│   │   ├── data_viewer.py                 # DataFrame preview utilities
│   │   ├── plot_viewer.py                 # Inline plot display
│   │   └── plot_viewer_comparison.py      # Before/after comparison plots
│   │
│   ├── m01_diagnostics/                   # Data profiling and structural diagnostics
│   ├── m02_validation/                    # Schema validation and certification gate
│   ├── m03_normalization/                 # Data cleaning and standardization
│   ├── m04_duplicates/                    # Duplicate detection and removal
│   ├── m05_detect_outliers/               # Outlier detection (IQR, z-score)
│   ├── m06_outlier_handling/              # Outlier imputation or transformation
│   ├── m07_imputation/                    # Missing data imputation
│   ├── m08_visuals/                       # Plotting utilities and dashboard rendering
│   │   ├── comparison_plots.py            # Before/after visual comparisons
│   │   ├── distributions.py               # Distribution and histogram plots
│   │   └── summary_plots.py               # Summary/overview charts
│   │
│   ├── m10_final_audit/                   # Final audit, edits, and pipeline certification
│   │
│   └── mcp_server/                        # MCP server — exposes toolkit as tools over JSON-RPC/stdio
│       ├── server.py                      # FastAPI /rpc dispatcher + stdio transport
│       ├── io.py                          # GCS/parquet/CSV data loading + report upload
│       ├── config_models.py               # Pydantic models for typed config validation
│       ├── schemas.py                     # TypedDicts and JSON Schema for tool I/O
│       ├── registry.py                    # Tool self-registration and dispatch
│       ├── state.py                       # StateStore — in-memory session management
│       ├── templates.py                   # Golden template loader and resolver
│       └── tools/                         # Self-registering tool modules (one per toolkit module)
│
├── 🧪 notebooks/                          # Interactive tutorial notebooks (modular & full run)
│
├── ⚙️ config/                             # YAML configuration files (one per module + full run)
│   └── golden_templates/                  # Best-practice configs for Fraud, Migration, Compliance
│
├── 📂 data/
│   ├── raw/                               # Original input datasets
│   ├── processed/                         # Final certified outputs (.csv)
│   └── features/                          # Optional engineered features
│
├── 📤 exports/
│   └── sample/                            # Sample HTML dashboards, reports, plots, and cleaned output
│
├── tests/                                 # Pytest test suite (MCP smoke, unit tests)
├── resource_hub/                          # Reference, guidebooks, documentation
├── Makefile                               # Common dev and ops commands
├── pyproject.toml                         # Build config and optional extras
├── environment.yaml                       # Conda environment definition
├── requirements-mcp.txt                   # MCP server pip requirements
├── Dockerfile.mcp                         # MCP server container
└── docker-compose.mcp.yml                 # Docker Compose for local MCP server

🤝 Contributing & Support


🤝 On Generative AI Use

Generative AI tools (Gemini 2.5-PRO, ChatGPT 4o - 4.1,Codex 5.4, Claude Sonnet 4.6) were used throughout this project as part of an integrated workflow — supporting code generation, documentation refinement, and idea testing. These tools accelerated development, but the logic, structure, and documentation reflect intentional, human-led design. This repository reflects a collaborative process: where automation supports clarity, and iteration deepens understanding.


📦 Licensing

This project is licensed under the MIT License.

from github.com/G-Schumacher44/analyst_toolkit

Установить Analyst Toolkit в Claude Desktop, Claude Code, Cursor

Рекомендуется · одна команда, все IDE
unyly install analyst-toolkit

Ставит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.

Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh

Или настроить вручную

Выполни в терминале:

claude mcp add analyst-toolkit -- uvx --from git+https://github.com/G-Schumacher44/analyst_toolkit analyst_toolkit

Пошаговые гайды: как установить Analyst Toolkit

FAQ

Analyst Toolkit MCP бесплатный?

Да, Analyst Toolkit MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Analyst Toolkit?

Нет, Analyst Toolkit работает без API-ключей и переменных окружения.

Analyst Toolkit — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Analyst Toolkit в Claude Desktop, Claude Code или Cursor?

Открой Analyst Toolkit на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Analyst Toolkit with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории development