Analyst Toolkit
БесплатноНе проверенModular data QA and preprocessing toolkit — run as a Jupyter notebook pipeline, CLI, or MCP server with Docker and GCS support.
Описание
Modular data QA and preprocessing toolkit — run as a Jupyter notebook pipeline, CLI, or MCP server with Docker and GCS support.
README
Self-Healing Data Audit · Data QA + Cleaning Engine · MCP Server
🧪 Analyst Toolkit
Modular data QA and preprocessing toolkit — run as a Jupyter notebook pipeline, CLI, or MCP server with Docker and GCS support.
🆕 Version 0.5.0: MCP Platform Upgrade
This release turns the MCP server into a fuller product surface, not just a dashboard add-on.
- Ingest First:
register_inputandupload_inputestablish canonicalinput_ids, deterministic error handling, and retry-aware session reuse. - Durable Sessions:
manage_sessionplus the optional SQLite backend make session lifecycle explicit: list, inspect, fork, rebind, clear, and survive restarts when operators opt in. - Artifact-First UX:
read_artifact, stable local artifact URLs, and the localhost artifact server make dashboards and reports usable across stdio, HTTP, and containerized clients. - Safer Contracts: config normalization, certification alignment, and retry/idempotency hardening now keep MCP behavior deterministic under reruns and partial failures.
👀 MCP Ecosystem
Ship the toolkit as an MCP server and plug it into Claude Desktop, FridAI, or any JSON-RPC 2.0 client.
- ⛓️ Pipeline Mode: Chain multiple tools in memory using
session_id— no intermediate saves. - 🕹️ Executive Cockpit: Get a 0-100 Data Health Score, artifact links, and a detailed Healing Ledger.
- 📀 Golden Templates: Example templates tuned for typical fraud/migration/compliance patterns.
- 📚 Template Resources: MCP
resources/list+resources/readexpose standard and golden YAML templates directly to clients/agents. - 🤖 Auto-Heal: One-click inference and repair — from raw data to certified output in a single tool call.
- 🧭 Artifact Server: Optional local dashboard serving lets cockpit and module outputs open as stable browser links instead of raw file paths.
- 📡 MCP Server Guide — full setup, tool reference, and host integrations
TL;DR
- Modular execution by stage (diagnostics, validation, normalization, etc.)
- Inline dashboards and exportable HTML + Excel reports
- Full pipeline execution (notebook or CLI)
- YAML-configurable logic per module
- Checkpointing and joblib persistence
- MCP server — expose all toolkit modules as tools to any MCP-compatible host
- Cockpit, diagnostics, final audit, and data dictionary dashboards as standalone HTML artifacts
- 🐧 Built using synthetic data from the dirty_birds_data_generator
- 📂 Sample output (HTML dashboards, reports, plots, cleaned dataset)
📎 Resource Hub (Start Here)
- 📡 MCP Server Guide — Setup, tool reference, FridAI + Claude Desktop integration
- 🧭 Config Guide — Overview of all YAML configuration files
- 📦 Config Templates — Full set of starter YAMLs for each module (in
config/) - 📘 Usage Guide — Running the toolkit via notebooks or CLI
- 📗 Notebook Usage Guide — Full breakdown of how each module is used in notebooks
- 🤝 Contributing Guide — Development workflow, quality gates, and PR expectations
- 📝 Changelog — Versioned, deterministic release notes
📚 Quick Start Notebooks
📸 Dashboard Snapshots
![]() |
![]() |
![]() |
|
Cockpit hub, diagnostics, and final audit — three of the ten standalone HTML dashboards produced per run.
For full HTML examples instead of screenshots:
- Open the sample cockpit dashboard
- Open the sample diagnostics dashboard
- Open the sample final audit dashboard
🧰 Installation
🔧 Local Development
git clone https://github.com/G-Schumacher44/analyst_toolkit.git
cd analyst_toolkit
make install-dev # editable install + dev tooling + notebook extras
With MCP server deps
pip install -e ".[mcp]"
The mcp extra installs analyst_toolkit_deploy from a pinned GitHub Release wheel, so MCP builds do not rely on git checkouts during dependency resolution.
With notebook extras
pip install "analyst_toolkit[notebook] @ git+https://github.com/G-Schumacher44/analyst_toolkit.git"
Install from GitHub (bare)
pip install git+https://github.com/G-Schumacher44/analyst_toolkit.git
🤖 MCP Server
The toolkit ships with a built-in MCP server that exposes every module as a tool callable by any MCP-compatible host — Claude Desktop, FridAI, VS Code, or any JSON-RPC 2.0 client.
Recent MCP-facing additions include:
- real MCP resources for quickstart, agent playbook, and capability catalog
- explicit module/workflow/runtime template exposure
- standalone data dictionary generation
- optional local artifact serving for cockpit and dashboard links
Pull from GHCR:
docker pull ghcr.io/g-schumacher44/analyst-toolkit-mcp:latest
Or build and start locally:
make mcp-up # docker compose up --build -d
make mcp-health # curl /health and pretty-print response
make mcp-logs # tail logs
make mcp-down # stop
# extra runtime checks:
curl http://localhost:8001/ready | python3 -m json.tool
curl http://localhost:8001/metrics | python3 -m json.tool
Call a tool:
curl -X POST http://localhost:8001/rpc \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/call","params":{"name":"outliers","arguments":{"gcs_path":"gs://my-bucket/data/"}}}'
Tools accept a gcs_path (GCS URI, local .parquet, or local .csv) and an optional config dict matching the module's YAML structure.
| Setting | Purpose |
|---|---|
export_html: true in config (or ANALYST_REPORT_BUCKET) |
Generate HTML dashboard reports |
ANALYST_MCP_ENABLE_ARTIFACT_SERVER=true |
Serve dashboards as browser-openable links |
ANALYST_MCP_STRUCTURED_LOGS=true |
Structured request lifecycle logging |
ANALYST_MCP_AUTH_TOKEN |
Bearer token auth for networked deployments |
ANALYST_MCP_RESOURCE_TIMEOUT_SEC |
Tune template/resource read timeouts |
ANALYST_MCP_SESSION_BACKEND |
Keep session state in memory by default or opt into durable local SQLite persistence |
Deployment Profiles
Use one of these operating modes intentionally:
| Profile | Bind/Auth Posture | Intended Use |
|---|---|---|
local-dev |
loopback bind, auth token optional | local testing, Claude Desktop, local FridAI integration |
internal-trusted |
explicit non-loopback bind, bearer token strongly recommended | team/internal network use behind normal network controls |
public-or-prod |
explicit non-loopback bind, bearer token strongly recommended, docs+ops review completed | managed or internet-reachable deployment |
Notes:
- Default HTTP posture is localhost-first. Do not treat
docker-composeport publishing as a reason to skip auth. - If you set a non-loopback host, set
ANALYST_MCP_AUTH_TOKENas an operator policy. Current runtime behavior warns when the token is unset; it does not hard-fail startup. - The artifact server is also localhost-first by default and should only be widened deliberately.
ANALYST_MCP_SESSION_BACKEND=sqlitewrites durable session state to the local filesystem. By default the database lives in a private user-local state directory, not underexports/. Treat that as an explicit trust-boundary expansion: keep filesystem permissions narrow and prefer the default memory backend unless operators intentionally want restart-persistent sessions.- In stdio/local mode, run one
analyst_toolkitMCP server instance per client. Repeated reconnects can orphan extra stdio workers, which can make session visibility and input bindings look inconsistent during QA. If behavior looks split-brain, restart the client and terminate stalepython -m analyst_toolkit.mcp_server.server --stdioprocesses before retesting.
See 📡 MCP Server Guide for full setup, tool reference, FridAI integration, Claude Desktop wiring, and environment variable reference.
MCP Retry Semantics
register_inputandupload_inputare input-idempotent when you provide a stableidempotency_key.- If
load_into_session=trueand you do not provide an explicitsession_id, anonymous retries now reuse the existing live bound session for that canonical input instead of minting a new session each time. - Repeating the same-run module and dashboard calls now reuses the primary remote artifact object when it already exists. Versioned fallback object names are reserved for true first-write collisions or permission failures, not normal retries.
🧾 Configuration
Each module is controlled by a YAML file stored in config/.
Example:
validation:
input_path: "data/raw/synthetic_penguins_v3.5.csv"
schema_validation:
run: true
rules:
expected_columns: [...]
For full structure and explanation, 📘 Read the Full Configuration Guide
🧪 Usage
📓 Notebook Use (Modular)
Run each module interactively inside a Jupyter notebook.
Example
from analyst_toolkit.m02_validation.run_validation_pipeline import run_validation_pipeline
from analyst_toolkit.m00_utils.config_loader import load_config
from analyst_toolkit.m00_utils.load_data import load_csv
# --- Load config and data ---
config = load_config("config/validation_config_template.yaml")
df = load_csv("path/to/your/data.csv")
# --- Extract global settings ---
notebook_mode = config.get("notebook", True)
run_id = config.get("run_id", "demo_run")
# --- Run Validation Module ---
df_validated = run_validation_pipeline(
config=config,
df=df,
notebook=notebook_mode,
run_id=run_id
)
Modules render dashboards inline if notebook: true is set in the YAML config.
See 📗 Notebook Usage Guide for a full breakdown
📓 Notebook Use (Full Pipeline)
Run the full pipeline interactively inside a Jupyter notebook.
Example
from analyst_toolkit.run_toolkit_pipeline import run_full_pipeline
final_df = run_full_pipeline(config_path="config/run_toolkit_config.yaml")
Each module reads its own YAML config file, with optional global overrides in config/run_toolkit_config.yaml. Example:
# --- Global Run Settings ---
run_id: "CLI_2_QA"
notebook: false
# --- Pipeline Entry Point ---
pipeline_entry_path: "data/raw/synthetic_penguins_v3.5.csv"
modules:
diagnostics:
run: true
config_path: "config/diag_config_template.yaml"
validation:
run: true
config_path: "config/validation_config_template.yaml"
See 📗 Notebook Usage Guide for a full breakdown
🔁 Full Pipeline (CLI)
make pipeline # uses config/run_toolkit_config.yaml
make pipeline CONFIG=config/my_config.yaml # custom config
# or directly:
python -m analyst_toolkit.run_toolkit_pipeline --config config/run_toolkit_config.yaml
For full structure and explanation, 📘 Read the Full Usage Guide
📝 Notes from the Dev
Why build a toolkit for analysts?
I built the Analyst Toolkit to eliminate the most frustrating part of the analytics workflow — wasting hours on boilerplate cleaning when we should be exploring, validating, and learning. This system gives you:
- A one-stop first-pass QA and cleaning run, fully executable in a single notebook
- Total modularity — run stage by stage or all at once
- YAML-driven control over everything from null handling to audit thresholds
Every step leaves behind artifacts: dashboards, exports, warnings, checkpoints. You don't just run the pipeline — you see it working. You know what changed, where it changed, and what the implications are downstream. This is auditable automation — and the current dashboard set is meant to be reviewed, not just generated.
It is overbuilt in the ways that matter: transparency, reproducibility, trust. It's designed for team collaboration, for portfolio projects, for production QA. It's for your current self — and your future self — when you need to revisit a workflow six months from now.
The system is human readable and YAML-driven — for your team, your stakeholders, and yourself.
🐧 Dirty Birds: Palmer Penguins Synthetic Dataset v3.5
This toolkit is developed and tested using the Dirty Birds v3.5 dataset — a fully synthetic recreation of the Palmer Penguins dataset, purposefully enriched with ambiguity, anomalies, and missing data. The dataset is generated using penguin_synthetic_data_generator.py, a synthetic data generator that simulates viable research data and injects realistic biological variance and field collection noise for robust QA testing.
🐧 Features include:
- Categorical anomalies (typos, whitespace, & swaps)
- Numeric outliers and skew (both in error and in biological boundaries)
- Nullable fields in both wide and narrow formats
- Simulated noise to match real-world field data collection
Version Release Notes
v0.5.0 — MCP Platform Upgrade
- Input Ingest + Retry Semantics: Added
register_inputandupload_inputas first-class MCP flows with canonicalinput_ids, stable idempotency keys, and conflict-safe retry handling. - Session Lifecycle Management: Added
manage_sessionwith retention policy visibility, on-demand config inspection, fork/rebind support, and an optional SQLite-backed durable session store. - Artifact Access + Delivery: Added
read_artifact, hardened local artifact server routing, and stabilized same-run remote artifact identities on retry. - Certification + Contract Alignment:
infer_configs,validation, andfinal_auditnow align inferred rules to transformed session state instead of failing on obvious runtime drift. - Operator Hardening: Improved trust boundaries around artifact serving, SQLite state paths, input registry/session behavior, and stdio troubleshooting guidance.
v0.4.4 — Full Dashboard Rollout
- Complete Dashboard Surface: All ten pipeline modules now produce standalone exportable HTML dashboards — cockpit, diagnostics, validation, normalization, duplicates, outlier detection, outlier handling, imputation, auto-heal, and data dictionary.
- Cockpit Hub: A unified landing page links every module dashboard for a run into a single navigable session view with run metadata, artifact links, and health score.
- Dashboard Renderer Modularization:
dashboard_html.pyrefactored from a monolith into a composable renderer stack (dashboard_core,dashboard_shared, per-module renderers) — easier to review, extend, and test. - Local Artifact Server: Optional server turns cockpit and module artifact file references into browser-openable links for local development and review.
- Data Dictionary Artifacts: Standalone data dictionary generation, surfaced as an MCP tool and as an HTML artifact with column-level schema, type, and sample coverage.
- MCP Template + Resource Inventory:
resources/listandresources/readnow expose the full template and resource inventory — quickstart, playbook, capability catalog, module YAML templates — directly to MCP clients and agents. - Export Destination Routing: Configurable artifact routing with explicit GCS, local, and (stubbed) Google Drive destinations.
- Observability + Auth Hardening: Structured request lifecycle logging (
ANALYST_MCP_STRUCTURED_LOGS),/metricsand/readyendpoints, and optional bearer token auth (ANALYST_MCP_AUTH_TOKEN). - CI Hardening: Tightened lint, mypy, and test gates; CodeRabbit review workflow added to PR process.
v0.4.0 — The Cockpit Upgrade
- State Management: Introduced
StateStorefor in-memory DataFrame persistence between tool calls viasession_id. - Data Health Score: Every run now generates a weighted 0-100 score (Completeness, Validity, Uniqueness, Consistency).
- Healing Ledger: Persistent JSON/GCS history tracking every transformation made during a run.
- Golden Templates: Example templates tuned for typical fraud/migration/compliance patterns (bundled in the image under
config/golden_templates/). - Autonomous Tools: Added
auto_heal(one-click cleaning) anddrift_detection(schema/statistical comparison). - Configuration Intelligence: Added
get_config_schemato return JSON Schemas for every module.
v0.3.0
- MCP Server: New
analyst_toolkit/mcp_server/package exposes all toolkit modules as MCP tools over JSON-RPC 2.0 (HTTP/rpc) and stdio transport. - HTML Reports: All modules can emit self-contained single-page HTML reports.
- Docker / GHCR: Image published to
ghcr.io/g-schumacher44/analyst-toolkit-mcpon every push to main. - CI + Quality: GitHub Actions: ruff lint, mypy, pytest, Docker build + GHCR push on main.
v0.2.1
- Normalization · Datetime parsing: Multi-format support, strict mode,
dayfirst/yearfirst/utcoptions. - Exports · Excel date stability: Explicit date formats for cross-platform rendering.
- Duplicates · Subset-focused clusters: Dashboard now focuses on chosen
subset_columnsfor clarity.
v0.2.0
- Standardized Configuration Handling: All modules now intelligently parse their own configuration blocks.
- Simplified Module API: Runners accept the full config object — no manual unpacking needed.
v0.1.3
- Refactored Duplicates Module (M04) with correct flag/remove modes and decoupled detection logic.
v0.1.2
- Core module scaffolding complete (M01–M10), full pipeline execution, inline dashboards, joblib checkpointing.
📂 Project Structure
📦 src/ # Source root
│
├── analyst_toolkit/ # 🔧 Main toolkit package
│ ├── run_toolkit_pipeline.py # CLI + notebook entrypoint
│ │
│ ├── m00_utils/ # Shared utilities
│ │ ├── config_loader.py # YAML config loading and merging
│ │ ├── load_data.py # CSV/parquet ingestion
│ │ ├── export_utils.py # Excel + HTML export helpers
│ │ ├── report_generator.py # Self-contained HTML report builder
│ │ ├── scoring.py # Data health scoring (0-100)
│ │ ├── rendering_utils.py # Shared display/rendering helpers
│ │ ├── data_viewer.py # DataFrame preview utilities
│ │ ├── plot_viewer.py # Inline plot display
│ │ └── plot_viewer_comparison.py # Before/after comparison plots
│ │
│ ├── m01_diagnostics/ # Data profiling and structural diagnostics
│ ├── m02_validation/ # Schema validation and certification gate
│ ├── m03_normalization/ # Data cleaning and standardization
│ ├── m04_duplicates/ # Duplicate detection and removal
│ ├── m05_detect_outliers/ # Outlier detection (IQR, z-score)
│ ├── m06_outlier_handling/ # Outlier imputation or transformation
│ ├── m07_imputation/ # Missing data imputation
│ ├── m08_visuals/ # Plotting utilities and dashboard rendering
│ │ ├── comparison_plots.py # Before/after visual comparisons
│ │ ├── distributions.py # Distribution and histogram plots
│ │ └── summary_plots.py # Summary/overview charts
│ │
│ ├── m10_final_audit/ # Final audit, edits, and pipeline certification
│ │
│ └── mcp_server/ # MCP server — exposes toolkit as tools over JSON-RPC/stdio
│ ├── server.py # FastAPI /rpc dispatcher + stdio transport
│ ├── io.py # GCS/parquet/CSV data loading + report upload
│ ├── config_models.py # Pydantic models for typed config validation
│ ├── schemas.py # TypedDicts and JSON Schema for tool I/O
│ ├── registry.py # Tool self-registration and dispatch
│ ├── state.py # StateStore — in-memory session management
│ ├── templates.py # Golden template loader and resolver
│ └── tools/ # Self-registering tool modules (one per toolkit module)
│
├── 🧪 notebooks/ # Interactive tutorial notebooks (modular & full run)
│
├── ⚙️ config/ # YAML configuration files (one per module + full run)
│ └── golden_templates/ # Best-practice configs for Fraud, Migration, Compliance
│
├── 📂 data/
│ ├── raw/ # Original input datasets
│ ├── processed/ # Final certified outputs (.csv)
│ └── features/ # Optional engineered features
│
├── 📤 exports/
│ └── sample/ # Sample HTML dashboards, reports, plots, and cleaned output
│
├── tests/ # Pytest test suite (MCP smoke, unit tests)
├── resource_hub/ # Reference, guidebooks, documentation
├── Makefile # Common dev and ops commands
├── pyproject.toml # Build config and optional extras
├── environment.yaml # Conda environment definition
├── requirements-mcp.txt # MCP server pip requirements
├── Dockerfile.mcp # MCP server container
└── docker-compose.mcp.yml # Docker Compose for local MCP server
🤝 Contributing & Support
- Contributing Guide — setup, branch workflow, and quality gates
- Security Policy — responsible vulnerability disclosure process
- Bug Report Template
- Feature Request Template
- Documentation Template
- Pull Request Template
🤝 On Generative AI Use
Generative AI tools (Gemini 2.5-PRO, ChatGPT 4o - 4.1,Codex 5.4, Claude Sonnet 4.6) were used throughout this project as part of an integrated workflow — supporting code generation, documentation refinement, and idea testing. These tools accelerated development, but the logic, structure, and documentation reflect intentional, human-led design. This repository reflects a collaborative process: where automation supports clarity, and iteration deepens understanding.
📦 Licensing
This project is licensed under the MIT License.
Установить Analyst Toolkit в Claude Desktop, Claude Code, Cursor
unyly install analyst-toolkitСтавит в Claude Desktop, Claude Code, Cursor и VS Code — сам разбирается с npx, uvx и сборкой из исходников.
Впервые? Поставь CLI: curl -fsSL https://unyly.org/install | sh
Или настроить вручную
Выполни в терминале:
claude mcp add analyst-toolkit -- uvx --from git+https://github.com/G-Schumacher44/analyst_toolkit analyst_toolkitПошаговые гайды: как установить Analyst Toolkit
FAQ
Analyst Toolkit MCP бесплатный?
Да, Analyst Toolkit MCP бесплатный — установка в пару кликов через Unyly без оплаты.
Нужен ли API-ключ для Analyst Toolkit?
Нет, Analyst Toolkit работает без API-ключей и переменных окружения.
Analyst Toolkit — hosted или self-hosted?
Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.
Как установить Analyst Toolkit в Claude Desktop, Claude Code или Cursor?
Открой Analyst Toolkit на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.
Похожие MCP
GitHub
PRs, issues, code search, CI status
автор: GitHubFilesystem
Secure file operations with configurable access controls.
Memory
Knowledge graph-based persistent memory system.
Template MCP Server
A CLI tool to create a new Model Context Protocol server project with TypeScript support, dual transport options, and an extensible structure
автор: mcpdotdirectAmap Maps Mcp Server
MCP server for using the AMap Maps API
автор: duxiaohuiSupabase
Database, auth and storage
автор: SupabaseEverything
Reference / test server with prompts, resources, and tools.
Git
Tools to read, search, and manipulate Git repositories.
Sequential Thinking
Dynamic and reflective problem-solving through thought sequences.
Time
Time and timezone conversion capabilities.
Compare Analyst Toolkit with
Не уверен что выбрать?
Найди свой стек за 60 секунд
Автор?
Embed-бейдж для README
Похожее
Все в категории development



