Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Open Masala

БесплатноНе проверен

Open, cited, machine-readable health reference ranges calibrated for South Asian ancestry — LOINC-coded, FHIR-shaped, every value traced to its source guideline

GitHubEmbed

Описание

Open, cited, machine-readable health reference ranges calibrated for South Asian ancestry — LOINC-coded, FHIR-shaped, every value traced to its source guideline or study.

README

The open, machine-readable, cited dataset of South Asian biomarker reference ranges and screening thresholds.

License: CC BY 4.0 DOI Hugging Face  ·  41 rows · every value cited · LOINC-coded · FHIR-shaped

South Asians develop cardiovascular disease and type-2 diabetes earlier, at lower BMI, through different metabolic pathways. The screening thresholds that reflect this — WHO Asian BMI cutoffs, IDF waist limits, earlier diabetes and coronary-calcium screening, Lp(a) prompts — exist, but they're scattered across guideline PDFs and hundreds of papers. Nobody has ever assembled them into one structured, cited, downloadable table. This is that table.

⚠️ v0 / early release. A curation and synthesis of published guidelines and studies — not a primary-data release, and not yet clinician-reviewed for clinical use. See Provenance and Disclaimer.

What this is (and isn't)

  • Is: a synthesis of published, citable thresholds into a machine-readable table — each row LOINC-coded, evidence-graded, provenance-tiered, and linked to its source.
  • Isn't: raw cohort data. The studies these numbers come from (MASALA, GenomeAsia, UK Biobank) are access-controlled for good reasons; we don't re-host them. We encode the published thresholds they produced.

Files

Path Format For
data/ancestry-reference-ranges.v0.json JSON (canonical) Everyone — source of truth, with nested sources + overclaim guards.
data/ancestry-reference-ranges.v0.csv CSV (flat) ML / data science / Hugging Face. Flattens to Parquet in one line.
data/ancestry-reference-ranges.fhir.json FHIR R4 Bundle Health-system engineers — ObservationDefinition.qualifiedInterval per row.
scripts/build.py Python (stdlib) Regenerates CSV + FHIR from the canonical JSON.
mcp-server/ MCP server Query the dataset as tool calls from any agent (Claude Desktop, Claude Code).
advisor/ Prompt + guide Run it as a personal, ancestry-aware health advisor on your own machine.
docs/schema.md Markdown Full data dictionary + FHIR/LOINC mapping.
docs/provenance.md Markdown How it was built, and how it stays current.
docs/governance.md Markdown Why we don't ship primary data; the eGFR/race lesson.

Coverage (v0.1)

  • 41 rows spanning cardiometabolic (BMI, waist, HDL, triglycerides / atherogenic pattern, ApoB, ApoB/ApoA-I ratio, Lp(a), non-HDL, remnant cholesterol, hs-CRP, HbA1c interpretation, fasting insulin, HOMA context, SCORE2 multiplier, CAC screening + South-Asian CAC percentiles, 10-year risk, eGFR), plus women's health (GDM early screening, menopause age, parity), pediatric (IAP BMI cut-offs), hepatic (BMI-23 MASLD, normal-ALT caveat), nutrition (vitamin D, vitamin B12), hematology (thalassemia carrier screening, benign ethnic neutropenia), endocrine (hypothyroidism emphasis), and pharmacogenomics (SLCO1B1 statins).
  • Provenance tiers: 18 guideline-endorsed · 19 study-derived · 4 contested-deprecated (do-not-apply showcases: eGFR race adjustment, APOL1, ALDH2/ADH1B, MTHFR-as-SA-variant).
  • Citation completeness: 100% — all 35 unique sources are backed by a curated citation with a resolvable identifier (DOI / PMID / PMCID / URL).
  • ⚠️ 20 rows added in v0.1 are review_status: proposed — pending clinician sign-off. They carry full provenance and overclaim guards but are not yet promoted to approved.

The provenance model (why you can trust a row)

Every row carries a provenance_tier — the single most important field:

Tier Meaning
guideline-endorsed A named body recommends this exact cutoff for this population (WHO Asian BMI, IDF waist, ADA SA screening). Highest confidence.
study-derived A real, cited effect estimate not yet codified into a guideline cutoff (e.g. the HbA1c under-read). Use with care; not settled practice.
contested-deprecated An adjustment the field is moving away from (the eGFR race coefficient, retired 2021). Included precisely so you can tell you must not apply it.

Every row also carries an overclaim_guard: the claim you must not make from it. Building an LLM feature? Read these — they're designed to stop confidently-wrong ancestry statements.

Quick start

import json

rows = json.load(open("data/ancestry-reference-ranges.v0.json"))["rows"]
sa_bmi = next(r for r in rows if r["id"] == "bmi-diabetes-screening-threshold")
print(sa_bmi["sa_value"], "vs general", sa_bmi["general_value"])
# {'comparator': '>=', 'value': 23} vs general {'comparator': '>=', 'value': 25}
# Regenerate the CSV + FHIR exports from the canonical JSON
python3 scripts/build.py

# CSV → Parquet, if you want it (needs pandas)
python3 -c "import pandas as pd; pd.read_csv('data/ancestry-reference-ranges.v0.csv').to_parquet('out.parquet')"

FHIR consumers: load the Bundle; each entry is an ObservationDefinition whose qualifiedInterval carries the population (appliesTo), gender, age, and numeric range. Provenance tier, evidence grade, overclaim guard, and sources ride on extensions (R4 ObservationDefinition has no native citation slot). See docs/schema.md.

Use it as a personal health advisor

Beyond the raw data, this repo ships an agent-callable layer:

  • mcp-server/ — an MCP server that exposes the dataset as tools (evaluate_value, get_reference, screening_for, …). Add it to Claude Desktop or Claude Code and ask, "I'm a South Asian man, 34, BMI 24 — what does this mean?" — you get a cited, ancestry-adjusted answer instead of a guess. It's a stateless reference oracle: it holds no user data.
  • advisor/ — system instructions that turn the MCP server into a careful, ancestry-aware health advisor that reviews your labs, honors the overclaim guards, and prepares you for a sharper conversation with your doctor. Your data stays on your machine.

How to cite

See CITATION.cff. Archived on Zenodo with a DOI: 10.5281/zenodo.21300418 (concept DOI — always resolves to the latest version). And cite the underlying primary sources — each row names its guideline/study; those authors did the original work.

Contributing

Corrections, new citations, and coverage requests are welcome — see CONTRIBUTING.md. Clinical values go through an evidence review before merge; a study-derived proposal stays flagged until it clears that bar.

License

CC-BY-4.0 — use freely, with attribution and the per-row citations preserved.

Disclaimer

Open Masala is a reference dataset for research and product development. It is not a medical device, not clinical advice, and not a substitute for a clinician's judgment. Reference ranges vary by lab and method; screening decisions require individual clinical context.

from github.com/masalahealth/open-masala

Установка Open Masala

У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.

▸ github.com/masalahealth/open-masala

FAQ

Open Masala MCP бесплатный?

Да, Open Masala MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Open Masala?

Нет, Open Masala работает без API-ключей и переменных окружения.

Open Masala — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Open Masala в Claude Desktop, Claude Code или Cursor?

Открой Open Masala на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Open Masala with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории development