Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Hoffman2 HPC

БесплатноНе проверен

MCP server for UCLA's Hoffman2 HPC cluster providing 37 tools for SLURM job management, including submitting, monitoring, and diagnosing jobs, querying cluster

GitHubEmbed

Описание

MCP server for UCLA's Hoffman2 HPC cluster providing 37 tools for SLURM job management, including submitting, monitoring, and diagnosing jobs, querying cluster resources, and accessing documentation.

README

MCP server for UCLA's Hoffman2 HPC cluster (SLURM). Lets any MCP-compatible client (Claude Code, Cursor, Windsurf, etc.) help researchers submit, monitor, debug, and optimize SLURM jobs on Hoffman2.

Quick Start

# On a Hoffman2 login node:
git clone <repo-url> h2mcp && cd h2mcp
npm install && npm run build

Add to your MCP client config (e.g. .mcp.json):

{
  "mcpServers": {
    "hoffman2": {
      "command": "node",
      "args": ["/path/to/h2mcp/dist/index.js"]
    }
  }
}

Tools (37)

Job Management

Tool Description
job_status List running/pending jobs
job_detail Full scontrol details for a job
job_history Query completed jobs from accounting DB
job_efficiency CPU/memory efficiency metrics (seff/sacct)
cancel_job Cancel a job (ownership verified)
hold_job / release_job Hold/release pending jobs
validate_job_script Pre-submit checks: account, QOS, GPU consistency, modules, paths
submit_job Submit via sbatch with pre-flight validation
diagnose_job One-call failure analysis: job details + error output + efficiency + pattern matching
why_pending Explain why a job is stuck in queue with resource mismatch detection
interactive_session Generate srun command for user to copy/paste
run_on_compute Execute a quick command on a compute node via srun (non-interactive)

Cluster & Queue

Tool Description
cluster_status Partition/node overview
partition_info Detailed partition config
node_availability Find idle/mixed nodes with optional filters
my_partitions User's accounts, partitions, QOS
my_account_info Full account overview: accounts, jobs, quota, fairshare — one call
queue_status Available resources per partition with GPU counts and queue depth
gpu_availability GPU node states
fairshare_info Scheduler priority weighting

Job Output

Tool Description
read_output / read_error Read stdout/stderr by job ID or path
tail_output Last N lines (useful for running jobs)
search_output Grep patterns in output files

Storage

Tool Description
check_quota Home directory quota usage
scratch_status Scratch usage, recent files, purge policy
find_large_files Find files eating quota

Modules

Tool Description
module_search Search modules with dependency chains (uses modules_lookup)
module_info Module details (paths, env vars)
module_list Currently loaded modules
module_list_all All available modules with load commands

Documentation & Templates

Tool Description
sitemap Hoffman2 doc site index
fetch_doc Fetch a page from hoffman2.idre.ucla.edu (fallback for latest info)
list_templates Available .job templates
read_template Read a template (serial, gpu, mpi, array, python-conda, matlab, r, jupyter, highp)

Migration

Tool Description
convert_uge_script Convert UGE/SGE scripts to SLURM (#$ -> #SBATCH, env vars, resource mappings)

Prompts (5)

Prompt Description
debug-failed-job Diagnose a failed job
submit-first-job Guided first job submission
check-my-usage Review recent jobs and efficiency
gpu-job-setup Set up a GPU job
migrate-from-uge Convert a UGE script to SLURM

Resources (2)

Resource URI
SLURM context hoffman2://slurm/context
Documentation sitemap hoffman2://kb/sitemap.yaml

The SLURM context (partitions, QOS, GPUs, routing, storage, example scripts) is served from an external Markdown file so it can be maintained alongside the live slurm.conf. The server resolves it from $HOFFMAN2_SLURM_CONTEXT, then /u/systems/slurm/config/etc/slurm/AGENTS.md, then the bundled kb/AGENTS.md.

Job Routing

Users generally do not need to specify --partition or --qos. The cluster's job_submit.lua routes automatically:

Scenario Routed to QOS
Default (no flags) campus or pi_shared campus24 / pi_shared24
--gres=gpu:N GPU-capable partition auto
--qos=highp User's pi_* partition highp (up to 72h)

All partitions have a 24h walltime limit except pi_* with highp (72h).

Knowledge Base

kb/
├── AGENTS.md             # SLURM context fallback (canonical copy lives in the slurm config dir)
├── sitemap.yaml          # Hoffman2 website doc index
└── templates/            # Job script templates (.job)
    ├── serial.job
    ├── openmp.job
    ├── mpi.job
    ├── gpu.job
    ├── array.job
    ├── python-conda.job
    ├── matlab.job
    ├── r.job
    ├── jupyter.job
    └── highp.job

Project Structure

h2mcp/
├── src/
│   ├── index.ts              # MCP server: tool/resource/prompt registration
│   ├── tools/
│   │   ├── jobs.ts           # Job lifecycle (submit, cancel, hold, etc.)
│   │   ├── diagnose.ts       # Job failure diagnosis with error pattern matching
│   │   ├── validate.ts       # Pre-submit job script validation
│   │   ├── queue.ts          # Queue analysis (why_pending, queue_status)
│   │   ├── account.ts        # Account overview (my_account_info)
│   │   ├── interactive.ts    # Interactive sessions and run_on_compute
│   │   ├── cluster.ts        # Cluster/partition/GPU info
│   │   ├── output.ts         # Job output/error file reading
│   │   ├── storage.ts        # Quota and scratch status
│   │   ├── modules.ts        # Module search/info (uses modules_lookup)
│   │   ├── docs.ts           # Doc fetching, templates
│   │   └── migrate.ts        # UGE/SGE to SLURM converter
│   └── util/
│       └── shell.ts          # Safe command execution (timeout, input sanitization)
├── kb/                       # AGENTS.md fallback, sitemap, templates (shipped with package)
├── package.json
├── tsconfig.json
└── ARCHITECTURE.md           # Original planning/architecture document

The cluster-side job_submit.lua routing plugin and the canonical SLURM context (AGENTS.md) live in the SLURM config dir (/u/systems/slurm/config/etc/slurm/), not in this repo.

Portability

To adapt for another SLURM cluster:

  1. Edit kb/AGENTS.md (or point $HOFFMAN2_SLURM_CONTEXT at your own) — partitions, QOS, GPUs, storage, routing, examples
  2. Update the instructions block in src/index.ts — the account/routing/storage policy sent to every client
  3. Edit kb/sitemap.yaml — point to your docs
  4. Update kb/templates/ — adjust module names, account placeholders
  5. Build and run

Requirements

  • Node.js >= 20
  • SLURM cluster access (login node)
  • modules_lookup script (for module dependency resolution)

from github.com/charliecpeterson/h2mcp

Установка Hoffman2 HPC

У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.

▸ github.com/charliecpeterson/h2mcp

FAQ

Hoffman2 HPC MCP бесплатный?

Да, Hoffman2 HPC MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Hoffman2 HPC?

Нет, Hoffman2 HPC работает без API-ключей и переменных окружения.

Hoffman2 HPC — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Hoffman2 HPC в Claude Desktop, Claude Code или Cursor?

Открой Hoffman2 HPC на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Hoffman2 HPC with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории development