Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

VoiceStudio

БесплатноНе проверен

MCP server wrapper for VoiceStudio (debpalash): local voice cloning & TTS via OmniVoice on DGX Spark. ASMR DSP pipeline included.

GitHubEmbed

Описание

MCP server wrapper for VoiceStudio (debpalash): local voice cloning & TTS via OmniVoice on DGX Spark. ASMR DSP pipeline included.

README

MCP server wrapper for VoiceStudio (debpalash) — local voice cloning & TTS via the OmniVoice engine.

Exposes 6 tools to any MCP client (Trae IDE, Claude Desktop, etc.):

Tool What it does
clone_voice_from_audio Save a voice profile from 5-30s of reference audio + transcript
synthesize_speech Generate speech with a cloned voice OR voice design keywords (or both)
design_voice Generate speech using only voice design (no cloned voice)
list_voices List all saved voice profiles
get_voice_info Get metadata of a single voice profile
delete_voice Delete a voice profile and its ref audio

Free & local. No API costs, all inference runs on your GPU (tested on NVIDIA GB10 sm_120 / DGX Spark). Uses OmniVoice (Apache 2.0) + higgs-audio-v2-tokenizer (MIT).

Parent project

This repo is a thin MCP wrapper around debpalash/VoiceStudio (AGPL-3.0) — the upstream local voice-cloning / TTS engine. All heavy inference (torch + CUDA) runs inside VoiceStudio's venv; this repo only exposes it over MCP. The underlying model is k2-fsa/OmniVoice.


Why this exists

VoiceStudio is a great local ElevenLabs alternative (AGPL-3.0, 1.3k+ stars) but it ships as a CLI + Gradio UI. This wrapper exposes it as an MCP server so you can:

  • Call voice cloning / TTS from any MCP-compatible agent
  • Programmatically manage voice profiles (create, list, delete)
  • Reuse a single VoiceStudio venv across multiple tools without spawning Gradio

Install (DGX Spark / aarch64+CUDA)

1. Clone this repo + VoiceStudio (sibling)

mkdir -p ~/Repositories
cd ~/Repositories
git clone https://github.com/debpalash/VoiceStudio.git
git clone https://github.com/jagones84/VoiceStudio_mcp.git   VoiceStudio_mcp

2. Build VoiceStudio venv (one-time, ~5min)

Run the following steps to build VoiceStudio's venv:

cd VoiceStudio
# patch pyproject (remove aarch64 marker, pin torch 2.11+cu128, etc.)
bash /home/jagones/Repositories/trash/patch_pyproject.sh
bash /home/jagones/Repositories/trash/fix_torchvision_pin.sh
# build venv
export PATH="$HOME/.local/bin:$PATH"
rm -rf .venv uv.lock
uv lock && uv sync --python 3.12
# install torchcodec + cu12 NPP
source .venv/bin/activate
bash /home/jagones/Repositories/trash/install_torchcodec.sh
bash /home/jagones/Repositories/trash/install_nvidia_cu12.sh
bash /home/jagones/Repositories/trash/upgrade_nvidia_cu128.sh
# verify
python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"
# True NVIDIA GB10

3. Build this MCP venv (lightweight, ~30s)

cd ../VoiceStudio_mcp
cp .env.template .env
# edit .env: set HF_TOKEN (free at https://huggingface.co/settings/tokens)
export PATH="$HOME/.local/bin:$PATH"
uv sync

The MCP venv reuses the heavy torch+CUDA libs from VoiceStudio's venv via the engine wrapper (mcp_voice_studio/core/engine.py). The MCP venv only needs mcp + pydantic.

4. Test

uv run python -c "from mcp_voice_studio.server import mcp; print('tools:', [t.name for t in mcp._tool_manager._tools.values()])"

Expected:

tools: ['tool_clone_voice_from_audio', 'tool_synthesize_speech', 'tool_design_voice', 'tool_list_voices', 'tool_get_voice_info', 'tool_delete_voice']

Register in MCP clients

Trae IDE (Windows)

Add to .mcp.json (project root or ~/.trae/mcp.json):

{
  "mcpServers": {
    "voice-studio": {
      "command": "ssh",
      "args": [
        "dgx",
        "cd /home/jagones/Repositories/VoiceStudio_mcp && /home/jagones/.local/bin/uv run --no-sync python -m mcp_voice_studio"
      ],
      "env": {
        "HF_TOKEN": "hf_xxx",
        "VOICESTUDIO_VENV": "/home/jagones/Repositories/VoiceStudio/.venv"
      }
    }
  }
}

Replace dgx with your SSH host alias, and hf_xxx with your real HF token (free).

Claude Desktop

claude_desktop_config.json:

{
  "mcpServers": {
    "voice-studio": {
      "command": "ssh",
      "args": ["dgx", "cd /home/jagones/Repositories/VoiceStudio_mcp && /home/jagones/.local/bin/uv run --no-sync python -m mcp_voice_studio"]
    }
  }
}

Usage examples (from an MCP client)

Clone a voice from a 6-second sample

> Use clone_voice_from_audio to save a voice called "claudia_asmr"
> from /home/jagones/Repositories/VoiceStudio/inputs/asmr_sample.wav
> with ref_text "Hola, soy Claudia. Esta es una muestra de mi voz en estilo ASMR."

Tool response:

{
  "status": "ok",
  "voice_name": "claudia_asmr",
  "profile_path": "/home/jagones/Repositories/VoiceStudio_mcp/mcp_voice_studio/data/voices/claudia_asmr/profile.json",
  "ref_audio_path": ".../data/voices/claudia_asmr/ref_audio.wav"
}

Generate ASMR speech with the cloned voice

> Use synthesize_speech with voice_name="claudia_asmr", text="Benvenuto, chiudi gli occhi, fai un respiro profondo..."

Tool response:

{
  "output_path": "/home/jagones/Repositories/VoiceStudio_mcp/mcp_voice_studio/data/outputs/synth_1757062500.wav",
  "duration_s": 24.04,
  "sample_rate": 24000,
  "channels": 1,
  "model": "k2-fsa/OmniVoice",
  "voice_name": "claudia_asmr",
  "generation_time_s": 47.3
}

ASMR enhancements (post-synth DSP pipeline)

synthesize_speech and design_voice accept 5 optional ASMR parameters. All are applied as a deterministic post-processing pipeline on the generated mono WAV (numpy + scipy, no extra GPU). The output is stereo whenever any pan/reverb effect is active.

The pipeline runs in this order: highpass (60Hz) → lowpass → stereo_pan → reverb (with HF damping) → binaural_beat → silence_padding. The highpass is always on by default to remove DC offset and sub-bass rumble that synthetic voices can carry.

Param Type Effect ASMR sweet spot
stereo_pan center | L | R | L<->R | L->R | R->L Stereo panning law (constant-power). L<->R = alternating L/R per period_s; L->R/R->L = sawtooth sweep. L<->R with period_s=2.03.0
silence_padding_ms int (0–5000) Inserts ms of silence at every . ? ! boundary (position weighted by sentence length). 400–800 ms
reverb none | small_room | large_room Schroeder reverb (4 comb + 2 allpass filters, 18 % mix) with HF damping filter (damping=0.5 default) on each comb to avoid the classic "metallic" ring. small_room
binaural_beat_hz float (0–40) Adds a sine wave to L (200 Hz) and a slightly-detuned sine to R. Perceived as a brainwave entrainment tone. Amplitude is fixed at 0.0005 (-66dBFS, true sub-audible carrier) — only the 4–8 Hz pulsation is heard, never the 200 Hz tone itself. OFF by default to keep the output clean; pass a positive value to enable. 4–8 Hz (theta-alpha)
lowpass_cutoff_hz float (0–20000) 2nd-order Butterworth lowpass for "warmth" / intimacy. 5000–7000 Hz

Example — full ASMR stack:

> Use synthesize_speech with:
    voice_name="claudia_asmr"
    text="Ascolta il mio respiro. Lascia andare ogni tensione. Sei al sicuro."
    stereo_pan="L<->R"
    silence_padding_ms=600
    reverb="small_room"
    lowpass_cutoff_hz=6500.0

Response:

{
  "output_path": ".../synth_1757064500.wav",
  "duration_s": 8.83,
  "sample_rate": 24000,
  "channels": 2,
  "asmr_applied": [
    "lowpass(6500Hz)", "stereo_pan(L<->R)", "reverb(small_room)",
    "binaural_beat(6Hz)", "silence_padding(600ms)"
  ],
  "generation_time_s": 17.6
}

The asmr_applied list reports exactly which effects ran (so you can distinguish "nothing applied" from "applied but no audible effect"). Order in the pipeline: lowpass → pan → reverb → binaural → silence padding.

Voice design without cloning

> Use design_voice with instruct="whisper, female, low pitch", text="Hello world"

Combine cloning + design

> Use synthesize_speech with voice_name="claudia_asmr", instruct="whisper", text="..."

(cloned voice + extra style instruction)

List / inspect / delete

> list_voices
> get_voice_info(voice_name="claudia_asmr")
> delete_voice(voice_name="claudia_asmr")

Voice design keywords (OmniVoice)

Only these are accepted by OmniVoice's --instruct (case-sensitive, comma+space separated, English OR Chinese, never mix):

English: american accent, australian accent, british accent, canadian accent, child, chinese accent, elderly, female, high pitch, indian accent, japanese accent, korean accent, low pitch, male, middle-aged, moderate pitch, portuguese accent, russian accent, teenager, very high pitch, very low pitch, whisper, young adult

Chinese (full-width comma ,): 东北话,中年,中音调,云南话,低音调,儿童,四川话,女,宁夏话,少年,极低音调,极高音调,桂林话,河南话,济南话,甘肃话,男,石家庄话,老年,耳语,贵州话,陕西话,青岛话,青年,高音调

For ASMR whisper: whisper, female, low pitch


Architecture

VoiceStudio_mcp/
├── pyproject.toml                 # mcp + pydantic + numpy + scipy (no torch!)
├── mcp_voice_studio/
│   ├── server.py                  # FastMCP entry, registers 6 tools
│   ├── core/
│   │   ├── config.py              # paths, env (HF_TOKEN, VOICESTUDIO_VENV, CUDA_VISIBLE_DEVICES)
│   │   ├── models.py              # Pydantic: VoiceProfile, SynthRequest, SynthResult
│   │   ├── storage.py             # JSON+ref_audio persistence per voice profile
│   │   ├── engine.py              # auto-fallback: subprocess omnivoice-infer → inproc import
│   │   └── asmr.py                # DSP post-processor: pan, padding, reverb, binaural, lowpass
│   ├── tools/
│   │   ├── clone_voice.py         # clone_voice_from_audio
│   │   ├── synthesize.py          # synthesize_speech, design_voice
│   │   └── manage.py              # list_voices, get_voice_info, delete_voice
│   └── data/                      # gitignored runtime data
│       ├── voices/<name>/         # per-voice: ref_audio.wav, ref_text.txt, profile.json
│       ├── outputs/               # generated WAVs
│       ├── logs/                  # synthesis logs
│       └── inputs/                # default reference audio
├── tests/                         # pytest
├── examples/                      # usage examples + mcp_config.json
└── docs/                          # architecture, API

Engine wrapper: auto-fallback

mcp_voice_studio/core/engine.py tries two execution modes:

  1. Subprocess (default): spawns uv run --no-sync omnivoice-infer ... from VoiceStudio's venv.
    • Pro: isolates GPU state, most robust, no version coupling
    • Con: ~1s spawn overhead per call
  2. In-process (fallback): adds VoiceStudio's site-packages to sys.path and imports omnivoice directly.
    • Pro: faster (no spawn)
    • Con: requires VoiceStudio venv to be importable in this venv

Auto-fallback: if subprocess fails because the binary is missing, switches to in-process.

LD_LIBRARY_PATH for cu12 NPP

torchcodec (used by torchaudio 2.11) loads libnppicc.so.12. DGX Spark only has CUDA 13 system libs. The engine wrapper sets LD_LIBRARY_PATH to point at the pip-installed nvidia-npp-cu12==12.4.1.87 (from VoiceStudio venv) BEFORE the system CUDA 13 path. This avoids the TLS clash caused by symlinks.


Standalone scripts (no MCP client needed)

Two scripts under scripts/ let you run the cloning + ASMR pipeline from a terminal (useful for batch jobs or quick testing).

apply_asmr_effects.py — DSP-only on an existing WAV (no GPU, runs anywhere)

python scripts/apply_asmr_effects.py INPUT.wav OUTPUT.wav --text "..." [options]
Flag Default Effect
--stereo-pan off center/L/R/L<->R/L->R/R->L
--period-s 2.0 L<->R/L->R period in seconds
--silence-padding-ms 0 silence padding at sentence boundaries (0-5000)
--reverb off none/small_room/large_room
--reverb-damping 0.5 HF damping 0..1 (0=classic Schroeder)
--binaural-beat-hz 0.0 0=off, 4-8=theta-alpha
--binaural-amplitude 0.0005 carrier peak (default -66dBFS sub-audible)
--lowpass-cutoff-hz 0 0=off, 5000-7000 sweet spot
--highpass-cutoff-hz 60.0 0=off, 60Hz = DC/sub-bass cleanup

Example:

python scripts/apply_asmr_effects.py voice.wav out.wav \
    --text "Benvenuto. Chiudi gli occhi. Respira." \
    --stereo-pan "L<->R" --period-s 2.5 \
    --silence-padding-ms 600 \
    --reverb small_room \
    --lowpass-cutoff-hz 6500.0

clone_and_speak.py — E2E: clone/design + synthesize + ASMR (needs DGX/GPU)

# Cloned voice
python scripts/clone_and_speak.py --voice claudia_asmr \
    --text "Ascolta il mio respiro. Sei al sicuro." \
    --out out.wav \
    --stereo-pan "L<->R" --silence-padding-ms 600 --reverb small_room

# Voice design (no clone)
python scripts/clone_and_speak.py --instruct "whisper, female, low pitch" \
    --text "Hello world" --out out.wav --reverb small_room

Both scripts call the same code paths the MCP tools use (synthesize_speech for TTS, apply_asmr_pipeline for DSP), so output is identical to what you'd get from an MCP client.


License

Copyright (c) 2026 jagones84.

AGPL-3.0 — this wrapper is licensed under the GNU Affero General Public License v3.0, the same license as the upstream debpalash/VoiceStudio it drives. See LICENSE and THIRD_PARTY_LICENSES.md for the full picture.

TL;DR for publishing on GitHub:

  • This repo (VoiceStudio_mcp) is AGPL-3.0 — if you run it (or a modified version) as a network service, AGPL section 13 requires you to offer the corresponding source to your users.
  • VoiceStudio is AGPL-3.0 — installed separately via git clone, NOT bundled here. It is the upstream engine this wrapper exposes over MCP.
  • OmniVoice is Apache 2.0, higgs-audio-v2-tokenizer is MIT — both downloaded from Hugging Face at runtime, NOT bundled.

from github.com/jagones84/VoiceStudio_mcp

Установка VoiceStudio

У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.

▸ github.com/jagones84/VoiceStudio_mcp

FAQ

VoiceStudio MCP бесплатный?

Да, VoiceStudio MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для VoiceStudio?

Нет, VoiceStudio работает без API-ключей и переменных окружения.

VoiceStudio — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить VoiceStudio в Claude Desktop, Claude Code или Cursor?

Открой VoiceStudio на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

LibreOffice Tools

Enables AI agents to read, write, and edit Office documents via LibreOffice with token-efficient design. Supports multiple formats including DOCX, XLSX, PPTX, a

passerbyflutterавтор: passerbyflutter

dannote/figma-use

Full Figma control: create shapes, text, components, set styles, auto-layout, variables, export. 80+ tools.

dannoteавтор: dannote

Logo.dev

Search and retrieve company logos by brand or domain. Customize size, format, and theme to match your design needs. Accelerate design, prototyping, and content

NOVA-3951автор: NOVA-3951

Design Inspiration Server

Searches top design platforms like Dribbble and Behance to provide UI inspiration, color palettes, and layout patterns via the Serper API. It allows users to re

YonasValentinавтор: YonasValentin

PIX4Dmatic

Enables GUI automation for controlling PIX4Dmatic on Windows through MCP. Supports launching, focusing, capturing screenshots, sending hotkeys, clicking UI elem

jangjo123автор: jangjo123

Figma

Extract design specs and assets

Figmaавтор: Figma

mcp-dockmaster

An Open-Sourced UI to install and manage MCP servers for Windows, Linux and macOS.

автор: Community

ariekogan/ateam-mcp

Build, validate, and deploy multi-agent AI solutions on the ADAS platform. Design skills with tools, manage solution lifecycle, and connect from any AI environm

ariekoganавтор: ariekogan

thinkchainai/mcpbundles

MCP Bundles: Create custom bundles of tools and connect providers with OAuth or API keys. Use one MCP server across thousands of integrations, with programmatic

thinkchainaiавтор: thinkchainai

arikusi/nakkas

MCP server that turns AI into an SVG artist. One rendering engine with JSON config, AI controls all design parameters. CSS @keyframes + SMIL animations, 16+ ele

arikusiавтор: arikusi

Compare VoiceStudio with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории design