Ai Media
FreeNot checkedEnables image and video generation across GPT-Image, Gemini, Grok, and Jimeng with file-based outputs, multi-reference support, and model capability lookup.
About
Enables image and video generation across GPT-Image, Gemini, Grok, and Jimeng with file-based outputs, multi-reference support, and model capability lookup.
README
A single-file, zero-dependency MCP server for image & video generation across GPT-Image / Gemini / Grok / Jimeng, with zero base64 — everything is saved to disk and only file paths are returned.
Node.js 18+ required. No
npm install, no build step. Just run the.mjs.
✨ Features
- 🖼 Image generation —
generate_image: one OpenAI-compatible path, model name pass-through (gpt-image-2/gemini-3-pro-image/grok-imagine-image-quality/doubao-seedream/ Jimeng image models …) - 🎬 Video generation —
generate_video(submit only) /get_video_status(poll) /generate_video_and_wait(submit → poll → download). Auto-routes Grok video & Jimeng video. - 🔑 Multi-key routing — different models automatically use their own API key & base URL (common with relay/aggregator gateways).
- 🚫 Zero base64 — images/videos are saved to local disk; responses only contain
Saved to: ...(~350 bytes vs ~3 MB with typical MCP wrappers). - 📚 Capability lookup —
list_model_capabilities: query supported resolutions / aspect ratios / durations per model; reverse-lookup (size: "720p",aspect_ratio: "21:9"). Zero cost. - 🖼 Multi-reference images (image-to-image) —
imagesarray (up to 5): Gemini native multi-image via multipleinlineData; GPT/Grok arrays with automaticedits → generationsfallback. The requested model is preserved; GPT, Gemini, and Grok have each passed two-reference gateway tests. - 🎬 Reference-image video (image-to-video) — single
image(Grokimage_url, Jimengfirst_frame_url) or multipleimages(Jimengreference_image_urlsup to 9, Grokimage_urlarray fallback). - ✍️ Auto prompt enhancement —
auto_prompt: true(default): if the prompt is < 15 chars and reference images are present, an internal "seamless fusion" template is applied (lighting / perspective / color grading / anti-cutout instructions verified on real generations). - 🧪
@chatbox-latest— passimage: "@chatbox-latest"to auto-read the most recent image dragged into the Chatbox window (chatbox-blobs\pictureinput-*). Repeated aliases inimagespreserve the UI attachment order. - 🔎 Verifiable references — every generated image reports ordered
reference_imagesmetadata derived from the exact bytes sent upstream: safe source label, MIME, byte count, dimensions, and SHA-256. - 📥 Sandbox delivery — callers can pass their writable sandbox as
output_dir, keepreturn_mode: "path", and turn returned files into download cards without moving files from unrelated external directories.
📦 Install
# Just clone/copy the mjs and run it directly
node ai-media-mcp.mjs
No npm dependencies. Add it to your MCP client (e.g. Chatbox / Claude Desktop) as a stdio server:
{
"command": "node",
"args": ["/path/to/ai-media-mcp.mjs"],
"env": { "...": "see .env.example" }
}
⚙️ Configuration
All variables are optional. A group falls back to the unified AI_MEDIA_API_KEY / AI_MEDIA_BASE_URL only when its group variable is absent. If a group key is explicitly present but empty (for example AI_MEDIA_JIMENG_API_KEY=), that group is disabled and will not reuse another provider's key.
| Variable | Group | Description |
|---|---|---|
AI_MEDIA_GPT_API_KEY / AI_MEDIA_GPT_BASE_URL |
GPT | gpt-image-2 / doubao-* / Jimeng image |
AI_MEDIA_GEMINI_API_KEY / AI_MEDIA_GEMINI_BASE_URL |
Gemini | native Gemini API (/v1beta/models/...:generateContent) |
AI_MEDIA_GROK_API_KEY / AI_MEDIA_GROK_BASE_URL |
Grok | grok-imagine-image-quality / grok-imagine-video |
AI_MEDIA_JIMENG_API_KEY / AI_MEDIA_JIMENG_BASE_URL |
Jimeng | as-sd2.0-fast / video-ds-2.0 |
AI_MEDIA_API_KEY / AI_MEDIA_BASE_URL |
unified | fallback for all groups |
AI_MEDIA_IMAGE_MODEL |
— | default image model (gpt-image-2) |
AI_MEDIA_VIDEO_MODEL |
— | default video model (grok-imagine-video) |
AI_MEDIA_IMAGE_OUTPUT_DIR / AI_MEDIA_VIDEO_OUTPUT_DIR |
— | output dirs |
AI_MEDIA_RETURN_MODE |
— | path (default, zero base64) / inline / both |
AI_MEDIA_AUTO_PROMPT |
— | auto (default) / never |
AI_MEDIA_TIMEOUT_MS / AI_MEDIA_POLL_INTERVAL_MS |
— | timeouts |
🧭 Routing
By model name prefix:
| Model prefix | Group | API |
|---|---|---|
gemini-* / imagen-* |
Gemini | native /v1beta/models/{model}:generateContent |
grok-* |
Grok | OpenAI-compatible |
video-ds* / as-sd* |
Jimeng | task-based /videos |
others (gpt-*, doubao-*, Jimeng image) |
GPT | OpenAI-compatible |
🛠 Tools
| Tool | Purpose |
|---|---|
generate_image |
Generate images (n, size/aspect_ratio/quality, image/images references, auto_prompt) |
generate_video |
Submit a video task only (Grok/Jimeng auto-route, image/images supported) |
get_video_status |
Query task status |
generate_video_and_wait |
Submit → poll → download to output_dir (or the configured default), return file path |
list_model_capabilities |
Query model capabilities (resolutions/ratios/durations), reverse-lookup by size or ratio |
probe_capabilities |
Probe actual supported sizes/ratios via minimal real calls, cache results |
📄 License
MIT
Install Ai Media in Claude Desktop, Claude Code & Cursor
unyly install ai-media-mcpInstalls into Claude Desktop, Claude Code, Cursor & VS Code — handles npx, uvx and build-from-source repos for you.
First time? Get the CLI: curl -fsSL https://unyly.org/install | sh
Or configure manually
Run in your terminal:
claude mcp add ai-media-mcp -- npx -y github:wojiaopanhaoran/ai-media-mcpStep-by-step: how to install Ai Media
FAQ
Is Ai Media MCP free?
Yes, Ai Media MCP is free — one-click install via Unyly at no cost.
Does Ai Media need an API key?
No, Ai Media runs without API keys or environment variables.
Is Ai Media hosted or self-hosted?
A hosted option is available: Unyly runs the server in the cloud, no local setup required.
How do I install Ai Media in Claude Desktop, Claude Code or Cursor?
Open Ai Media on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.
Related MCPs
ARA
Generate images, video and audio from any AI agent — one connector.
by ARAOmni Video
An MCP server that transforms LLM-enabled IDEs into professional video editors by pre-processing footage into text proxies, generating motion graphics via HTML/
by buildwithtazaYouTube
Transcripts, channel stats, search
by YouTubeEverArt
AI image generation using various models.
by modelcontextprotocolgpu-bridge/mcp-server
Unified GPU inference API with 30 AI services (LLM, image gen, video, TTS, whisper, embeddings, reranking, OCR) as MCP tools. Pay-per-use via x402 USDC or API k
by gpu-bridgehamflx/imagen3-mcp
A powerful image generation tool using Google's Imagen 3.0 API through MCP. Generate high-quality images from text prompts with advanced photography, artistic,
by hamflxmerterbak/Grok-MCP
MCP server for xAI's [Grok API](https://docs.x.ai/docs/overview) with agentic tool calling, image generation, vision, and file support.
by merterbakSureScaleAI/openai-gpt-image-mcp
OpenAI GPT image generation/editing MCP server.
by SureScaleAIYangLiangwei/PersonalizationMCP
Comprehensive personal data aggregation MCP server with Steam, YouTube, Bilibili, Spotify, Reddit and other platforms integrations. Features OAuth2 authenticati
by YangLiangweiAceDataCloud/MCPFlux
Flux AI image generation and editing (Black Forest Labs) via Ace Data Cloud API.
by AceDataCloudCompare Ai Media with
Not sure what to pick?
Find your stack in 60 seconds
Author?
Embed badge for your README
Browse similar
All media MCPs
