About
MCP server for AI image generation via Microsoft Copilot
README
An MCP server that connects GitHub Copilot CLI to Microsoft Copilot's image generation backend.
Generate images from text prompts and iteratively refine them — all from the terminal. Images are saved locally and organized by session.
Requirements
- Python 3.10+
- GitHub Copilot CLI
- Microsoft 365 Copilot license with image generation enabled
- A Chromium-based browser (Edge, Chrome, or Chromium) for first-time sign-in
Quick Start
Install and register
pip install copilot-image-gen-mcp copilot plugin install msartem/copilot_image_gen_mcpLaunch Copilot CLI
copilotSign in (first time only)
Copilot will automatically call
sign_inwhen needed. An Edge browser window opens — sign in with your Microsoft 365 account. The auth code is captured automatically via Playwright (no manual copy-paste).Try it out
Generate an image of an elephant in Times Square Make the elephant golden Change the background to a sunset over mountains
After first sign-in, auth is silent via cached refresh tokens (~90 day lifetime, auto-renewing). No browser window on subsequent uses.
How It Works
- You ask Copilot CLI to generate or modify an image
- Copilot routes it to the
generate_imageorrefine_imagetool - The MCP server opens a WebSocket to Microsoft Copilot's backend
- The backend generates the image (DALL-E) and returns it as base64 PNG
- The image is saved to
~/.copilot-images/and the file path is returned
Multi-turn refinement works automatically — the server maintains a conversation
so each refine_image call builds on the previous image.
You → Copilot CLI → image gen MCP → M365 Copilot (Sydney) → DALL-E → PNG saved locally
Tools
| Tool | Description |
|---|---|
generate_image(prompt, orientation) |
Generate a new image from text. Blocks ~15-30s, returns file path. |
refine_image(prompt) |
Modify the last generated image. Same conversation context. |
sign_in() |
Sign in to Microsoft 365 (opens browser, one-time). |
new_session() |
Start a fresh conversation (discard previous image context). |
Orientation Options
landscape(default)portraitsquare
Image Storage
Images are saved to ~/.copilot-images/ organized by session. Each session
groups an initial image with its refinements:
~/.copilot-images/
├── 20250115_143022_elephant_in_times_square/
│ ├── session.json ← metadata (prompts, timestamps)
│ ├── 001_elephant_in_times_square.png ← initial image
│ ├── 002_make_the_elephant_golden.png ← first refinement
│ └── 003_change_the_background_to_a_sunset.png ← second refinement
└── 20250116_091200_sunset_over_mountains/
├── session.json
└── 001_sunset_over_mountains.png
Session directories are created lazily on first image save — no side effects until you actually generate something.
Authentication
Authentication uses Playwright to automate browser sign-in via Microsoft Edge. This works identically on macOS and Windows — no platform-specific code.
First sign-in
- Copilot calls
sign_in(or you trigger it manually) - An Edge window opens to the Microsoft sign-in page
- Sign in with your M365 account (SSO may auto-complete this)
- The auth code is captured automatically — the browser closes
- Tokens are cached locally
Subsequent runs
Cached refresh tokens are used — no browser window, no interaction needed.
Browser selection
Playwright uses Microsoft Edge by default. To use a different Chromium-based browser:
export COPILOT_BROWSER=chrome # or: msedge (default), chromium
Note:
chromiumrequiresplaywright install chromium. Edge and Chrome use your installed browser directly — no extra install needed.
Manual auth
python auth.py # Interactive sign-in
python auth.py logout # Clear cached tokens
Token storage
| Platform | Cache location |
|---|---|
| macOS / Linux | ~/.copilot-image-gen-mcp/token_cache.json |
| Windows | %LOCALAPPDATA%\copilot-image-gen-mcp\token_cache.json |
Environment Variables
| Variable | Default | Description |
|---|---|---|
COPILOT_IMAGES_DIR |
~/.copilot-images |
Image output directory |
COPILOT_BROWSER |
msedge |
Browser for sign-in: msedge, chrome, chromium |
COPILOT_TENANT |
common |
Azure AD tenant ID |
COPILOT_TIMEOUT |
90 |
Image generation timeout in seconds |
COPILOT_VARIANTS |
(built-in) | Feature flag overrides (advanced) |
Technical Details
See TECHNICAL.md for details on the SignalR WebSocket protocol, authentication flow, image delivery format, and multi-turn refinement architecture.
Disclaimer
This is an independent community project — not affiliated with or supported by Microsoft.
It works by communicating with the same undocumented web APIs that power the M365 Copilot web/desktop app image generation. These APIs may change or break without notice. Use at your own risk.
Installing Copilot Image Gen
This server has no published package — it is built from source. Open the repository and follow its README.
▸ github.com/msartem/copilot_image_gen_mcpFAQ
Is Copilot Image Gen MCP free?
Yes, Copilot Image Gen MCP is free — one-click install via Unyly at no cost.
Does Copilot Image Gen need an API key?
No, Copilot Image Gen runs without API keys or environment variables.
Is Copilot Image Gen hosted or self-hosted?
Self-hosted: the server runs locally on your machine via the install command above.
How do I install Copilot Image Gen in Claude Desktop, Claude Code or Cursor?
Open Copilot Image Gen on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.
Related MCPs
ARA
Generate images, video and audio from any AI agent — one connector.
by ARAOmni Video
An MCP server that transforms LLM-enabled IDEs into professional video editors by pre-processing footage into text proxies, generating motion graphics via HTML/
by buildwithtazaYouTube
Transcripts, channel stats, search
by YouTubeEverArt
AI image generation using various models.
by modelcontextprotocolCompare Copilot Image Gen with
Not sure what to pick?
Find your stack in 60 seconds
Author?
Embed badge for your README
Browse similar
All media MCPs
