Command Palette

Search for a command to run...

UnylyUnyly
Весь каталог

Six Eyes

БесплатноНе проверен

MCP server that gives text-only AI agents the ability to understand images via vision tools, including multi-image analysis, OCR, comparison, and structured ext

GitHubEmbed

Описание

MCP server that gives text-only AI agents the ability to understand images via vision tools, including multi-image analysis, OCR, comparison, and structured extraction. It uses providers like OpenAI, Anthropic, Gemini, and OpenRouter to return plain text descriptions.

README

mcp-six-eyes logo

mcp-six-eyes

MCP server that gives text-only AI agents the ability to understand images, including multi-image chats like “refer image 1 and 2” or “compare these screenshots.”

Text-only models cannot see pixels. This server bridges that gap: agents call vision tools, the server talks to a multimodal API, and the agent gets plain text back.

Agent (text-only)
   │  tool call: analyze / compare / refer / ocr / …
   ▼
mcp-six-eyes (this server)
   │  1..N images: path | URL | base64  (labels: 1, 2, before, …)
   ▼
Vision API (OpenAI / Anthropic / Gemini / OpenRouter / custom)
   │
   ▼
Plain-text description / OCR / comparison / structured extract
   │
   ▼
Agent continues reasoning with text

Why this works

MCP exposes tools an agent can call. The agent never needs native vision:

  1. User uploads or points at one or more images
  2. Agent calls a vision tool with those sources (and optional labels)
  3. Server loads the image(s) and sends them to a multimodal model
  4. Server returns only text, with stable image labels
  5. The text-only agent uses that text like any other tool result

Tools

Tool Purpose
analyze_image General Q&A over one or more images
describe_image Dense scene/UI description (great “context dump” for agents)
ocr_image Extract visible text (per-image sections when multi)
compare_images Diff 2+ images (before/after, A/B, variants)
refer_images Answer questions that cite “image 1”, “both figures”, etc.
inspect_ui UI/UX screenshot review and multi-step flows
read_chart Charts, plots, tables, dashboards
explain_diagram Architecture / flowchart / ERD / whiteboard explainers
extract_from_images Structured JSON from forms, receipts, tables, labels
vision_status Show configured provider/model and limits

Image inputs

Every image tool accepts:

  • Single: image: local path, file://, http(s), data URL, or base64
  • Multi: images: array of sources or { source, label?, mimeType? } objects
  • You can pass both; they are merged

Labels default to "1", "2", … so agent prompts like “compare image 1 and 2” map cleanly. Custom labels work too ("before", "after", "fig-a").

# one image
analyze_image({ image: "./shot.png", prompt: "What failed?" })

# multi-image with default labels 1..n
compare_images({
  images: ["./a.png", "./b.png"],
  prompt: "What changed in the error state?"
})

# multi-image with explicit labels (best for long threads)
refer_images({
  images: [
    { source: "./login.png", label: "1" },
    { source: "./dashboard.png", label: "2" }
  ],
  prompt: "Using image 1 and image 2, is the user authenticated?"
})

Supported source forms:

  • local file path (/path/to/image.png or C:\path\to\image.png)
  • file:// URI
  • http(s) URL
  • data URL (data:image/png;base64,...)
  • raw base64 (pass mimeType when possible)

Requirements

  • Node.js 20+
  • A vision-capable API key (OpenAI, Anthropic, Google, OpenRouter, or any OpenAI-compatible endpoint)

Install

Published on npm as mcp-six-eyes.

npx -y mcp-six-eyes

Or install globally / as a project dependency:

npm install -g mcp-six-eyes
# or
npm install mcp-six-eyes

Most people wire it into an MCP client instead of running it by hand. Example Claude Desktop / Cursor config:

{
  "mcpServers": {
    "mcp-six-eyes": {
      "command": "npx",
      "args": ["-y", "mcp-six-eyes"],
      "env": {
        "VISION_PROVIDER": "openai",
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

Why npx is popular here:

  • no global install
  • client starts the server on demand
  • -y skips the install prompt on first run
  • npm caches the package for later launches

Local development

npm install
npm run build

Then either:

{
  "mcpServers": {
    "mcp-six-eyes": {
      "command": "npx",
      "args": ["-y", "."],
      "env": {
        "VISION_PROVIDER": "openai",
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

or point Node at the built entrypoint:

{
  "mcpServers": {
    "mcp-six-eyes": {
      "command": "node",
      "args": ["./build/index.js"],
      "env": {
        "VISION_PROVIDER": "openai",
        "OPENAI_API_KEY": "sk-..."
      }
    }
  }
}

Environment

Set provider keys in the MCP client env block (recommended) or a local .env for development.

Minimal OpenAI setup:

VISION_PROVIDER=openai
OPENAI_API_KEY=sk-...

Optional model / limits:

VISION_MODEL=gpt-4o-mini
VISION_MAX_IMAGES=10
VISION_MAX_IMAGE_BYTES=20971520

The server speaks MCP over stdio. Do not write application logs to stdout.

Client notes

Claude Desktop

Config file:

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json
  • Windows: %AppData%\Claude\claude_desktop_config.json

Use the npx block from Quick start with npx.

Cursor

Add the same server block to .cursor/mcp.json (project) or your global Cursor MCP config.

Other stdio MCP hosts

Any host that can spawn:

npx -y mcp-six-eyes

and pass environment variables will work.

Providers

Provider VISION_PROVIDER Key env var Default model
OpenAI openai OPENAI_API_KEY gpt-4o-mini
Anthropic anthropic ANTHROPIC_API_KEY claude-sonnet-4-5
Google Gemini google GOOGLE_API_KEY gemini-2.0-flash
OpenRouter openrouter OPENROUTER_API_KEY openai/gpt-4o-mini
Custom OpenAI-compatible custom VISION_API_KEY + VISION_BASE_URL set VISION_MODEL

Optional fallback:

VISION_FALLBACK_PROVIDER=anthropic
ANTHROPIC_API_KEY=sk-ant-...

Example agent usage

Single screenshot

User: What's wrong in this screenshot? ./screenshots/build-error.png

Agent → ocr_image({ image: "./screenshots/build-error.png" })
Agent → analyze_image({
  image: "./screenshots/build-error.png",
  prompt: "Explain the error and suggest a fix"
})
Agent → answers in plain text

Multi-image: refer / compare

User: I uploaded two shots. Compare image 1 and 2. Did the fix work?

Agent → compare_images({
  images: [
    { source: "./before.png", label: "1" },
    { source: "./after.png", label: "2" }
  ],
  prompt: "Did the red error banner disappear after the fix?"
})
User: Refer image 1 and image 2. Which CTA is primary?

Agent → refer_images({
  images: [
    { source: "./landing-a.png", label: "1" },
    { source: "./landing-b.png", label: "2" }
  ],
  prompt: "Which image has the stronger primary CTA and why?"
})

UI flow, chart, diagram, structured extract

inspect_ui({
  images: ["./step1.png", "./step2.png", "./step3.png"],
  prompt: "Describe the checkout flow and any friction"
})

read_chart({
  image: "https://example.com/revenue.png",
  prompt: "Summarize the trend and call out outliers"
})

explain_diagram({
  image: "./architecture.png",
  prompt: "List services and data flow"
})

extract_from_images({
  image: "./receipt.jpg",
  schema: "{\"merchant\":string,\"date\":string,\"total\":number,\"items\":[{\"name\":string,\"price\":number}]}"
})

Architecture

src/
  index.ts                 MCP server + tools
  config.ts                env/provider config
  image.ts                 path/URL/base64 loader + multi-image labels
  prompts.ts               task prompts (analyze/describe/ocr/compare/...)
  providers/
    index.ts               provider router + fallback
    openai-compatible.ts   OpenAI / OpenRouter / custom (multi-image)
    anthropic.ts           Claude vision (multi-image)
    google.ts              Gemini vision (multi-image)
    types.ts               shared contracts
test/                      unit tests (node:test, mocked providers)
assets/
  logo.png                 project logo

Design notes

  • Tools, not resources: image understanding is an action with side effects (API cost), so it is exposed as tools.
  • Text-only output: host models without vision only need text content blocks.
  • Labeled multi-image: agents in chat UIs talk about “image 1/2”; labels keep that grounding stable.
  • Task-specific tools: compare / refer / UI / chart / diagram / extract beat one mega-prompt for tool selection.
  • Stdio transport: simplest local integration for desktop agents.
  • No stdout logging: stdout is reserved for JSON-RPC; diagnostics go to stderr.
  • Provider abstraction: swap backends without changing tool names the agent learns.

Development

npm install
npm test
npm start
Script Purpose
npm run build Compile TypeScript to build/
npm run typecheck Typecheck only
npm test Build + full unit test suite
npm run test:unit Run tests against current build/
npm run smoke Quick image-loader smoke script
npm start Run MCP server on stdio

Debug with the MCP Inspector:

npx @modelcontextprotocol/inspector node ./build/index.js

See CONTRIBUTING.md for PR and coding guidelines.

Links

Release workflow

Maintainer path after local changes:

# one-time
npm login

# bump version + CHANGELOG, then ship
npm test
npm publish --access public

Optional helper (tests, then npm publish):

npm run release

Security

  • API keys stay in environment variables / client config, never in tool responses
  • Remote URL fetches are explicit tool inputs; treat untrusted URLs carefully
  • Large images are rejected via VISION_MAX_IMAGE_BYTES (default 20MB)
  • Image count per call is capped via VISION_MAX_IMAGES (default 10)

Full policy: SECURITY.md.

Contributing

Issues and pull requests are welcome. Please run npm test before opening a PR and read CONTRIBUTING.md.

License

MIT

from github.com/RimunAce/mcp-six-eyes

Установка Six Eyes

У этого сервера нет опубликованного пакета — он собирается из исходников. Открой репозиторий и следуй инструкции в README.

▸ github.com/RimunAce/mcp-six-eyes

FAQ

Six Eyes MCP бесплатный?

Да, Six Eyes MCP бесплатный — установка в пару кликов через Unyly без оплаты.

Нужен ли API-ключ для Six Eyes?

Нет, Six Eyes работает без API-ключей и переменных окружения.

Six Eyes — hosted или self-hosted?

Self-hosted: сервер запускается локально на твоей машине командой из раздела установки.

Как установить Six Eyes в Claude Desktop, Claude Code или Cursor?

Открой Six Eyes на unyly.org, выбери вкладку своего клиента (Claude Desktop, Claude Code, Cursor) и нажми Install — конфиг сгенерируется автоматически, без правки JSON.

Похожие MCP

Compare Six Eyes with

Не уверен что выбрать?

Найди свой стек за 60 секунд

Автор?

Embed-бейдж для README

Похожее

Все в категории media