Command Palette

Search for a command to run...

UnylyUnyly
Browse all

VisionatorMCP

FreeNot checked

An MCP server for agents that enables models without vision/multimodal support to analyze images using local Ollama vision models.

GitHubEmbed

About

An MCP server for agents that enables models without vision/multimodal support to analyze images using local Ollama vision models.

README

An MCP (Model Context Protocol) server for GitHub Copilot (or other Agents/Harnesses that support MCP) that enables agentic models without native vision/multimodal support to analyze screenshots and images using local Ollama vision models.

Built with C# / .NET 10, the MCP C# SDK, and SkiaSharp (License) for image preprocessing.

Why ?

I needed a vision/vision-proxy MCP mcp that works without python and venv, or node. Just a simple dll or exe that can be started and it works, no matter the enviroment or system configuration.

🎯 Focus

Primary focus is Unity scene and game screenshot analysis for:

  • Visual validation and regression testing
  • Debugging rendering issues (missing textures, broken shaders)
  • UI inspection and layout verification
  • Automated QA of game screenshots

Works with any screenshots/images — not limited to Unity.

🔧 Tools

7 tools:

Tool Modes / Description
vision_analyze analyze (default): QA-focused screenshot analysis · ask: free-form prompt · describe: area-focused · detect: structured issue scan
vision_text extract (default): OCR text extraction (VLM or Tesseract) · find: locate specific text with bounding boxes
vision_validate Check if an object/element is present. Optional area param to crop to a 3×2 grid region first.
vision_compare Compare two screenshots — pixel-level diff + VLM semantic analysis
vision_color palette (default): dominant color extraction · match: percentage of pixels in custom color ranges
vision_health Verify Ollama connection, list available models, check config
vision_read_image Read a local image file and convert to base64 (convenience helper)

📋 Prerequisites

  • .NET 10 SDK
  • Ollama running locally with a vision model (e.g., llava, llama3.2-vision, minicpm-v)
  • SkiaSharp (auto-installed via NuGet) for image resizing, validation, and pixel comparison
  • (Optional) Tesseract OCR for deterministic text extraction

Install a vision model

ollama pull llava:7b
# or
ollama pull llama3.2-vision:11b
# or
ollama pull minicpm-v:8b

🚀 Quick Start

1. Build

cd VisionatorMCP
dotnet build

2. Configure in VS Code

The included .vscode/mcp.json configures the server for GitHub Copilot. Adjust environment variables as needed.

3. Run standalone (for testing)

dotnet run --project VisionatorMCP/VisionatorMCP.csproj

The server communicates via stdio — it is not a web server. It's designed to be launched by MCP clients like GitHub Copilot.

MCP Configuration:

build DLL with dotnet

{
    "servers": {
        "VisionatorMCP": {
            "type": "stdio",
            "command": "dotnet",
            "args": [
                "${workspaceFolder}/VisionatorMCP/bin/Debug/net10.0/VisionatorMCP.dll"
            ],
            "env": {
                "OLLAMA_HOST": "127.0.0.1",
                "OLLAMA_PORT": "11434",
                "VISION_MODEL": "llava:7b",
                "OCR_MODEL": "llava:7b",
                "MAX_TOKENS": "1024",
                "TEMPERATURE": "0.15",
                "REQUEST_TIMEOUT_SECONDS": "120",
                "IMAGE_AUTO_NORMALIZE": "true",
                "IMAGE_MAX_DIMENSION": "2048",
                "IMAGE_JPEG_QUALITY": "85",
                "STRIP_THINKING": "true"
            }
        }
    }
}

Self contained executable

dotnet publish VisionatorMCP/VisionatorMCP.csproj -c Release -o ./publish
{
    "servers": {
        "VisionatorMCP": {
            "type": "stdio",
            "command": "${workspaceFolder}/publish/VisionatorMCP.exe",
            "args": [],
            "env": {
                "OLLAMA_HOST": "127.0.0.1",
                "OLLAMA_PORT": "11434",
                "VISION_MODEL": "llava:7b",
                "OCR_MODEL": "llava:7b"
            }
        }
    }
}

⚙️ Configuration

All configuration is via environment variables:

Variable Default Description
OLLAMA_HOST 127.0.0.1 Ollama server host
OLLAMA_PORT 11434 Ollama server port
VISION_MODEL llava:7b Default vision model for analysis
OCR_MODEL llava:7b Default model for OCR
MAX_TOKENS 1024 Max output tokens per generation
TEMPERATURE 0.15 Generation temperature (lower = more deterministic)
REQUEST_TIMEOUT_SECONDS 120 HTTP request timeout
IMAGE_AUTO_NORMALIZE true Auto-resize & compress images before Ollama
IMAGE_MAX_DIMENSION 2048 Max pixels for longest side after resize
IMAGE_JPEG_QUALITY 85 JPEG quality for re-encoded images (1-100)
STRIP_THINKING true Strip thinking/reasoning tags from model responses
OLLAMA_AUTO_START false Auto-start Ollama if not running at MCP server start
OLLAMA_STARTUP_TIMEOUT_SECONDS 30 Max seconds to wait for Ollama to become ready

🔄 Auto-starting Ollama

VisionatorMCP can automatically start Ollama if it's not running when the MCP server starts. This is opt-in — set OLLAMA_AUTO_START=true to enable it.

How it works:

  1. On startup, the server pings Ollama at the configured host/port
  2. If Ollama is not reachable, it checks if the ollama CLI is on your PATH
  3. If found, it spawns ollama serve as a managed child process
  4. It polls every 500ms until Ollama responds (up to OLLAMA_STARTUP_TIMEOUT_SECONDS)
  5. When the MCP server shuts down, it gracefully terminates the child process

When to use this:

  • Development environments where you want a zero-config experience
  • CI/CD pipelines that need a transient Ollama instance
  • Workstations where you don't run Ollama as a system service

When NOT to use this:

  • If Ollama is already running as a system service (leave OLLAMA_AUTO_START=false)
  • If you manage Ollama via Docker or another orchestrator

Example MCP config with auto-start:

{
  "servers": {
    "VisionatorMCP": {
      "type": "stdio",
      "command": "dotnet",
      "args": [
        "${workspaceFolder}/VisionatorMCP/bin/Debug/net10.0/VisionatorMCP.dll"
      ],
      "env": {
        "OLLAMA_AUTO_START": "true",
        "OLLAMA_STARTUP_TIMEOUT_SECONDS": "45",
        "VISION_MODEL": "llava:7b"
      }
    }
  }
}

🖼️ Image Processing

VisionatorMCP uses SkiaSharp to preprocess images before sending them to Ollama, solving common issues models face with image handling:

  • Auto-resize — Large screenshots (e.g. 4K) are proportionally scaled to fit within IMAGE_MAX_DIMENSION (default 2048px). This dramatically reduces base64 payload size (often 90%+ reduction) and speeds up vision model inference.

  • Format normalization — All images are re-encoded to JPEG with configurable quality. PNG screenshots are converted to save bandwidth.

  • Validation — Base64 strings are validated as actual images before hitting Ollama. Models get clear, actionable error messages if the data is corrupt, too small, or not a supported image format.

  • Pixel-level comparisonvision_compare uses SkiaSharp for precise pixel-diff analysis (with tolerance for compression artifacts). It generates a visual diff image with changed areas highlighted in red, then optionally sends this to the VLM for semantic interpretation. This is much faster and more accurate than asking a VLM to compare two full images directly.

Disable auto-normalization by setting IMAGE_AUTO_NORMALIZE=false if you need pixel-perfect fidelity.

🧠 Thinking Tag Stripping

Many reasoning-capable vision models (DeepSeek R1, Qwen, etc.) embed "chain-of-thought" or "thinking" blocks in their responses. These consume token budget and clutter the output — often cutting off the actual useful answer.

VisionatorMCP strips these automatically by default (STRIP_THINKING=true). Supported tag formats:

Pattern Models
<think>...</think> DeepSeek R1, Qwen
<|channel|>thought...<|channel|> DeepSeek vision models
<thinking>...</thinking> Anthropic-style
[THINK]...[/THINK] Various local models

To see raw thinking output, set STRIP_THINKING=false in your MCP config or environment.

� Agent Usage Guide

When using VisionatorMCP through an AI agent (e.g., GitHub Copilot), all vision tools accept images in two ways:

Input mode Parameter Use when
File path image_path You have a local image file on disk
Base64 image_b64 You already have base64-encoded image data

If both are provided, image_path takes precedence.

Single-step workflow (recommended)

Just pass the file path directly — no manual base64 conversion needed:

vision_analyze(image_path="screenshot.png", prompt="Check for UI issues")
vision_analyze(mode="ask", image_path="screenshot.png", prompt="How many enemies?")
vision_analyze(mode="describe", image_path="screenshot.png", area="top_left")
vision_analyze(mode="detect", image_path="screenshot.png", categories="missing_textures,ui_problems")
vision_text(mode="extract", image_path="error.png", engine="vlm")
vision_text(mode="find", image_path="menu.png", query="Play")
vision_validate(query="red 'Exit' button", image_path="ui.png")
vision_validate(query="health bar", image_path="gameplay.png", area="top_left")
vision_compare(image_a_path="before.png", image_b_path="after.png", focus="UI layout")
vision_color(mode="palette", image_path="art.png")
vision_color(mode="match", image_path="art.png", ranges='[{"name":"green","min_g":100,"g_gt_r":20}]')
vision_health(test_model="llava:7b")

Two-step workflow (when you need the base64)

Use vision_read_image to get base64 data plus metadata, then pass it to other tools:

Step 1: vision_read_image(file_path="C:\\screenshots\\game.png")
        → returns { "base64": "/9j/4AAQ...", "metadata": {...} }
Step 2: vision_analyze(image_b64=<base64 from step 1>, prompt="Describe the scene")

Prompting agents to use VisionatorMCP

When instructing an AI agent to use these tools, provide clear context. Here are exemplary prompts you can use:


For Unity QA / visual regression:

You have access to VisionatorMCP tools for screenshot analysis.
Use vision_analyze(mode="detect") to scan for common Unity rendering problems
(missing textures, broken shaders, UI issues).
Then use vision_analyze(mode="analyze") for a deeper scene inspection.
Always pass image_path with the file path — never manually convert to base64.

For UI verification:

Use VisionatorMCP to verify the UI in this screenshot:
- vision_text(mode="find", query="Settings") to locate button labels or error messages
- vision_validate(query="expected element") to check if elements are present
- vision_text(mode="extract") to extract all visible text
Always use image_path=<file> rather than image_b64.

For screenshot comparison / regression testing:

Compare these two screenshots using vision_compare:
- Pass both files via image_a_path and image_b_path
- Ask it to focus on specific areas of change
- Report the pixel difference percentage and severity

For general image analysis:

Analyze this image using VisionatorMCP tools.
Use vision_analyze with a specific prompt about what to look for.
Use vision_health first if you're unsure whether Ollama is running.

Common agent workflows

Unity QA check:

vision_analyze(mode="detect", image_path="screenshot.png")
→ vision_analyze(mode="analyze", image_path="screenshot.png", prompt="Describe scene composition and visual anomalies")

UI text verification:

vision_text(mode="extract", image_path="ui.png", engine="vlm", structured=true)
→ vision_text(mode="find", image_path="ui.png", query="Settings")

Area-based analysis (3×2 grid):

Both vision_analyze and vision_validate accept an optional area parameter to crop to one of 6 equal-sized regions:

top_left    | top_center    | top_right
bottom_left | bottom_center | bottom_right
vision_analyze(mode="describe", image_path="screenshot.png", area="top_left", prompt="What UI elements are visible?")
vision_analyze(mode="describe", image_path="screenshot.png", area="bottom_center", prompt="Is there an error message?")
vision_validate(query="red 'Exit' button", area="top_right", image_path="ui.png")
vision_validate(query="health bar", area="top_left", image_path="gameplay.png")

Free-form image Q&A:

vision_analyze(mode="ask", image_path="gameplay.png", prompt="How many enemies are visible? Count and describe each.")
vision_analyze(mode="ask", image_path="screenshot.png", prompt="Is this a main menu or gameplay?")

Pixel color analysis

vision_color with mode palette extracts dominant colors — no Python/PIL scripts needed:

# Top 12 colors from the full image (default)
vision_color(image_path="screenshot.png")

# Fine-grained palette with 16-size buckets
vision_color(image_path="screenshot.png", bucket_size=16, top_n=20)

# Analyze only the bottom-left region (relative 0.0–1.0 coordinates)
vision_color(image_path="screenshot.png", region="0,0.6,0.4,1.0")

vision_color with mode match calculates what percentage of pixels match custom color ranges:

# Detect green, yellow, and brown pixel percentages
vision_color(
  mode="match",
  image_path="gameplay.png",
  ranges='[
    {"name":"green",  "min_g":100, "g_gt_r":20, "g_gt_b":20},
    {"name":"yellow", "min_r":140, "min_g":120, "max_b":110, "r_gt_b":40},
    {"name":"brown",  "min_r":90, "min_g":70, "r_gt_g":10, "r_gt_b":25}
  ]')

# Check sky region for dominant blue
vision_color(
  mode="match",
  image_path="landscape.png",
  region="0,0,1,0.3",
  ranges='[{"name":"sky_blue", "max_r":100, "max_g":180, "min_b":150}]')

📁 Project Structure

VisionatorMCP/
├── Configuration/
│   └── VisionatorConfig.cs    # Strongly-typed config
├── Services/
│   ├── ImageProcessor.cs      # SkiaSharp image preprocessing
│   └── OllamaService.cs       # Ollama HTTP API client
├── Tools/
│   ├── AnalyzeTool.cs          # vision_analyze (analyze/ask/describe/detect modes)
│   ├── TextTool.cs             # vision_text (extract/find modes)
│   ├── ValidateTool.cs         # vision_validate (optional area crop)
│   ├── CompareTool.cs          # vision_compare (pixel diff + VLM)
│   ├── ColorTool.cs            # vision_color (palette/match modes)
│   ├── HealthTool.cs           # vision_health
│   ├── ReadImageTool.cs        # vision_read_image (file→base64 helper)
│   └── ToolHelpers.cs          # Shared parsing utilities
└── Program.cs                  # MCP server entry point

📄 License

MIT

Inspired by

Vision-MCP Ocular

from github.com/Tyrannmisu/VisionatorMCP

Installing VisionatorMCP

This server has no published package — it is built from source. Open the repository and follow its README.

▸ github.com/Tyrannmisu/VisionatorMCP

FAQ

Is VisionatorMCP MCP free?

Yes, VisionatorMCP MCP is free — one-click install via Unyly at no cost.

Does VisionatorMCP need an API key?

No, VisionatorMCP runs without API keys or environment variables.

Is VisionatorMCP hosted or self-hosted?

Self-hosted: the server runs locally on your machine via the install command above.

How do I install VisionatorMCP in Claude Desktop, Claude Code or Cursor?

Open VisionatorMCP on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.

Related MCPs

Compare VisionatorMCP with

Not sure what to pick?

Find your stack in 60 seconds

Author?

Embed badge for your README

Browse similar

All ai MCPs