Ollama Mcp Go
FreeNot checkedMCP server for Ollama — chat, generate, embed, list models. Pure Go, zero dependencies.
About
MCP server for Ollama — chat, generate, embed, list models. Pure Go, zero dependencies.
README
MCP server for Ollama. Lets Claude Code, Cursor, or any MCP client talk to your local Ollama models.
One binary. Zero dependencies. Runs over stdio.
Why this one
There are several Ollama MCP servers out there. This one exists because:
- Go, single binary — no npm, no pip, no runtime. Download and run.
- Agent-oriented — exposes
generate(with system prompts),chat(multi-turn), andembed. No admin tools (pull/push/delete) cluttering the tool list. - Designed for agent-mesh — works as an upstream MCP server behind policy, tracing, and approval workflows. Also works standalone with any MCP client.
- Watchable — MCP delivers a tool result in one piece, so a long reply is invisible until it is done. This server streams from Ollama and mirrors every token to a trace file as it arrives, reasoning included.
tail -fand you see the model write.
Tools
| Tool | Description |
|---|---|
list_models |
List all available Ollama models |
generate |
One-shot generation with optional system prompt |
chat |
Multi-turn conversation (system/user/assistant messages) |
embed |
Generate embeddings (for semantic search, RAG, etc.) |
generate and chat stream internally and mirror their output to a trace
file — see Watching a reply as it is written.
The MCP client still receives the reply in one piece, along with the
prompt_eval_count / eval_count usage counters.
Conversation state lives entirely on the client: /api/chat is stateless, so
every call sends the whole history. Nothing accumulates server-side, and
nothing needs clearing here.
Install
From source
git clone https://github.com/KTCrisis/ollama-mcp-go.git
cd ollama-mcp-go
go build -o ollama-mcp-go .
Requires Go 1.22+ and a running Ollama instance.
Pre-built binaries
Coming soon.
Setup with Claude Code
Add to your Claude Code MCP config (~/.claude/settings.json or project .mcp.json):
{
"mcpServers": {
"ollama": {
"command": "/path/to/ollama-mcp-go"
}
}
}
If Ollama runs on a non-default host:
{
"mcpServers": {
"ollama": {
"command": "/path/to/ollama-mcp-go",
"env": {
"OLLAMA_HOST": "http://192.168.1.10:11434"
}
}
}
}
Setup with agent-mesh
mcp_servers:
- name: ollama
transport: stdio
command: /path/to/ollama-mcp-go
policies:
- name: agents
agent: "*"
rules:
- tools: ["ollama.*"]
action: allow
Usage examples
Once connected, your MCP client can call:
Generate with system prompt:
{
"name": "generate",
"arguments": {
"model": "qwen3:14b",
"prompt": "Explain service meshes in 2 sentences.",
"system": "You are a concise technical writer."
}
}
Multi-turn chat:
{
"name": "chat",
"arguments": {
"model": "llama3:8b",
"messages": [
{"role": "system", "content": "You answer in French."},
{"role": "user", "content": "What is the capital of Germany?"}
]
}
}
Embeddings:
{
"name": "embed",
"arguments": {
"model": "nomic-embed-text",
"text": "AI agent governance"
}
}
Configuration
| Env variable | Default | Description |
|---|---|---|
OLLAMA_HOST |
http://localhost:11434 |
Ollama API URL |
OLLAMA_MCP_TRACE |
/tmp/ollama-mcp-trace.log |
Where replies are mirrored as they stream. Set to off to disable. |
Watching a reply as it is written
generate and chat stream from Ollama, but MCP has no way to deliver a
partial tool result: the client receives the reply in one piece, at the end.
So the server mirrors every token to a trace file the moment it arrives.
tail -f /tmp/ollama-mcp-trace.log
Each call writes a header, the prompt, the reply as it lands, and a footer with the token counts:
=== 21:15:33 chat gpt-oss:20b ===
> Dis bonjour en trois mots exactement.
---
--- pense ---
Three words greeting: "Bonjour à tous" - Bonjour(1) à(2) tous(3)...
--- repond ---
Bonjour à tous.
[done: 75 prompt / 201 eval]
Models that emit a separate reasoning channel (gpt-oss and friends) have it
mirrored under --- pense ---. Reasoning never reaches the MCP client —
it goes to the trace and nowhere else, so the answer stays clean while the
deliberation stays visible.
Failing to open the trace file is not fatal — tracing turns itself off and generation carries on.
Protocol
Implements Model Context Protocol (MCP) over stdio transport using JSON-RPC 2.0. Compatible with protocol version 2024-11-05.
License
MIT
Installing Ollama Mcp Go
This server has no published package — it is built from source. Open the repository and follow its README.
▸ github.com/KTCrisis/ollama-mcp-goFAQ
Is Ollama Mcp Go MCP free?
Yes, Ollama Mcp Go MCP is free — one-click install via Unyly at no cost.
Does Ollama Mcp Go need an API key?
No, Ollama Mcp Go runs without API keys or environment variables.
Is Ollama Mcp Go hosted or self-hosted?
Self-hosted: the server runs locally on your machine via the install command above.
How do I install Ollama Mcp Go in Claude Desktop, Claude Code or Cursor?
Open Ollama Mcp Go on unyly.org, pick your client tab (Claude Desktop, Claude Code, Cursor) and press Install — the config is generated automatically, no JSON editing.
Related MCPs
Gmail
Read, send and search emails from Claude
by GoogleSlack
Send, search and summarize Slack messages
by SlackRunbear
No-code MCP client for team chat platforms, such as Slack, Microsoft Teams, and Discord.
Discord Server
A community discord server dedicated to MCP by [Frank Fiegel](https://github.com/punkpeye)
Compare Ollama Mcp Go with
Not sure what to pick?
Find your stack in 60 seconds
Author?
Embed badge for your README
Browse similar
All communication MCPs
