writeonce/docs/examples/mcp-think/README.md
shoney.arickathil 49872a4b11 docs: six-bucket status board; recover lost vision doc; stop ignoring docs/
- 00-kanban.md rebuilt: ▶ NEXT PLAN pointer (iteration 4 — emitter, corpus,
  `woc build`) then six buckets — stories, in progress, done, pending,
  discarded, learnings. It tracked only the Rust runtime before, so the whole
  OOP track (wovm shipped, woc Tasks 1-8 shipped) was invisible.
- Board now records iteration 3's known gaps instead of silently owing them:
  `?T` plumbed but unenforced; E205/E201/E203/E204 dead, so structural
  interface satisfaction is unchecked.
- New discarded.md — settled rejections with reasons so they are not
  re-proposed: inheritance, `abstract` newtypes, Money/SKU/Float, Dynamic/cast/
  macro/extern, AOT-to-C, Menhir, shared engine state, external deployer.
- RECOVERED docs/plan/exploration/blue-green-vm/00-vision.md — gone from disk,
  never committed (gitignored path), cited by five docs incl. principle 12.
- Root cause was broader: all seven forward-roadmap plans in
  docs/superpowers/plans/ were untracked and ignored, on one disk only. Rules
  were half-fiction — 33 of 34 exploration files were already tracked, so they
  swallowed only *new* files.
- Dropped the docs ignore rules (exploration, oop-vm, superpowers/plans,
  examples/agent-loop, examples/mcp-think) with a do-not-re-add comment; added
  __pycache__/*.pyc. `tests/` stays ignored but warns that the next plan lands
  the corpus there.
2026-08-10 21:52:59 +02:00

4.4 KiB

mcp-think — a local-model MCP tool for Claude

A minimal, working MCP server that Claude (Claude Code or Claude Desktop) can call as a tool — and whose work runs entirely on your machine, on a local model served by Ollama. One Python file, stdio transport, standard library only (plus the official mcp package).

Why offload from Claude to a local model:

Reason Example
Privacy summarize a confidential document — the text never leaves the box
Cost brainstorm 20 ideas or condense a 50k-line log at zero token cost
Diversity critique from a different model family — an independent second opinion
Bandwidth Claude stays on the main task while the local model grinds a side job

Tools

Tool What it does
think(task, context?) Reason through a problem; conclusion first, reasoning after
critique(work, focus?) Adversarial review — strongest objections, ranked
brainstorm(topic, n?) n genuinely distinct ideas
summarize(text, max_words?) Faithful local summary of large/sensitive material

Prerequisites

# 1. Ollama, serving a model
ollama serve                # if not already running as a service
ollama pull qwen3:4b        # the default model (or any other — see Configuration)

# 2. uv (provisions the `mcp` package on the fly) — or `pip install mcp`

For a quick low-disk trial, ollama pull qwen3:0.6b (~500 MB) works — set THINK_MODEL=qwen3:0.6b.

Add to Claude Code

claude mcp add think -- uv run --with mcp python3 \
    /abs/path/to/docs/examples/mcp-think/server.py

(replace with your absolute path; add -e THINK_MODEL=... before -- to override the model). Or declare it in a project's .mcp.json:

{
  "mcpServers": {
    "think": {
      "command": "uv",
      "args": ["run", "--with", "mcp", "python3",
               "/abs/path/to/docs/examples/mcp-think/server.py"],
      "env": { "THINK_MODEL": "qwen3:4b" }
    }
  }
}

Claude Desktop: the same command/args/env block goes under mcpServers in claude_desktop_config.json.

Then just ask: "use the think tool to weigh epoll vs io_uring for the WAL path", "brainstorm 10 names for this feature", "summarize this log with the local model" — or let Claude reach for the tools on its own (the tool descriptions say when each applies).

Configuration

Variable Default Meaning
OLLAMA_URL http://localhost:11434 Ollama daemon base URL
THINK_MODEL qwen3:4b any pulled Ollama model
THINK_TIMEOUT 300 per-call timeout, seconds

Errors (daemon down, model not pulled) come back as readable tool results, so Claude can tell you exactly what to run.

Smoke test without Claude

python3 docs/examples/mcp-think/test_client.py       # uses uv + THINK_MODEL if set

Drives the server over stdio exactly like Claude does — initialize → tools/list → one think call — and prints the local model's answer.

How this relates to writeonce

This example is the consumer side of MCP: Claude ← stdio → this server → local model. The server side is plan 15: the writeonce runtime itself becomes an MCP server over Streamable HTTP, generating tools/resources from .wo declarations — after 15d, a class method like Product.set_price is callable as an MCP tool with no Python in between. The two meet naturally: a writeonce app exposes its data as MCP tools, and a local model (through a server like this one) works over it.

Ideas to build next

Same pattern (one file, stdio, local model), different capability:

  • mcp-redact — strip PII/secrets from text locally before it is sent to any cloud model.
  • mcp-embed — local embeddings (ollama embed) + a small vector index over a repo or notes; gives Claude semantic search without a cloud vector DB.
  • mcp-vision — describe screenshots/diagrams via a local vision model (llava, qwen2.5-vl) for machines where images must stay local.
  • mcp-translate — private translation of documents.
  • mcp-testgen — bulk-generate test fixtures/edge cases on the local model, free of token cost.
  • mcp-judge — a local LLM-as-judge for scoring outputs in eval loops, so the judge is independent of the model being judged.