writeonce/docs/examples/mcp-think/README.md
shoney.arickathil 49872a4b11 docs: six-bucket status board; recover lost vision doc; stop ignoring docs/
- 00-kanban.md rebuilt: ▶ NEXT PLAN pointer (iteration 4 — emitter, corpus,
  `woc build`) then six buckets — stories, in progress, done, pending,
  discarded, learnings. It tracked only the Rust runtime before, so the whole
  OOP track (wovm shipped, woc Tasks 1-8 shipped) was invisible.
- Board now records iteration 3's known gaps instead of silently owing them:
  `?T` plumbed but unenforced; E205/E201/E203/E204 dead, so structural
  interface satisfaction is unchecked.
- New discarded.md — settled rejections with reasons so they are not
  re-proposed: inheritance, `abstract` newtypes, Money/SKU/Float, Dynamic/cast/
  macro/extern, AOT-to-C, Menhir, shared engine state, external deployer.
- RECOVERED docs/plan/exploration/blue-green-vm/00-vision.md — gone from disk,
  never committed (gitignored path), cited by five docs incl. principle 12.
- Root cause was broader: all seven forward-roadmap plans in
  docs/superpowers/plans/ were untracked and ignored, on one disk only. Rules
  were half-fiction — 33 of 34 exploration files were already tracked, so they
  swallowed only *new* files.
- Dropped the docs ignore rules (exploration, oop-vm, superpowers/plans,
  examples/agent-loop, examples/mcp-think) with a do-not-re-add comment; added
  __pycache__/*.pyc. `tests/` stays ignored but warns that the next plan lands
  the corpus there.
2026-08-10 21:52:59 +02:00

92 lines
4.4 KiB
Markdown

# mcp-think — a local-model MCP tool for Claude
A minimal, working **MCP server** that Claude (Claude Code or Claude Desktop) can call as a tool — and whose work runs entirely on **your machine**, on a local model served by [Ollama](https://ollama.com). One Python file, stdio transport, standard library only (plus the official `mcp` package).
Why offload from Claude to a local model:
| Reason | Example |
| --- | --- |
| **Privacy** | `summarize` a confidential document — the text never leaves the box |
| **Cost** | `brainstorm` 20 ideas or condense a 50k-line log at zero token cost |
| **Diversity** | `critique` from a different model family — an independent second opinion |
| **Bandwidth** | Claude stays on the main task while the local model grinds a side job |
## Tools
| Tool | What it does |
| --- | --- |
| `think(task, context?)` | Reason through a problem; conclusion first, reasoning after |
| `critique(work, focus?)` | Adversarial review — strongest objections, ranked |
| `brainstorm(topic, n?)` | `n` genuinely distinct ideas |
| `summarize(text, max_words?)` | Faithful local summary of large/sensitive material |
## Prerequisites
```bash
# 1. Ollama, serving a model
ollama serve # if not already running as a service
ollama pull qwen3:4b # the default model (or any other — see Configuration)
# 2. uv (provisions the `mcp` package on the fly) — or `pip install mcp`
```
For a quick low-disk trial, `ollama pull qwen3:0.6b` (~500 MB) works — set `THINK_MODEL=qwen3:0.6b`.
## Add to Claude Code
```bash
claude mcp add think -- uv run --with mcp python3 \
/abs/path/to/docs/examples/mcp-think/server.py
```
(replace with your absolute path; add `-e THINK_MODEL=...` before `--` to override the model). Or declare it in a project's `.mcp.json`:
```json
{
"mcpServers": {
"think": {
"command": "uv",
"args": ["run", "--with", "mcp", "python3",
"/abs/path/to/docs/examples/mcp-think/server.py"],
"env": { "THINK_MODEL": "qwen3:4b" }
}
}
}
```
Claude Desktop: the same `command`/`args`/`env` block goes under `mcpServers` in `claude_desktop_config.json`.
Then just ask: *"use the think tool to weigh epoll vs io_uring for the WAL path"*, *"brainstorm 10 names for this feature"*, *"summarize this log with the local model"* — or let Claude reach for the tools on its own (the tool descriptions say when each applies).
## Configuration
| Variable | Default | Meaning |
| --- | --- | --- |
| `OLLAMA_URL` | `http://localhost:11434` | Ollama daemon base URL |
| `THINK_MODEL` | `qwen3:4b` | any pulled Ollama model |
| `THINK_TIMEOUT` | `300` | per-call timeout, seconds |
Errors (daemon down, model not pulled) come back as readable tool results, so Claude can tell you exactly what to run.
## Smoke test without Claude
```bash
python3 docs/examples/mcp-think/test_client.py # uses uv + THINK_MODEL if set
```
Drives the server over stdio exactly like Claude does — `initialize` → `tools/list` → one `think` call — and prints the local model's answer.
## How this relates to writeonce
This example is the **consumer side** of MCP: Claude ← stdio → this server → local model. The **server side** is [plan 15](../../plan/15-mcp-streamable-http.md): the writeonce runtime itself becomes an MCP server over Streamable HTTP, generating tools/resources from `.wo` declarations — after 15d, a class method like [`Product.set_price`](../pricing/types/product.wo) is callable as an MCP tool with no Python in between. The two meet naturally: a writeonce app exposes its data as MCP tools, and a local model (through a server like this one) works over it.
## Ideas to build next
Same pattern (one file, stdio, local model), different capability:
- **mcp-redact** — strip PII/secrets from text locally *before* it is sent to any cloud model.
- **mcp-embed** — local embeddings (`ollama embed`) + a small vector index over a repo or notes; gives Claude semantic search without a cloud vector DB.
- **mcp-vision** — describe screenshots/diagrams via a local vision model (`llava`, `qwen2.5-vl`) for machines where images must stay local.
- **mcp-translate** — private translation of documents.
- **mcp-testgen** — bulk-generate test fixtures/edge cases on the local model, free of token cost.
- **mcp-judge** — a local LLM-as-judge for scoring outputs in eval loops, so the judge is independent of the model being judged.