Houtini Lm
UnclaimedMCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.
Install
claude mcp add houtini-lm -- npx -y @houtini/lmSet up this server
More in AI & ML
Browse the full directoryUnclaimed listing
Is this your MCP server?
This listing was auto-indexed from the public record. Claim it to edit the page, set compatibility and unlock growth tools. Takes under two minutes.
Claim this serverSecurity profile
Claimed and verified servers get a weekly static scan that shows what the code can reach: external services, environment variables, shell commands, agent configuration folders, plus any dependencies with known advisories. Claim this listing to get one. How the security profile works
8 of 8 tools
Documented tools (8)
From project documentation. A server handshake does not verify each tool’s description or behavior.
chat
The workhorse. Send a task, get an answer. The description includes planning triggers that nudge Claude to identify offloadable work when it's starting a big task.
code_task
Built for code analysis. Pre-configured system prompt with temperature and output constraints tuned per model family via the routing layer.
code_task_files
Like codetask, but the local LLM reads files directly from disk — source never passes through the MCP client's context window. Use this when reviewing multiple related files, or a single large file that's awkward to paste. Files are read in parallel with Promise.allSettled, so one unreadable file do
custom_prompt
Three-part prompt: system, context, instruction. Keeping them separate prevents context bleed - consistently outperforms stuffing everything into one message, especially with local models. I tested this properly one weekend - took the same batch of review tasks and ran them both ways. Splitting thin
discover
Health check and speed readout. Returns model name, context window, capability profile, connection latency (labelled explicitly — this is the /v1/models fetch round-trip, not inference speed), and the active model's measured tok/s and TTFT averaged over the session. Before any real call has run, mea
embed
Generate text embeddings via the OpenAI-compatible /v1/embeddings endpoint. Requires an embedding model to be available - Nomic Embed is a solid choice. Returns the vector, dimension count, and usage stats.
list_models
Lists everything on the LLM server - loaded and downloaded - with full metadata: architecture, quantisation, context window, capabilities, and HuggingFace enrichment data. Shows capability profiles describing what each model is best at, so Claude can make informed delegation decisions.
stats
Compact markdown dump of your offload stats — session and lifetime totals, per-model performance history, reasoning-token overhead — without the model catalog that discover prints. Cheap to call repeatedly to watch the 💰 counter climb.
Tool change history
FAQ
Questions about Houtini Lm MCP Server
- How do I connect Houtini Lm MCP Server to Claude?
- The listing records `claude mcp add houtini-lm -- npx -y @houtini/lm` as its setup step. Run it, then follow the repository's instructions for the client configuration; the listing names Claude Desktop, Claude Code as compatible clients.
- Is Houtini Lm MCP Server free?
- The listed licence is Apache-2.0. Check the upstream terms for permitted use and commercial requirements; a public repository does not by itself mean the software is free or open source. Connected APIs and hosted services may have separate charges.
- What can Houtini Lm MCP Server do?
- Houtini Lm MCP Server documents 8 tools to the agent, including chat, code_task, code_task_files. The descriptions above come from project documentation. A live handshake does not test individual tool behavior.