Houtini Lm logo

Houtini Lm

未认领

作者:houtini-ai

MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.

ai-agentsclaudeclaude-mcpcode-generationdeveloper-tooldeveloper-toolslm-studiolm-studio-mcplocal-llmmcpmcp-clientmcp-servermodel-context-protocolollamaopenai-apitoken-savingsvibe-coding-assistantvllm

安装

$claude mcp add houtini-lm -- npx -y @houtini/lm

此服务器需要专门配置。目前没有可复用的公开启动命令,请遵循项目说明。

项目说明

更多AI & ML服务器

浏览完整目录

未认领列表

这个 MCP 服务器是你的吗?

此列表根据公开信息自动生成。认领后,你可以编辑页面、设置兼容性并解锁增长工具。全程不超两分钟。

认领此服务器

安全概况

已认领和已认证的服务器每周接受一次静态扫描,显示代码能触及的范围(外部服务、环境变量、shell 命令、代理配置目录)以及带有已知公告的依赖。认领此列表即可获得。 安全概况的工作原理

8 个工具中显示 8 个

文档中列出的工具 (8)

内容来自项目文档。服务器握手不会验证每个工具的说明或行为。

chat

The workhorse. Send a task, get an answer. The description includes planning triggers that nudge Claude to identify offloadable work when it's starting a big task.

code_task

Built for code analysis. Pre-configured system prompt with temperature and output constraints tuned per model family via the routing layer.

code_task_files

Like codetask, but the local LLM reads files directly from disk — source never passes through the MCP client's context window. Use this when reviewing multiple related files, or a single large file that's awkward to paste. Files are read in parallel with Promise.allSettled, so one unreadable file do

custom_prompt

Three-part prompt: system, context, instruction. Keeping them separate prevents context bleed - consistently outperforms stuffing everything into one message, especially with local models. I tested this properly one weekend - took the same batch of review tasks and ran them both ways. Splitting thin

discover

Health check and speed readout. Returns model name, context window, capability profile, connection latency (labelled explicitly — this is the /v1/models fetch round-trip, not inference speed), and the active model's measured tok/s and TTFT averaged over the session. Before any real call has run, mea

embed

Generate text embeddings via the OpenAI-compatible /v1/embeddings endpoint. Requires an embedding model to be available - Nomic Embed is a solid choice. Returns the vector, dimension count, and usage stats.

list_models

Lists everything on the LLM server - loaded and downloaded - with full metadata: architecture, quantisation, context window, capabilities, and HuggingFace enrichment data. Shows capability profiles describing what each model is best at, so Claude can make informed delegation decisions.

stats

Compact markdown dump of your offload stats — session and lifetime totals, per-model performance history, reasoning-token overhead — without the model catalog that discover prints. Cheap to call repeatedly to watch the 💰 counter climb.

工具变更历史

比较相同配置的完整检查。只列出工具,未调用工具。此处不测量输入结构的变化。

尚无完整工具检查。