{"repo":"houtini-ai/houtini-lm","free":true,"listed":false,"github":"https://github.com/houtini-ai/houtini-lm","clone":"git clone https://github.com/houtini-ai/houtini-lm.git","description":"MCP server that saves Claude Code tokens by delegating bounded tasks to local or cloud LLMs. Works with LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras.","language":"JavaScript","stars":108,"topics":["ai-agents","claude-mcp","code-generation","developer-tool","developer-tools","lm-studio","local-llm","mcp","mcp-client","mcp-server"],"license":"Apache-2.0","category":"mcp-servers","readme_excerpt":"@houtini/lm Houtini LM - Save Tokens by Offloading Tasks from Claude Code to Your Local LLM Server (LM Studio / Ollama), Openrouter or a Cloud API Quick Navigation How it works Quick start What gets offloaded Tools Performance tracking Structured JSON output Model routing Self-test (shakedown) Configuration Compatible endpoints Developer guide I built this because I kept leaving Claude Code running overnight on big refactors and the token bill was painful. A huge chunk of that spend goes on bounded tasks any decent model handles fine - generating boilerplate, code review, commit messages, format conversion. Stuff that doesn't need Claude's reasoning or tool access. Houtini LM connects Claude Code to a local LLM on your network - or any OpenAI-compatible API (LM Studio, Ollama, vLLM, DeepSeek, Groq, Cerebras, and OpenRouter's 300+ models through one endpoint). Claude keeps doing the hard work - architecture, planning, multi-file changes - and offloads the grunt work to whatever cheaper model you've got running. No Claude quota burn. No rate limits. Private if local, cheap if cloud. The trade is wall-clock time: local inference is typically 3-30× slower than frontier models, so delegation wins on bounded, self-contained tasks rather than everything. I wrote a full walkthrough of why I built this and how I use it day to day. The manual This README is the overview. The depth lives in focused pages: Page What's in it --- --- Getting started Local models from zero: LM Studio or Doc","default_branch":null,"files":null,"tree":[],"storefront":"/r/houtini-ai","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/houtini-ai/houtini-lm/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}