{"repo":"RightNow-AI/local-kimi","free":true,"listed":false,"github":"https://github.com/RightNow-AI/local-kimi","clone":"git clone https://github.com/RightNow-AI/local-kimi.git","description":"Optimized local serving engine for Kimi-Linear-48B: INT4 quantizer, fused decode kernels for a measured 3.18x, and an OpenAI-compatible server. Ships with k3, a bridge that detects the client per request so Claude Code, Codex, Cline, Aider and opencode all work unchanged.","language":"Python","stars":37,"topics":["anthropic-api","claude-code","coding-agents","int4","kimi","llama-cpp","llm-inference","local-llm","openai-api","quantization"],"license":"Apache-2.0","category":"mcp-servers","readme_excerpt":"local-kimi Sketch is illustrative. This repository serves Kimi-Linear-48B, measured on an NVIDIA L40S inside a hard 32 GiB cap. The full Kimi K3 is 2.78T parameters and runs on an 8x B300 node, linked below. TL;DR. Run Kimi on your own consumer GPU, not someone's API. A single 32 GB card (RTX 5090, or any datacenter card) serves Kimi-Linear-48B at 113.83 tok/s , up 3.18x from fused INT4 kernels, and Claude Code, Codex, Cline, Aider and opencode all connect to it unchanged. Need the full Kimi K3 in production? 2.8 trillion parameters will not fit a workstation. If you want it running inside your own network instead of behind a closed-source API, RunInfra ships it as a deployable package with a pinned vLLM build, a benchmark receipt and weight verification: Kimi K3 on 8x B300 → Point any coding agent at a local Kimi model and it works: k3 translates between the agent's protocol and the model's. The problem Local model servers usually speak OpenAI Chat Completions. Claude Code speaks Anthropic Messages. Codex speaks OpenAI Responses. Point the wrong client at the wrong endpoint and the request fails. k3 sits between them, detects the caller on each request, and translates the request and response. Protocol support The protocol landscape has changed. Current releases of vLLM, llama.cpp, and Ollama now document all three protocol families. k3 is different because it is a standalone adapter for an existing OpenAI Chat Completions backend. Capability vLLM llama.cpp Ollama k3 --- ---","default_branch":null,"files":null,"tree":[],"storefront":"/r/RightNow-AI","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/RightNow-AI/local-kimi/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}