{"repo":"youngharold/tightwad","free":true,"listed":false,"github":"https://github.com/youngharold/tightwad","clone":"git clone https://github.com/youngharold/tightwad.git","description":"Mixed-vendor GPU inference cluster manager with speculative decoding","language":"Python","stars":31,"topics":["cuda","deep-learning","distributed-inference","gpu-cluster","gpu-inference","gpu-pooling","homelab","llama-cpp","llm-inference","local-llm"],"license":"MIT","category":"machine-learning","readme_excerpt":"Tightwad Mixed-vendor GPU inference cluster manager with speculative decoding proxy. Pools CUDA and ROCm GPUs across machines using llama.cpp RPC, and accelerates inference via application-layer speculative decoding across network-separated servers. AI Agents: See AGENTS.md for a complete setup playbook — hardware assessment, decision tree, step-by-step config, model selection guide, and troubleshooting. Designed for Claude, ChatGPT, Copilot, OpenClaw, and any AI assistant helping users deploy Tightwad. How It Works in 10 Seconds It's not 2 matching GPUs. It's your entire junk drawer of compute unified into one API. That dusty 770 in your closet? Put it to work. The small model is fast but sometimes wrong. The big model is slow but always right. Tightwad uses the small model to do most of the work, and the big model to catch mistakes. Because catching mistakes is cheap — it's one batch operation, not N serial ones. What Does This Look Like as a User? You change nothing about your workflow. Tightwad is invisible. Before After --- --- --- Your chat app Open WebUI, ChatBot UI, etc. Same app, no changes Points at http://192.168.1.10:11434 (Ollama on one machine) http://192.168.1.10:8088 (Tightwad proxy) Model you talk to Qwen3-32B Qwen3-32B (same model, same output) What you see Normal chat responses Normal chat responses, just faster The small model Doesn't exist Hidden — drafting on a different machine entirely Other machines Idle, wasted RTX 2070, old Xeon, laptop — all contri","default_branch":null,"files":null,"tree":[],"storefront":"/r/youngharold","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/youngharold/tightwad/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}