{"repo":"yazon/flexllama","free":true,"listed":false,"github":"https://github.com/yazon/flexllama","clone":"git clone https://github.com/yazon/flexllama.git","description":"🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GPU support","language":"Python","stars":59,"topics":[],"license":"BSD-3-Clause","category":"self-hosted-apps","readme_excerpt":"FlexLLama - \"One to rule them all\" FlexLLama is a lightweight, extensible, and user-friendly self-hosted tool that easily runs multiple llama.cpp server instances with OpenAI v1 API compatibility . It's designed to manage multiple models across different GPUs, making it a powerful solution for local AI development and deployment. Key Features of FlexLLama - 🚀 Multiple llama.cpp instances - Run different models simultaneously - 🎯 Multi-GPU support - Distribute models across different GPUs - 🔌 OpenAI v1 API compatible - Drop-in replacement for OpenAI endpoints - 📊 Real-time dashboard - Monitor model status, live GPU telemetry, and per-model token throughput in a web interface - 🤖 Chat & Completions - Full chat and text completion support - 🔍 Embeddings & Reranking - Supports models for embeddings and reranking - 🎙️ Audio endpoints - Speech-to-text and text-to-speech proxied to audio models (Voxtral, Qwen3-Omni, and more) - 🛠️ MCP proxy - Optional unified Model Context Protocol endpoint that routes requests to mcp -tagged models - ⚡ Auto-start - Automatically start default runners on launch - 🔄 Model switching - Dynamically load/unload models as needed - ⏱️ Auto model unload - Automatically unload models after a configurable idle timeout Quickstart 🚀 Want to get started in 5 minutes? Check out our QUICKSTART.md for a simple Docker setup with the Qwen3-4B model! 📦 Local Installation 1. Install FlexLLama: From GitHub: From local source (after cloning): 1. Create your co","default_branch":null,"files":null,"tree":[],"storefront":"/r/yazon","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/yazon/flexllama/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}