{"repo":"thushan/olla","free":true,"listed":false,"github":"https://github.com/thushan/olla","clone":"git clone https://github.com/thushan/olla.git","description":"High-performance lightweight proxy and load balancer for LLM infrastructure. Intelligent routing, automatic failover and unified model discovery across local and remote inference backends.","language":"Go","stars":281,"topics":["ai","llm-inference","lmstudio","ollama","proxy","vllm","golang","llamacpp","llm-proxy","llm-router"],"license":"Apache-2.0","category":"networking-infra","readme_excerpt":"Recorded with VHS - see demo tape &nbsp; &nbsp; Olla is a high-performance, low-overhead, low-latency proxy and load balancer for managing LLM infrastructure. It intelligently routes LLM requests across self-hosted inference nodes with a wide variety of natively supported endpoints and extensible enough to support others. Olla provides model discovery and unified model catalogues within each provider, enabling seamless routing to available models on compatible endpoints. Olla works alongside API gateways like LiteLLM or orchestration platforms like GPUStack, focusing on making your existing LLM infrastructure reliable through intelligent routing and failover. You can choose between two proxy engines: Sherpa for simplicity and maintainability or Olla for maximum performance with advanced features like circuit breakers and connection pooling. Single CLI application and config file is all you need to go Olla! For large GPU deployments, enterprise and data centre use, see TensorFoundry FoundryOS. For an inference control plane, consider Alloy. Key Features - 🔄 Smart Load Balancing : Priority-based routing with automatic failover and connection retry - 📌 Sticky Sessions : KV-cache-aware affinity routing that pins multi-turn conversations to the same backend - 🔍 Smart Model Unification : Per-provider unification + OpenAI-compatible cross-provider routing - ⚡ Dual Proxy Engines : Sherpa (simple) and Olla (high-performance) - 🎯 Advanced Filtering : Profile and model filtering wit","default_branch":null,"files":null,"tree":[],"storefront":"/r/thushan","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/thushan/olla/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}