{"repo":"toverainc/willow-inference-server","free":true,"listed":false,"github":"https://github.com/toverainc/willow-inference-server","clone":"git clone https://github.com/toverainc/willow-inference-server.git","description":"Open source, local, and self-hosted highly optimized language inference server supporting ASR/STT, TTS, and LLM across WebRTC, REST, and WS","language":"Python","stars":512,"topics":["cuda","deep-learning","llama","llm","privacy","speech-recognition","speech-to-text","text-to-speech","vicuna","webrtc"],"license":"Apache-2.0","category":"machine-learning","readme_excerpt":"Willow Inference Server Watch the WIS WebRTC Demo Willow Inference Server (WIS) is a focused and highly optimized language inference server implementation. Our goal is to \"automagically\" enable performant, cost-effective self-hosting of released state of the art/best of breed models to enable speech and language tasks: - Primarily targeting CUDA with support for low-end (cheap) devices such as the Tesla P4, GTX 1060, and up. Don't worry - it screams on an RTX 4090 too! (See benchmarks). Can also run CPU-only. - Memory optimized - all three default Whisper (base, medium, large-v2) models loaded simultaneously with TTS support inside of 6GB VRAM. - ASR. Heavy emphasis - Whisper optimized for very high quality as-close-to-real-time-as-possible speech recognition via a variety of means (Willow, WebRTC, POST a file, integration with devices and client applications, etc). Results in hundreds of milliseconds or less for most intended speech tasks. - TTS. Primarily provided for assistant tasks (like Willow!) and visually impaired users. - Support for a variety of transports. REST, WebRTC, Web Sockets. - Performance and memory optimized. Leverages CTranslate2 for Whisper support. - Willow support. WIS powers the Tovera hosted best-effort example server Willow users enjoy. - Support for WebRTC - stream audio in real-time from browsers or WebRTC applications to optimize quality and response time. Heavily optimized for long-running sessions using WebRTC audio track management. Leave your","default_branch":null,"files":null,"tree":[],"storefront":"/r/toverainc","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/toverainc/willow-inference-server/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}