{"repo":"alesaccoia/VoiceStreamAI","free":true,"listed":false,"github":"https://github.com/alesaccoia/VoiceStreamAI","clone":"git clone https://github.com/alesaccoia/VoiceStreamAI.git","description":"Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS","language":"Python","stars":960,"topics":["ai","speech-recognition","speech-to-text","websocket"],"license":"MIT","category":"networking-infra","readme_excerpt":"VoiceStreamAI VoiceStreamAI is a Python 3 -based server and JavaScript client solution that enables near-realtime audio streaming and transcription using WebSocket. The system employs Huggingface's Voice Activity Detection (VAD) and OpenAI's Whisper model (faster-whisper being the default) for accurate speech recognition and processing. Features - Real-time audio streaming through WebSocket. - Modular design for easy integration of different VAD and ASR technologies. - Factory and strategy pattern implementation for flexible component management. - Unit testing framework for robust development. - Customizable audio chunk processing strategies. - Support for multilingual transcription. - Supports Secure Sockets with optional cert and key file arguments Demo Video https://github.com/alesaccoia/VoiceStreamAI/assets/1385023/9b5f2602-fe0b-4c9d-af9e-4662e42e23df Demo Client Running with Docker This will not guide you in detail on how to use CUDA in docker, see for example here. Still, these are the commands for Linux: You can build the container image with: After getting your VAD token (see next sections) run: The \"volume\" stuff will allow you not to re-download the huggingface models each time you re-run the container. If you don't need this, just use: Normal, Manual Installation To set up the VoiceStreamAI server, you need Python 3.8 or later and the following packages: 1. transformers 2. pyannote.core 3. pyannote.audio 4. websockets 5. asyncio 6. sentence-transformers 7. faster-","default_branch":null,"files":null,"tree":[],"storefront":"/r/alesaccoia","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/alesaccoia/VoiceStreamAI/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}