{"repo":"FluidInference/FluidAudio","free":true,"listed":false,"github":"https://github.com/FluidInference/FluidAudio","clone":"git clone https://github.com/FluidInference/FluidAudio.git","description":"Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.","language":"Swift","stars":2657,"topics":["coreml","ios","macos","speaker-diarization","speaker-embedding","speaker-identification","speaker-recognition","swift","audio","avfoundation"],"license":"Apache-2.0","category":"mobile-apps","readme_excerpt":"FluidAudio - Transcription, Text-to-speech, VAD, Speaker diarization with CoreML Models FluidAudio is a Swift SDK for fully local, low-latency audio AI on Apple devices, with inference offloaded to the Apple Neural Engine (ANE), resulting in less memory and generally faster inference. The SDK includes state-of-the-art speaker diarization, transcription, and voice activity detection via open-source models (MIT/Apache 2.0) that can be integrated with just a few lines of code. Models are optimized for background processing, ambient computing and always on workloads by running inference on the ANE, minimizing CPU usage and avoiding GPU/MPS entirely. For custom use cases, feedback, additional model support, or platform requests, join our Discord. We're also bringing visual, language, and TTS models to device and will share updates there. Below are some featured local AI apps using Fluid Audio models on macOS and iOS: Want to convert your own model? Check möbius Highlights - Automatic Speech Recognition (ASR) : Parakeet TDT v3 (0.6b) and other TDT/CTC models for batch transcription supporting 25 European languages and Japanese, plus SenseVoice and Paraformer for Mandarin Chinese; Parakeet EOU (120m) for streaming ASR with end-of-utterance detection (English only). See all ASR models. - Inverse Text Normalization (ITN) : Post-process ASR output to convert spoken-form to written-form (\"two hundred\" → \"200\"). See text-processing-rs - Text-to-Speech (TTS) : Kokoro (82m) for parallel sy","default_branch":null,"files":null,"tree":[],"storefront":"/r/FluidInference","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/FluidInference/FluidAudio/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}