{"repo":"modelscope/FunASR","free":true,"listed":false,"github":"https://github.com/modelscope/FunASR","clone":"git clone https://github.com/modelscope/FunASR.git","description":"Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.","language":"Python","stars":19897,"topics":["pytorch","speech-recognition","paraformer","punctuation","speaker-diarization","voice-activity-detection","asr","multilingual-asr","speech-to-text","transcription"],"license":"MIT","category":"machine-learning","readme_excerpt":"(简体中文 English 日本語 한국어) Industrial speech recognition toolkit for offline, streaming, and edge deployment. ASR · VAD · punctuation · speaker pipelines · emotion and audio-event models · OpenAI-compatible serving Quick Start · Colab · Benchmark · Model selection · Migration guide · Use cases · Community integrations · Deployment matrix · Deployment hub · Troubleshooting · Models · Agent Integration · OpenClaw · Docs · Contribute --- Quick Start No local setup? Open the Colab quickstart to transcribe a public sample or upload your own audio in a browser. For GPU quickstarts, install the PyTorch and torchaudio wheels that match your NVIDIA driver from pytorch.org before installing FunASR. After installation, confirm the GPU is visible: Only use device=\"cuda\" when this prints True ; otherwise use device=\"cpu\" or reinstall PyTorch with the correct CUDA wheel. Flagship model — Fun-ASR-Nano (LLM-ASR for Chinese, English, and Japanese, plus Chinese dialect groups and regional accents; needs a GPU): For the separate 31-language checkpoint, use Fun-ASR-MLT-Nano-2512. Language coverage is checkpoint-specific, so Nano and MLT-Nano should be treated as distinct model choices. On CPU (or for five-language ASR plus emotion and audio-event tags), use SenseVoiceSmall . The pipeline below composes SenseVoiceSmall with FSMN-VAD and CAM++; diarization is provided by the separate CAM++ model, not by the SenseVoiceSmall checkpoint: See the SenseVoice paper, Hugging Face checkpoint, and GGUF ed","default_branch":null,"files":null,"tree":[],"storefront":"/r/modelscope","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/modelscope/FunASR/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}