{"repo":"mudassar531/hearsay","free":true,"listed":false,"github":"https://github.com/mudassar531/hearsay","clone":"git clone https://github.com/mudassar531/hearsay.git","description":"crawl4ai for video & audio — one command turns any YouTube video, podcast, or local recording into clean, timestamped, LLM-ready markdown","language":"Python","stars":16,"topics":["cli","faster-whisper","llm","markdown","mcp","podcast","python","rag","speech-to-text","transcription"],"license":"MIT","category":"ai-agents","readme_excerpt":"hearsay crawl4ai for video & audio. One command turns any YouTube video, podcast episode, or local recording into clean, timestamped, LLM-ready markdown — or a TTS/STT training dataset of sliced audio clips paired with verbatim transcripts. Captions-first, runs locally, no plumbing. One input — a link or a file — and two kinds of output. Read it (RAG, notes, agents) or train on it (text-to-speech, speech recognition): 📄 Clean markdown — for RAG, notes &amp; agents 🎙️ TTS/STT dataset — for training Captions when they exist (fast, no download); local Whisper or Apple-Silicon Parakeet transcription when they don't. Single videos, whole playlists, and podcast feeds. Nothing leaves your machine. - How it works · Install - 🎙️ Build TTS/STT datasets · 📄 Clean markdown - Web UI · Transcription engines · MCP server - How it compares · CLI reference · Requirements How it works One pipeline, two outputs. hearsay gets a word-timestamped transcript the cheapest way it can — existing captions if the source has them, otherwise local transcription — then either reflows it into readable markdown or slices the audio into training clips: - Captions-first. Uses the source's captions when available — fast, no media download. - Falls back to transcription automatically (CPU Whisper, or Parakeet on Apple Silicon). - Local & private. Everything runs on your machine; hearsay hosts nothing and ships no data. - Scales. One video, a whole YouTube playlist, or a podcast RSS feed — batched into one ou","default_branch":null,"files":null,"tree":[],"storefront":"/r/mudassar531","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/mudassar531/hearsay/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}