{"repo":"sigoden/rag-crawler","free":true,"listed":false,"github":"https://github.com/sigoden/rag-crawler","clone":"git clone https://github.com/sigoden/rag-crawler.git","description":"Crawl a website to generate knowledge file for RAG","language":"TypeScript","stars":55,"topics":["crawler","knowledge","llm","rag"],"license":"MIT","category":"ai-agents","readme_excerpt":"rag-crawler Crawl a website to generate knowledge file for RAG. Installation Usage Output to stdout Output to JSON file Output to separates files Crawl Markdown files in GitHub Tree Many documentation sites host their source Markdown files on GitHub. The crawler has been optimized to crawl these files directly from GitHub. Preset A preset consists of predefined crawl options. You can review the predefined presets at ./src/preset.ts. Why Use Preset? Let's use GitHub Wiki as an example. To enhance scraping quality, we need to configure both --exclude and --extract . Since all GitHub Wiki websites share these crawl options, we can define a preset for reusability. This allows for a simplified command: When the preset is set to auto , rag-crawler will automatically determine the appropriate preset. It does this by checking if the startUrl matches the test regex. Custom Presets You can add custom presets by editing the /.rag-crawler.json file: License The project is under the MIT License, Refer to the LICENSE file for detailed information.","default_branch":null,"files":null,"tree":[],"storefront":"/r/sigoden","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/sigoden/rag-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}