{"repo":"paulpierre/markdown-crawler","free":true,"listed":false,"github":"https://github.com/paulpierre/markdown-crawler","clone":"git clone https://github.com/paulpierre/markdown-crawler.git","description":"A multithreaded 🕸️ web crawler that recursively crawls a website and creates a 🔽 markdown file for each page, designed for LLM RAG","language":"Python","stars":467,"topics":["html-to-markdown","html-to-markdown-converter","html2md","llm","llmops","markdown","markdown-parser","rag","web-scraper","markdown-crawler"],"license":"MIT","category":"ai-agents","readme_excerpt":"📝 Overview This is a multithreaded web crawler that crawls a website and creates markdown files for each page. It was primarily created for large language model document parsing to simplify chunking and processing of large documents for RAG use cases. Markdown by nature is human readable and maintains document structure while keeping a small footprint. ✨ Features include - 🧵 Threading support for faster crawling - ⏯️ Continue scraping where you left off - ⏬ Set the max depth of children you wish to crawl - 📄 Support for tables, images, etc. - ✅ Validates URLs, HTML, filepaths - ⚙️ Configure list of valid base paths or base domains - 🚫 Exclude specific paths from crawling ( --exclude-paths ) - 🎨 Configurable markdown heading style ( --heading-style ) - 🌐 Browser-like User-Agent to bypass JS checks - 🍲 Uses BeautifulSoup to parse HTML - 🪵 Verbose logging option - 👩‍💻 Ready-to-go CLI interface - 🧪 Comprehensive test suite (95% coverage) 📦 What's New in v0.0.9 This release resolves 6 open issues and introduces several community-requested features: Issue Fix ------- ----- #7 UnicodeEncodeError on non-ASCII content — UTF-8 encoding on all file writes #8 Bypass JavaScript checks — browser-like User-Agent header on all requests #9 Exclude paths from crawling — new --exclude-paths / -x CLI flag #10 target content CSS selector handling — null href links now skipped #11 UnboundLocalError fix in get target content #17 Dependencies added to pyproject.toml for uvx support #20 C","default_branch":null,"files":null,"tree":[],"storefront":"/r/paulpierre","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/paulpierre/markdown-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}