{"repo":"obeone/crawler-to-md","free":true,"listed":false,"github":"https://github.com/obeone/crawler-to-md","clone":"git clone https://github.com/obeone/crawler-to-md.git","description":"Convert web content to Markdown & JSON files to fuel your GPTs !","language":"Python","stars":40,"topics":["chatgpt","markdown","scaper"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"Web Scraper to Markdown 🌐✍️ This Python-based web scraper fetches content from URLs and exports it into Markdown and JSON formats, specifically designed for simplicity, extensibility, and for uploading JSON files to GPT models. It is ideal for those looking to leverage web content for AI training or analysis. 🤖💡 🚀 Quick Start (Or even better, use Docker! 🐳 ) Recommended installation using pipx (isolated environment) Alternatively, install with pip Then run the scraper: 🌟 Features - Scrapes web pages for content and metadata. 📄 - Filters links by base URL. 🔍 - Excludes URLs containing certain strings. ❌ - Automatically finds links or can use a file of URLs to scrape. 🔗 - Rate limiting and delay support. 🕘 - Exports data to Markdown and JSON, ready for GPT uploads. 📤 - Exports each page as an individual Markdown file if --export-individual is used. 📝 - Uses SQLite for efficient data management. 📊 - Configurable via command-line arguments. ⚙️ - Include or exclude specific HTML elements using CSS-like selectors (#id, .class, tag) during Markdown conversion. 🧩 - Docker support. 🐳 📋 Requirements Python 3.10 or higher is required. Project dependencies are managed with pyproject.toml . Install them with: 🛠 Usage Start scraping with the following command: Options: - --url , -u : The starting URL. 🌍 - --urls-file : Path to a file containing URLs to scrape, one URL per line. If '-', read from stdin. 📁 - --output-folder , -o : Where to save Markdown files (default: ./o","default_branch":null,"files":null,"tree":[],"storefront":"/r/obeone","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/obeone/crawler-to-md/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}