{"repo":"clemlesne/scrape-it-now","free":true,"listed":false,"github":"https://github.com/clemlesne/scrape-it-now","clone":"git clone https://github.com/clemlesne/scrape-it-now.git","description":"Web scraper made for AI and simplicity in mind. It runs as a CLI that can be parallelized and outputs high-quality markdown content.","language":"Python","stars":543,"topics":["ai","azure","cli","markdown","scraper"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"🛰️ Scrape It Now! Web scraper made for AI and simplicity in mind. It runs as a CLI that can be parallelized and outputs high-quality markdown content. Features Shared: - 🏗️ Decoupled architecture with Azure Queue Storage or local sqlite - ⚙️ Idempotent operations that can be run in parallel - 💾 Scraped content is stored in Azure Blob Storage or local disk Scraper: - 🛑 Avoid re-scraping a page if it hasn't changed - 🚫 Block ads to lower network costs with The Block List Project - 🔗 Explore pages in depth by detecting links and de-duplicating them - ✍️ Extract markdown content from a page with Pandoc - 🏷️ Extract metadata elements from the page - 🖥️ Load dynamic JavaScript content with Playwright and Chromium - 🕵️‍♂️ Preserve anonymity with a random user agent, random viewport size, and no client hints headers - 📊 Show progress with a status command - 🖼️ Store images collected on the page - 📸 Store screenshot of the page - 📡 Track progress of total network usage Indexer: - 🧠 AI Search index is created automatically - ✂️ Chunk markdown while keeping the content coherent - 📈 Embed chunks with OpenAI embeddings - 🔍 Indexed content is semantically searchable with Azure AI Search Installation From PyPI To configure the CLI (including authentication to the backend services), use environment variables, a .env file or command line options. From sources Application must be run with Python 3.13 or later. If this version is not installed, an easy way to install it is pyenv","default_branch":null,"files":null,"tree":[],"storefront":"/r/clemlesne","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/clemlesne/scrape-it-now/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}