{"repo":"discourselab/scrapai-cli","free":true,"listed":false,"github":"https://github.com/discourselab/scrapai-cli","clone":"git clone https://github.com/discourselab/scrapai-cli.git","description":"AI-powered web scraping CLI. Describe what you want, get a production-ready Scrapy spider. Write once, reuse forever.","language":"Python","stars":114,"topics":["ai","claude-code","cli","cloudflare","cloudflare-bypass","cloudflare-scraper","crawler","python","scraper","scrapy"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"scrapai A CLI where you describe what you want to scrape in plain English, an AI agent builds the scraper, and Scrapy runs it. Minutes later you have a tested, production-ready scraper stored in a database. No Python, no CSS selectors, no Scrapy knowledge. The AI agent analyzes the site, writes extraction rules, verifies quality, and saves a reusable config. Run it tomorrow or next year. Same command, no AI costs. Built by DiscourseLab. Used in production across 500+ websites. Table of Contents - Who This Is For - Why scrapai? - How It Works - Features - Quick Start - Using with AI Agents - Migrating Existing Scrapers - For Developers - Architecture - Security - CLI Reference - Configuration - Limitations - Documentation - Contributing - Responsible Use - License Who This Is For Good fit: - Teams that need to scrape many websites and don't want to write individual scrapers - Non-technical users who can describe what they want in plain English - Organizations where scraping is a means to an end, not the core competency - Anyone building datasets from public web content (news, research, documentation) Not a good fit: - Single-site scraping where you want fine-grained control (use Scrapling or crawl4ai) - Sites with hard CAPTCHAs (we handle Cloudflare challenges, not Capsolver-level CAPTCHAs) - Login-required or paywall content (not supported yet) See COMPARISON.md for a detailed comparison with Scrapling and crawl4ai. Why scrapai? We needed data for our work. Hundreds of websit","default_branch":null,"files":null,"tree":[],"storefront":"/r/discourselab","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/discourselab/scrapai-cli/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}