{"repo":"paulrobello/par_scrape","free":true,"listed":false,"github":"https://github.com/paulrobello/par_scrape","clone":"git clone https://github.com/paulrobello/par_scrape.git","description":"AI assisted web scraping and data extraction","language":"Python","stars":218,"topics":["ai","markdown","webscraping"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"PAR Scrape PAR Scrape is a versatile web scraping tool with options for Selenium or Playwright, featuring AI-powered data extraction and formatting. Table of Contents - Features - Known Issues - Prompt Cache - How it works - Site Crawling - Crawl state - Incremental crawls - Prerequisites - Installation - Usage - Custom extraction prompts - Roadmap - What's New - Contributing - License Screenshots Features - Web scraping using Playwright or Selenium - AI-powered data extraction and formatting - Can be used to crawl and extract clean markdown without AI - Supports multiple output formats (JSON, Excel, CSV, Markdown) - Customizable field extraction - Token usage and cost estimation - Prompt cache for Anthropic provider - Uses my PAR AI Core Known Issues - Selenium silent mode on windows still shows message about websocket. There is no simple way to get rid of this. - Providers other than OpenAI are hit-and-miss depending on provider / model / data being extracted. Prompt Cache - OpenAI will auto cache prompts that are over 1024 tokens. - Anthropic will only cache prompts if you specify the --prompt-cache flag. Due to cache writes costing more only enable this if you intend to run multiple scrape jobs against the same url, also the cache will go stale within a couple of minutes so to reduce cost run your jobs as close together as possible. How it works - Data is fetched from the site using either Selenium or Playwright - HTML is converted to clean markdown - If you specify an ou","default_branch":null,"files":null,"tree":[],"storefront":"/r/paulrobello","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/paulrobello/par_scrape/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}