{"repo":"crwlrsoft/crawler","free":true,"listed":false,"github":"https://github.com/crwlrsoft/crawler","clone":"git clone https://github.com/crwlrsoft/crawler.git","description":"Library for Rapid (Web) Crawler and Scraper Development","language":"PHP","stars":369,"topics":["crawling","php","scraper","scraping","scraping-websites","web-crawler","web-crawling","web-scraping","hacktoberfest","crawler"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Library for Rapid (Web) Crawler and Scraper Development This library provides kind of a framework and a lot of ready to use, so-called steps , that you can use as building blocks, to build your own crawlers and scrapers with. To give you an overview, here's a list of things that it helps you with: Crawler Politeness &#128519; (respecting robots.txt, throttling,...) Load URLs using a (PSR-18) HTTP client (default is of course Guzzle) or a headless browser (chrome) to get source after Javascript execution Get absolute links from HTML documents &#x1F517; Get sitemaps from robots.txt and get all URLs from those sitemaps Crawl (load) all pages of a website &#x1F577; Use cookies (or don't) &#x1F36A; Use any HTTP methods (GET, POST,...) and send any headers or body Easily iterate over paginated list pages &#x1F501; Extract data from: HTML and also XML (using CSS selectors or XPath queries) JSON (using dot notation) CSV (map columns) Extract schema.org structured data in JSON-LD format from HTML documents Keep memory usage low by using PHP Generators &#x1F4AA; Cache HTTP responses during development, so you don't have to load pages again and again after every code change Get logs about what your crawler is doing (accepts any PSR-3 LoggerInterface) And a lot more... Documentation You can find the documentation at crwlr.software. Contributing If you consider contributing something to this package, read the contribution guide (CONTRIBUTING.md).","default_branch":null,"files":null,"tree":[],"storefront":"/r/crwlrsoft","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/crwlrsoft/crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}