{"repo":"postmodern/spidr","free":true,"listed":false,"github":"https://github.com/postmodern/spidr","clone":"git clone https://github.com/postmodern/spidr.git","description":"A versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use.","language":"Ruby","stars":835,"topics":["spider","ruby","spider-links","crawler","web","scraper","web-scraping","web-spider","web-crawler","web-scraper"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"Spidr Homepage Source Issues Mailing List Description Spidr is a versatile Ruby web spidering library that can spider a site, multiple domains, certain links or infinitely. Spidr is designed to be fast and easy to use. Features Follows: a tags. iframe tags. frame tags. Cookie protected links. HTTP 300, 301, 302, 303 and 307 Redirects. Meta-Refresh Redirects. HTTP Basic Auth protected links. Black-list or white-list URLs based upon: URL scheme. Host name Port number Full link URL extension Optional /robots.txt support. Provides callbacks for: Every visited Page. Every visited URL. Every visited URL that matches a specified pattern. Every origin and destination URI of a link. Every URL that failed to be visited. Provides action methods to: Pause spidering. Skip processing of pages. Skip processing of links. Restore the spidering queue and history from a previous session. Custom User-Agent strings. Custom proxy settings. HTTPS support. Examples Start spidering from a URL: Spider a host: Spider a domain (and any sub-domains): Spider a site: Spider multiple hosts: Do not spider certain links: Do not spider links on certain ports: Do not spider links blacklisted in robots.txt: Print out visited URLs: Build a URL map of a site: Print out the URLs that could not be requested: Finds all pages which have broken links: Search HTML and XML pages: Print out the titles from every page: Print out every HTTP redirect: Find what kinds of web servers a host is using, by accessing the headers: ","default_branch":null,"files":null,"tree":[],"storefront":"/r/postmodern","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/postmodern/spidr/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}