{"repo":"s0rg/crawley","free":true,"listed":false,"github":"https://github.com/s0rg/crawley","clone":"git clone https://github.com/s0rg/crawley.git","description":"The unix-way web crawler","language":"Go","stars":341,"topics":["cli","crawler","golang-application","unix-way","web-scraping","web-spider","golang","web-crawler","pentest-tool","go"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"crawley Crawls web pages and prints any link it can find. features - fast html SAX-parser (powered by x/net/html) - js/css lexical parsers (powered by tdewolff/parse) - extract api endpoints from js code and url() properties - small (below 1500 SLOC), idiomatic, more than 80% test covered codebase - grabs most of useful resources urls (pics, videos, audios, forms, etc...) - found urls are streamed to stdout and guranteed to be unique (with fragments omitted) - scan depth (limited by starting host and path, by default - 0) can be configured - can be polite - crawl rules and sitemaps from robots.txt - brute mode - scan html comments for urls (this can lead to bogus results) - make use of HTTP PROXY / HTTPS PROXY environment values + handles proxy auth (use HTTP PROXY=\"socks5://127.0.0.1:1080/\" crawley for socks5) - directory-only scan mode (aka fast-scan ) - user-defined cookies, in curl-compatible format (i.e. -cookie \"ONE=1; TWO=2\" -cookie \"ITS=ME\" -cookie @cookie-file ) - user-defined headers, same as curl: -header \"ONE: 1\" -header \"TWO: 2\" -header @headers-file - tag filter - allow to specify tags to crawl for (single: -tag a -tag form , multiple: -tag a,form , or mixed) - url ignore - allow to ignore urls with matched substrings from crawling (i.e.: -ignore logout ) - subdomains support - allow depth crawling for subdomains as well (e.g. crawley http://some-test.site will be able to crawl http://www.some-test.site ) examples installation - binaries / deb / rpm for Linux, F","default_branch":null,"files":null,"tree":[],"storefront":"/r/s0rg","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/s0rg/crawley/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}