{"repo":"spider-rs/spider-nodejs","free":true,"listed":false,"github":"https://github.com/spider-rs/spider-nodejs","clone":"git clone https://github.com/spider-rs/spider-nodejs.git","description":"Spider ported to Node.js","language":"Rust","stars":59,"topics":["crawler","indexer","spider","headless-chrome","nodejs","typescript","distributed-systems","scraper"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"spider-rs The spider project ported to Node.js Getting Started 1. npm i @spider-rs/spider-rs --save Collect the resources for a website. Run the crawls in the background on another thread. Use headless Chrome rendering for crawls. Cron jobs can be done with the following. Use the crawl shortcut to get the page content and url. Benchmarks View the benchmarks to see a breakdown between libs and platforms. Test url: https://espn.com libraries pages speed :--------------------------- :-------- :------ spider(rust): crawl 150,387 1m spider(nodejs): crawl 150,387 153s spider(python): crawl 150,387 186s scrapy(python): crawl 49,598 1h crawlee(nodejs): crawl 18,779 30m The benches above were ran on a mac m1, spider on linux arm machines performs about 2-10x faster. Development Install the napi cli npm i @napi-rs/cli --global . 1. yarn build:test","default_branch":null,"files":null,"tree":[],"storefront":"/r/spider-rs","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/spider-rs/spider-nodejs/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}