{"repo":"scrapy/scrapy-bench","free":true,"listed":false,"github":"https://github.com/scrapy/scrapy-bench","clone":"git clone https://github.com/scrapy/scrapy-bench.git","description":"A CLI for benchmarking Scrapy.","language":"Python","stars":32,"topics":["scrapy","scrapy-bench","benchmark-suite","command-line-tool","web-crawler","python"],"license":"MIT","category":"cli-tools","readme_excerpt":"Benchmarking CLI for Scrapy (The project is still in development.) A command-line interface for benchmarking Scrapy, that reflects real-world usage. Why? Currently, the scrapy bench option present just spawns a spider which aggressively crawls randomly generated links at a high speed. The speed thus obtained, which maybe useful for comparisons, does not actually reflects a real-world scenario. The actual speed varies with the python version and scrapy version. Current Features Spawns a CPU-intensive spider which follows a fixed number of links of a static snapshot of the site Books to Scrape. Follows a real-world scenario where various information of the books is extracted, and stored in a .csv file. A broad crawl benchmark that uses 1000 copies of the site Books to Scrape which are dynamically generated using twisted . The server file is present here. A micro benchmark that tests LinkExtractor() function by extracting links from a collection of html pages. A micro benchmark that tests extraction using css from a collection of html pages. A micro benchmark that tests extraction using xpath from a collection of html pages Profile the benchmarkers with vmprof and upload to their website Options --n-runs option for performing more than one iteration of spider to improve the precision. --only result option for viewing the results only. --upload result option to upload the results to local codespeed for better comparison. Spider settings SCRAPY BENCH RANDOM PAYLOAD SIZE : Adds a r","default_branch":null,"files":null,"tree":[],"storefront":"/r/scrapy","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/scrapy/scrapy-bench/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}