{"repo":"lizongying/go-crawler","free":true,"listed":false,"github":"https://github.com/lizongying/go-crawler","clone":"git clone https://github.com/lizongying/go-crawler.git","description":"A web crawling framework implemented in Golang, it is simple to write and delivers powerful performance. It comes with a wide range of practical middleware and supports various parsing and storage methods. Additionally, it supports distributed deployment. 基于golang实现的爬虫框架，编写简单，性能强劲。内置了丰富的实用中间件，支持多种解析、保存方式，支持分布式部署。","language":"Go","stars":149,"topics":["crawler","golang","scrapy","spider","scraping"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"go-crawler A web crawling framework implemented in Golang, it is simple to write and delivers powerful performance. It comes with a wide range of practical middleware and supports various parsing and storage methods. Additionally, it supports distributed deployment. go-crawler document 中文 Contents 1. Feature 2. Install 3. Usage 1. Basic Architecture 2. Options 3. Item 4. Middleware 5. Pipeline 6. Request 7. Response 8. Signals 9. Proxy 10. Media Downloads 11. Mock Server 12. Configuration 13. Startup 14. Web Page Parsing Based on Field Tags 4. api 5. Q&A 6. Example 7. Tools 1. Certificate 2. MITM 8. TODO Feature Simple to write, yet powerful in performance. Built-in various practical middleware for easier development. Supports multiple parsing methods for simpler page parsing. Supports multiple storage methods for more flexible data storage. Provides numerous configuration options for richer customization. Allows customizations for components, providing more freedom for feature extensions. Includes a built-in mock Server for convenient debugging and development. It supports distributed deployment. Support Summary Parsing supports CSS, XPath, Regex, and JSON. Output supports JSON, CSV, MongoDB, MySQL, Sqlite, and Kafka. Supports Chinese decoding for gb2312, gb18030, gbk, big5 character encodings. Supports gzip, deflate, and brotli decompression. Supports distributed processing. Supports Redis and Kafka as message queues. Supports automatic handling of cookies and redirects. Su","default_branch":null,"files":null,"tree":[],"storefront":"/r/lizongying","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/lizongying/go-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}