{"repo":"xuxueli/xxl-crawler","free":true,"listed":false,"github":"https://github.com/xuxueli/xxl-crawler","clone":"git clone https://github.com/xuxueli/xxl-crawler.git","description":"A lightweight web crawler framework.（Java爬虫框架）","language":"Java","stars":761,"topics":["crawler","web","spider","object-oriented","flexible","xxl-crawler","java","distributed"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"XXL-CRAWLER XXL-CRAWLER, a lightweight web crawler framework. -- Home Page -- Introduction XXL-CRAWLER is a lightweight web crawler framework. A line of code to develop a multi-threaded crawler, fully annotated way to collect page data to Java objects, with \"multi-threaded, fully annotated, JS rendering, proxy, distributed extension\" and other features; XXL-CRAWLER 是一个轻量级Java爬虫框架。一行代码开发一个多线程爬虫，全注解方式采集页面数据至Java对象，拥有\"多线程、全注解、JS渲染、代理、分布式扩展\"等特性； Documentation - 中文文档 Features - 1、简洁：API直观简洁，可快速上手； - 2、轻量级：底层实现仅强依赖jsoup，简洁高效； - 3、模块化：模块化的结构设计，可轻松扩展； - 4、全注解：支持通过注解提取页面数据，高效映射页面数据到PageVO对象，底层自动完成PageVO对象的数据抽取和封装返回；单个页面支持抽取一个或多个PageVO； - 5、多线程：线程池方式运行，提高采集效率； - 6、扩散全站：支持以现有URL为起点扩散爬取整站； - 7、JS渲染：通过扩展 \"PageLoader\" 模块，支持采集JS动态渲染数据。原生提供 Jsoup(非JS渲染，速度更快)、Selenium+ChromeDriver(JS渲染，兼容性高) 等多种实现，支持自由扩展其他实现； - 8、代理IP：对抗反采集策略规则WAF； - 9、动态代理：支持运行时动态调整代理池，以及自定义代理池路由策略； - 10、失败重试：请求失败后重试，并支持设置重试次数； - 11、异步：支持同步、异步两种方式运行； - 12、幂等去重：防止重复爬取； - 13、URL扩散过滤：支持设置页面白名单正则，过滤URL； - 14、分布式支持：通过扩展 \"RunUrlPool\" 模块，并结合Redis或DB共享运行数据可实现分布式。默认提供LocalRunUrlPool单机版爬虫； - 15、自定义请求信息，如：请求参数、Cookie、Header、UserAgent轮询、Referrer等； - 16、动态参数：支持运行时动态调整请求参数； - 17、超时控制：支持设置爬虫请求的超时时间； - 18、主动停顿：爬虫线程处理完页面之后进行主动停顿，避免过于频繁被拦截； Example code 注意：仅供学习测试使用，如有侵犯请联系删除 如下测试代码可以前往仓库查看：测试代码目录 序号 爬虫名称 功能描述 测试用例代码文件 ---- ----------------------------- ------------------------------------------------------------------------------------- ------------------ 1 CNBlog精华文章数据爬虫【页面提取数据】 一行代码启动多线程爬虫，分页方式扩散爬取“CNBlog精华文章”，通过“注解式”自动提取页面数据，封装成PageVo输出； X","default_branch":null,"files":null,"tree":[],"storefront":"/r/xuxueli","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/xuxueli/xxl-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}