{"repo":"blogdaren/PHPCreeper","free":true,"listed":false,"github":"https://github.com/blogdaren/PHPCreeper","clone":"git clone https://github.com/blogdaren/PHPCreeper.git","description":"A new generation of multi-process async event-driven spider engine based on workerman. Support headless browser. 🌿基于workerman实现的多进程异步事件驱动型PHP爬虫引擎，支持无头浏览器🌿","language":"PHP","stars":140,"topics":["multi-process","asynchronous","event-driven","high-performance","proxy-pool","spider","crawler","socket","workerman"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"PHPCreeper What is it PHPCreeper is a new generation of multi-process asynchronous event-driven spider engine based on workerman. Focus on efficient agile development, and make the crawler job become more easy Solve the performance and scalability bottlenecks of traditional crawler frameworks Take advantage of crawling fully in Multi-Process + Distributed + Separated environment Support headless browser which can execute JavaScript codes for crawling dynamic pages 爬山虎是基于workerman开发的全新一代多进程异步事件驱动型PHP爬虫引擎, 它有助于： 专注于高效敏捷开发，让爬取工作变得更加简单。 解决传统型PHP爬虫框架的性能和扩展瓶颈问题。 充分发挥多进程+分布式+分离式部署环境下的爬取优势。 支持无头浏览器即支持运行JavaScript代码及其渲染页。 Documentation The chinese document is relatively complete, and the english document will be kept up-to-date constantly here. 注意： 爬山虎中文开发文档相对比较完善，各位小伙伴直接点击下方链接阅读即可. 爬山虎中文官方网站：http://www.phpcreeper.com 中文开发文档主节点：http://www.phpcreeper.com/docs/ 爬山虎是一个免费开源的佛系爬虫项目，欢迎小星星Star支持，让更多的人发现、使用并受益。 爬山虎源码根目录下有一个 Examples/start.php 样例脚本，开发之前建议先阅读它而后运行它。 爬山虎提供的例子如果未能按照预期工作，请检查修改爬取规则，因为源站DOM极可能更新了。 技术交流 下方绿色二维码为微信技术交流群：phpcreeper【进群之前需先加此专属微信并备明来意或附上备注：爬山虎】 下方橙色二维码为作者的淘宝店铺，有需要的小伙伴可以购买支持作者的全套原创视频《深入PHP内核源码》 《深入PHP内核源码》原创视频的配套文档是作者一个字一个字随堂认真敲写而来，文字总数高达有近30000字，并且附有大量自绘原创插图，所以如果你是通过workerman社区联系到作者本人，计划在视频录制结束并二次完善文档之后的合适时间 免费赠予有缘小伙伴 ，点此观看视频配套文档和完整目录章节视频 。 微信群主要围绕 爬山虎 和 workerman 和 深入PHP内核源码 开展技术交流，观看PHP内核源码视频请移步至B站。 &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Screenshot Features Inherit almost all features from workerman Support headless browser for crawling dynamic pages Support ","default_branch":null,"files":null,"tree":[],"storefront":"/r/blogdaren","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/blogdaren/PHPCreeper/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}