{"repo":"chyroc/WechatSogou","free":true,"listed":false,"github":"https://github.com/chyroc/WechatSogou","clone":"git clone https://github.com/chyroc/WechatSogou.git","description":"基于搜狗微信搜索的微信公众号爬虫接口","language":"Python","stars":6378,"topics":["wechat","sogou","python","crawler","pypi","scrapy"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"基于搜狗微信搜索的微信公众号爬虫接口 === 我的另外一个作品: https://github.com/chyroc/lark ，基于代码生成的 Lark/飞书 Go SDK，欢迎 star 。 项目简介 基于搜狗微信搜索的微信公众号爬虫接口，可以扩展成基于搜狗搜索的爬虫 如果有问题，请提issue CHANGELOG 交流分享 - QQ群（只需加一个） - 一群 132955136（已满） - 二群 819084985 - 微信群 赞助作者 甲鱼说，咖啡是灵魂的饮料，买点咖啡 谢谢这些人的☕️ 问题集锦 Q:没有得到原始文章url / 提示链接已经过期？ A:微信屏蔽此接口，请在临时链接有效期内保存文章内容。 Q:获取文章只能10篇？ A:是的，仅显示最近10条群发。 Q:使用的是python 2 还是 3？ A:都支持，若出错，请报BUG。 安装 使用 初始化 API 获取特定公众号信息 - get gzh info - 使用 - 返回数据结构 搜索公众号 - 使用 - 数据结构 list of dict, dict: 搜索微信文章 - 使用 - 数据结构 list of dict, dict: 解析最近文章页 - get gzh article by history - 使用 - 数据结构 解析 首页热门 页 - get gzh article by hot - 使用 - 数据结构 获取关键字联想词 - 使用 - 数据结构 关键词列表 --- TODO - [x] 相似文章的公众号获取 - [ ] 主页热门公众号获取 - [ ] 文章详情页信息 - [x] 所有类型的解析 - [ ] 验证码识别 - [ ] 接入爬虫框架 - [x] 兼容py2 ---","default_branch":null,"files":null,"tree":[],"storefront":"/r/chyroc","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/chyroc/WechatSogou/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}