{"repo":"gxcsoccer/wechat-article-crawler","free":true,"listed":false,"github":"https://github.com/gxcsoccer/wechat-article-crawler","clone":"git clone https://github.com/gxcsoccer/wechat-article-crawler.git","description":"抓取微信公众号文章，导出结构化数据和 Markdown，自动下载图片绕过防盗链 | Claude Code skill for crawling WeChat articles","language":"Python","stars":12,"topics":["claude-code","claude-code-skill","crawler","markdown","python","web-scraping","wechat","wechat-article","wechat-crawler","weixin"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"wechat-article-crawler A Claude Code skill for crawling WeChat public account (微信公众号) articles into structured data and clean markdown — with images downloaded locally to bypass hotlink protection. Features - Spoofs WeChat in-app browser User-Agent to get full article content - Fixes lazy-loaded images ( data-src → src ) - Extracts structured metadata: title, author, publish time, account description - Generates clean markdown with inline images (SVG placeholders removed) - Downloads images locally with correct Referer header to bypass CDN hotlink protection - Concurrent downloads with rate limiting (semaphore) Quick Start As a Claude Code Skill Copy this directory to .claude/skills/crawl-wechat in your project: Then in Claude Code, just paste a WeChat article URL or say \"抓取这篇微信文章\". As a Standalone Script As a Python Library Output Field Description ---------------- ------------------------------- title Article title author Public account name publish time Publication timestamp account desc Account description/bio markdown Clean markdown with images html Raw HTML of article body url Final URL after redirects CLI Options How It Works 1. User-Agent spoofing — Sets MicroMessenger/8.0.43 so WeChat serves the full article 2. Dynamic wait — wait for=\"css:#js content\" ensures the body is rendered 3. Lazy-image JS injection — Copies data-src → src on all tags before scraping 4. Structured extraction — JsonCssExtractionStrategy targets WeChat's DOM ( #activity-name , #js name , #publi","default_branch":null,"files":null,"tree":[],"storefront":"/r/gxcsoccer","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/gxcsoccer/wechat-article-crawler/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}