{"repo":"jxtse/auto-paper-harvester","free":true,"listed":false,"github":"https://github.com/jxtse/auto-paper-harvester","clone":"git clone https://github.com/jxtse/auto-paper-harvester.git","description":"Batch-download paper PDFs by DOI. Routes through publisher TDM APIs (Wiley/Elsevier/Springer) → OA fallbacks (Crossref/OpenAlex/Unpaywall) → optional Playwright browser fallback that reuses your institutional cookies for ACS/RSC/IEEE/AIP/IOP/APS. Ships as both a CLI and an agent skill.","language":"Python","stars":26,"topics":["claude-skill","paper-download","playwright","agent-skill","pdf-download"],"license":null,"category":"scrapers-browser-automation","readme_excerpt":"Auto Paper Harvester Batch-download paper PDFs (and supplementary files) by DOI. Routes each DOI through publisher TDM APIs → open-access aggregators → optional institutional-browser fallback, so you actually get the PDFs your institution is paying for instead of a wall of 403s. Each article ends up in downloads/pdfs/ / .pdf with any supplementary PDFs detected on the landing page saved alongside it. Throughput is throttled to satisfy publisher TDM rate limits (≥ 1 s/file by default). v0.2.0 highlights : 24 DOI-prefix routing table covering 19 publisher families (see docs/SUPPORTED PUBLISHERS.md); new --use-browser-fallback Playwright pass for paywalled publishers without a public TDM API (ACS, RSC, IEEE, AIP, IOP, APS, ...); failed-DOI tracking with structured residual-failure summary; pre-packaged agent skill at .claude/skills/paper-download/ . --- 🚀 Quick start For paywalled publishers without a TDM API (ACS, RSC, IEEE, AIP, IOP, APS): PDFs land under downloads/pdfs/ / . Re-running skips files already on disk. The repo distinguishes two audiences — the rest of this README has one section for each. Both share the same install and the same .env ; the only thing that differs is how you invoke the tool . --- 👤 For Humans You're a researcher or developer who wants to run this on your own DOI list. Install Alternative: uv sync if you prefer uv over pip . Configure credentials At minimum set ONE of CROSSREF MAILTO / OPENALEX MAILTO to a real email — public APIs require this for","default_branch":null,"files":null,"tree":[],"storefront":"/r/jxtse","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/jxtse/auto-paper-harvester/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}