{"repo":"myklovenyzforever/chem-pdf-extractor","free":true,"listed":false,"github":"https://github.com/myklovenyzforever/chem-pdf-extractor","clone":"git clone https://github.com/myklovenyzforever/chem-pdf-extractor.git","description":"LLM-powered scientific PDF data extraction tool for chemistry and chemical engineering literature.","language":"Python","stars":15,"topics":["catalysis","chemical-engineering","data-extraction","excel-export","literature-review","llm","materials-science","ollama","pdf-extraction","research-tools"],"license":"MIT","category":"ai-agents","readme_excerpt":"Chem-PDF-Extractor Chem-PDF-Extractor is an open-source tool for extracting structured experimental data from chemical engineering, catalysis, materials, energy, and environmental research PDFs into Excel/CSV tables. It is designed for literature reviews, preliminary dataset construction, and manual verification workflows. Chem-PDF-Extractor 是一个面向化工、催化、材料、能源与环境领域论文的开源 PDF 数据抽取工具，可将文献中的实验条件和结果整理为 Excel/CSV 表格，适用于文献综述、初步数据集构建和后续人工核验流程。 Language / 语言: English 中文 Overview The project provides an inspectable, local-first workflow for converting PDF papers to Markdown/text, defining configurable LLM-based extraction fields, extracting multiple records from one paper, and exporting structured results for review. It supports local Ollama models and optional OpenAI-compatible cloud APIs, but cloud use is opt-in and requires the user to provide their own API configuration. The output is intended as first-pass extraction for literature review and dataset preparation. It should be checked against the original paper before being used for scientific conclusions. Why This Project Experimental information in scientific papers is often scattered across paragraphs, tables, figures, captions, and supplementary materials. Manually collecting feedstocks, catalysts, reaction temperature, pressure, conversion, selectivity, yield, and product information is slow and error-prone. This project provides a local, inspectable workflow for building first-pass structured datasets that still require human r","default_branch":null,"files":null,"tree":[],"storefront":"/r/myklovenyzforever","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/myklovenyzforever/chem-pdf-extractor/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}