{"repo":"currentslab/extractnet","free":true,"listed":false,"github":"https://github.com/currentslab/extractnet","clone":"git clone https://github.com/currentslab/extractnet.git","description":"A fork of Dragnet that also extract author, headline, date, keywords from context, as well as built in metadata extraction all in one package","language":"HTML","stars":300,"topics":["content-extraction","author-extraction","date-extraction","webscraping","web-scraping","text-cleaning","text-mining","news-extractor","news-extraction","news"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"ExtractNet ======= Based on the popular content extraction package Dragnet, ExtractNet extend the machine learning approach to extract other attributes such as date, author and keywords from news article. Example code: Simply use the following command to install the latest released version: Start extract content and other meta data passing the result html to function Why don't just use existing rule-base extraction method: We discover some webpage doesn't provide the real author name but simply populate the author tag with a default value. For example ltn.com.tw, udn.com always populate the same author value for each news article while the real author can only be found within the content. ExtractNet uses machine learning approach to extract these relevant data through visible section of the webpage just like a human. What ExtractNet is and isn't ExtractNet is a platform to extract any interesting attributes from any webpage, not just limited to content based article. The core of ExtractNet aims to convert unstructured webpage to structured data without relying hand crafted rules ExtractNet do not support boilerplate content extraction ExtractNet allows user to add custom pipelines that returns additional data through a list of callbacks function Performance Results of the body extraction evaluation: We use the same body extraction benchmark from article-extraction-benchmark Model Precision Recall F1 Accuracy Open Source --- --- --- --- --- --- AutoExtract 0.984 ± 0.003 0.956 ","default_branch":null,"files":null,"tree":[],"storefront":"/r/currentslab","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/currentslab/extractnet/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}