{"repo":"fhamborg/news-please","free":true,"listed":false,"github":"https://github.com/fhamborg/news-please","clone":"git clone https://github.com/fhamborg/news-please.git","description":"news-please - an integrated web crawler and information extractor for news that just works","language":"Python","stars":2479,"topics":["news-crawler","news-extractor","crawler","extractor","news","news-websites","elasticsearch","json","python","nlp"],"license":"Apache-2.0","category":"scrapers-browser-automation","readme_excerpt":"news-please # news-please is an open source, easy-to-use news crawler that extracts structured information from almost any news website. It can recursively follow internal hyperlinks and read RSS feeds to fetch both most recent and also old, archived articles. You only need to provide the root URL of the news website to crawl it completely. news-please combines the power of multiple state-of-the-art libraries and tools, such as scrapy, Newspaper, and readability. news-please also allows Python developers to use the crawling and extraction functionality within their own program. Moreover, news-please allows to conveniently crawl and extract articles from the (very) large news archive at commoncrawl.org. If you want to contribute to news-please, please first read here. Extracted information news-please extracts the following attributes from news articles. An examplary json file as extracted by news-please can be found here. headline lead paragraph main text main image name(s) of author(s) publication date language Features works out of the box : install with pip, add URLs of your pages, run :-) run news-please conveniently using its CLI mode use it as a library within your own software extract articles from commoncrawl.org's news archive Modes and use cases news-please supports three main use cases, which are explained in more detail in the following. CLI mode stores extracted results in JSON files, PostgreSQL, ElasticSearch, Redis, or your own storage simple but extensive conf","default_branch":null,"files":null,"tree":[],"storefront":"/r/fhamborg","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/fhamborg/news-please/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}