{"repo":"flairNLP/fundus","free":true,"listed":false,"github":"https://github.com/flairNLP/fundus","clone":"git clone https://github.com/flairNLP/fundus.git","description":"A very simple news crawler with a funny name","language":"Python","stars":473,"topics":["cc-news","commoncrawl","corpus","crawler","news-crawler","news-scraping","nlp","python","rss","scraper"],"license":"MIT","category":"scrapers-browser-automation","readme_excerpt":"A very simple news crawler in Python. Developed at Humboldt University of Berlin . Quick Start Tutorials News Sources Paper --- Disclaimer : Although we try to provide an indication of whether a publisher has not explicitly objected to the training of AI models on its data, we would like to point out that this information must be verified independently before their content is used. More details can be found here. --- Fundus is: A static news crawler. Fundus lets you crawl online news articles with only a few lines of Python code! Be it from live websites or the CC-NEWS dataset. An open-source Python package. Fundus is built on the idea of building something together. We welcome your contribution to help Fundus grow! Quick Start To install from pip, simply do: Fundus requires Python 3.8+. Example 1: Crawl a bunch of English-language news articles Let's use Fundus to crawl 2 articles from publishers based in the US. That's already it! If you run this code, it should print out something like this: This printout tells you that you successfully crawled two articles! For each article, the printout details: - the number of images included in the article - the \"Title\" of the article, i.e. its headline - the \"Text\", i.e. the main article body text - the \"URL\" from which it was crawled - the news source it is \"From\" Example 2: Crawl a specific news source Maybe you want to crawl a specific news source instead. Let's crawl news articles from The New Yorker only: Example 3: Crawl 1 Milli","default_branch":null,"files":null,"tree":[],"storefront":"/r/flairNLP","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/flairNLP/fundus/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}