{"repo":"SecondDim/crawler-news","free":true,"listed":false,"github":"https://github.com/SecondDim/crawler-news","clone":"git clone https://github.com/SecondDim/crawler-news.git","description":"Use python scrapy build crawler for real-time Taiwan NEWS website.","language":"Python","stars":13,"topics":["crawler","taiwan-news-website","python","news","scrapy","taiwan","news-crawler","docker","docker-compose","database"],"license":"MIT","category":"deployment-docker-iac","readme_excerpt":"Python Crawler for News A real-time news crawler built with Python and Scrapy, designed to fetch latest news from major Taiwan news websites. 使用 Python Scrapy 建置的台灣新聞網站即時爬蟲。 Features - Real-time Crawling : Fetches the latest news updates efficiently. - Multiple Sources : Supports various major Taiwan news outlets. - Flexible Storage : Supports saving data to Cassandra, MySQL, or JSON (via pipelines). - Extensible : Easy to add new spiders for additional news sites. - Docker Support : Ready-to-use Docker configuration for easy deployment. Supported News Sites News Site Website Status ----------- --------- -------- 自由時報 (Liberty Times) Website ✅ Active 東森新聞 (EBC) Website ✅ Active 聯合新聞網 (UDN) Website ✅ Active 今日新聞 (NOWnews) Website ✅ Active ETtoday Website ✅ Active 中時電子報 (China Times) Website ✅ Active TVBS Website ✅ Active 三立新聞網 (SETN) Website ✅ Active 中央通訊社 (CNA) Website ✅ Active 巴哈姆特 (Gamer) Website 🚧 TODO 風傳媒 (Storm) Website 🚧 TODO Requirements - Python 3.9+ - Redis (for deduplication and queue management) - Database (Optional): Cassandra or MySQL Installation 1. Clone the repository 2. Install dependencies 3. Configuration Copy the example settings file and configure your environment: Edit crawler news/settings.py to set up your database connections (Redis, MySQL, Cassandra) and other preferences. Usage Run All Spiders To run all spiders sequentially: Run a Single Spider To run a specific spider (e.g., ettoday ): Available spiders: chinatimes , cna , ebc , ettoday , libert","default_branch":null,"files":null,"tree":[],"storefront":"/r/SecondDim","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/SecondDim/crawler-news/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}