- Install feedparser
pip install feedparser- Run the script either locally or https://colab.research.google.com/drive/16O4qd4nMc0UWtbmjReHpyaDkmhQSHRYy?usp=sharing
- Language Detection (via regex because there weren't many different languages, can be improved)
- Storing data in Database with duplicate detection.
- Cron jobs can be easily implemented on server directly with 1 line.
- Flask was not implemented because testing was done on colab for easier client access.
Starting Global News RSS Scraper...
Configured to scrape from 20 countries
Saving data...
Total Articles: 1,578
Countries Covered: 20
News Sources: 30
Date Range: 2017-03-30T00:00:00 to 2025-05-28T16:15:14+10:00
- Russia: 349 articles
- Germany: 149 articles
- Argentina: 110 articles
- Brazil: 100 articles
- South Korea: 100 articles
- Ukraine: 100 articles
- Singapore: 82 articles
- South Africa: 80 articles
- United Kingdom: 78 articles
- Pakistan: 70 articles
- India: 68 articles
- China: 61 articles
- United States: 50 articles
- Australia: 45 articles
- Japan: 37 articles
- Middle East: 25 articles
- France: 24 articles
- Canada: 20 articles
- Nigeria: 20 articles
- Poland: 10 articles
- Lenta.ru : ΠΠΎΠ²ΠΎΡΡΠΈ: 199 articles
- Deutsche Welle: 149 articles
- Folha: 100 articles
- La Nacion: 100 articles
- News Agency UNIAN: 100 articles
- RT: 100 articles
- Yonhap News: 100 articles
- The Express Tribune: 70 articles
- Straits Times: 62 articles
- Daily Maverick: 60 articles
- English (en): 1,147 articles
- Basque (eu): 223 articles
- Russian (ru): 200 articles
- Chinese (zh): 8 articles
- Database:
news_data/news_database.db - CSV:
news_data/news_articles_20250528_062335.csv - JSON:
news_data/news_articles_20250528_062335.json
β Scraping completed successfully!