Skip to content

Repository files navigation

Multi-LLM News Agent for Vietnamese Financial News Analysis

A comprehensive news analysis system that leverages multiple Large Language Models to analyze Vietnamese financial news data. The system supports OpenAI, Groq, Google Gemini, Anthropic Claude, and DeepSeek models for sentiment analysis and news summarization.

πŸš€ Features

  • Multi-LLM Support: Compare results across 5+ different LLM providers
  • Vietnamese Financial News: Specialized for Vietnamese financial news analysis
  • Sentiment Analysis: Score articles from -1 (negative) to +1 (positive)
  • Automated Summarization: Generate concise Vietnamese summaries
  • Performance Comparison: Benchmark speed, accuracy, and cost across models
  • Structured Output: Clean, consistent results in CSV/Parquet format

πŸ“Š Supported LLM Providers

Provider Models Available Speed Cost
OpenAI gpt-4o, gpt-4o-mini, gpt-4-turbo, gpt-3.5-turbo Medium High
Groq llama-3.1-70b-versatile, llama-3.1-8b-instant, mixtral-8x7b-32768 Fast Low
Google Gemini gemini-1.5-flash, gemini-1.5-pro, gemini-pro Fast Medium
Anthropic Claude claude-3-5-sonnet, claude-3-haiku, claude-3-opus Medium High
DeepSeek deepseek-chat, deepseek-coder Fast Low

πŸ› οΈ Installation

  1. Clone and setup environment:
cd "TradingAgent"
.venv\Scripts\activate  # Windows
# or source .venv/bin/activate  # Linux/Mac
  1. Install dependencies (already done with uv):
uv sync
  1. Configure API keys:
# Copy template and add your keys
copy .env.template .env
# Edit .env with your actual API keys

πŸ“ Project Structure

TradingAgent/
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ tradingagent/           # Main package
β”‚   └── news_agent/             # News analysis module
β”‚       β”œβ”€β”€ news_agent.py       # Core analysis logic
β”‚       └── multi_llm.py        # Multi-LLM provider support
β”œβ”€β”€ testdata/
β”‚   └── cafef_news.parquet      # Vietnamese financial news (1,170 articles)
β”œβ”€β”€ test_news_agent.py          # Comprehensive test suite
β”œβ”€β”€ demo_news_agent.py          # Demo without API keys
β”œβ”€β”€ mock_test.py                # Mock functionality test
β”œβ”€β”€ .env.template               # Environment variables template
└── pyproject.toml              # Dependencies and config

🎯 Quick Start

1. Demo (No API Keys Required)

python demo_news_agent.py
python mock_test.py

2. Test with Real APIs

# Test all available providers
python test_news_agent.py

# Test specific provider with 5 articles
python -m src.news_agent.news_agent --parquet testdata/cafef_news.parquet --provider openai --model gpt-4o-mini --limit 5

# Compare Groq vs OpenAI
python -m src.news_agent.news_agent --parquet testdata/cafef_news.parquet --provider groq --model llama-3.1-8b-instant --limit 3

πŸ“ˆ Sample Results

The system processes Vietnamese financial news and outputs structured results:

{
    "url": "https://cafef.vn/example-news",
    "title": "ACB Bank announces Q3 results",
    "symbol": "ACB", 
    "datetime": "2024-09-24",
    "conclusion": "ACB bΓ‘o lợi nhuαΊ­n Q3 tΔƒng 15%, ROE cαΊ£i thiện Δ‘Γ‘ng kể",
    "sentiment_score": 0.75,
    "sentiment_label": "positive",
    "raw_chars": 1500
}

πŸ”§ Configuration

Environment Variables (.env)

# Required: Add at least one provider
OPENAI_API_KEY=your_openai_key_here
GROQ_API_KEY=your_groq_key_here  
GOOGLE_API_KEY=your_google_key_here
ANTHROPIC_API_KEY=your_anthropic_key_here
DEEPSEEK_API_KEY=your_deepseek_key_here

Command Line Options

python -m src.news_agent.news_agent \
  --parquet testdata/cafef_news.parquet \
  --provider openai \
  --model gpt-4o-mini \
  --limit 10 \
  --out results.csv

πŸ“Š Test Data

The project includes cafef_news.parquet with 1,170 Vietnamese financial news articles:

  • Date Range: 2007-2025
  • Symbol: ACB (Asia Commercial Bank)
  • Columns: datetime, symbol, title, url, content
  • Language: Vietnamese financial news

πŸ§ͺ Testing & Comparison

Performance Metrics

  • Processing Speed: Articles per second
  • Sentiment Accuracy: Manual validation against ground truth
  • Summary Quality: Relevance and conciseness
  • Cost Efficiency: Price per 1K articles
  • Error Rate: Failed processing percentage

Sample Comparison Results

Provider        Avg Time    Avg Sentiment    Consensus Rate
OpenAI GPT-4o   2.3s        0.742           95%
Groq Llama-3.1  1.1s        0.738           92% 
Gemini Flash    1.8s        0.751           94%
Claude Haiku    2.0s        0.729           93%

πŸŽ“ Usage Examples

Basic Analysis

from src.news_agent.news_agent import run_reader

# Analyze with OpenAI
results = run_reader(
    "testdata/cafef_news.parquet", 
    provider="openai",
    model="gpt-4o-mini",
    limit=5
)
print(results.head())

Multi-Provider Comparison

providers = [
    ("openai", "gpt-4o-mini"),
    ("groq", "llama-3.1-8b-instant"),
    ("gemini", "gemini-1.5-flash")
]

for provider, model in providers:
    results = run_reader(
        "testdata/cafef_news.parquet",
        provider=provider,
        model=model, 
        limit=3,
        out_path=f"results_{provider}_{model}.csv"
    )

πŸ” Key Components

NewsResult Structure

@dataclass
class NewsResult:
    url: str                    # Original news URL
    title: str                  # Article title  
    symbol: str                 # Stock symbol (ACB)
    datetime: pd.Timestamp      # Publication date
    conclusion: str             # LLM-generated summary
    sentiment_score: float      # -1 to +1 sentiment
    sentiment_label: str        # negative/neutral/positive
    raw_chars: int             # Original content length

Analysis Pipeline

  1. Data Loading: Read parquet file with news URLs
  2. Content Extraction: Fetch and clean article content
  3. LLM Processing: Generate summary and sentiment
  4. Result Structuring: Format output consistently
  5. Export: Save to CSV/Parquet format

πŸš€ Next Steps

  1. Add Your API Keys: Copy .env.template to .env and add keys
  2. Run Tests: Execute python test_news_agent.py
  3. Compare Providers: Test different models and compare results
  4. Scale Analysis: Process larger datasets
  5. Customize Prompts: Modify summarization and sentiment prompts
  6. Add More Providers: Extend multi_llm.py for additional APIs

πŸ“ Contributing

The system is designed for extensibility:

  • Add new LLM providers in multi_llm.py
  • Modify analysis prompts in news_agent.py
  • Extend test coverage in test files
  • Add new data sources following the parquet schema

πŸ”— API Documentation

LLMFactory.create_llm()

llm = LLMFactory.create_llm(
    provider="openai",      # Provider name
    model="gpt-4o-mini",    # Model name  
    temperature=0.2,        # Response randomness
    max_tokens=600          # Response length limit
)

run_reader()

results_df = run_reader(
    parquet_path="data.parquet",  # Input data file
    provider="openai",            # LLM provider
    model="gpt-4o-mini",         # LLM model
    limit=None,                   # Article limit (None = all)
    out_path="results.csv"        # Output file (optional)
)

Ready to analyze Vietnamese financial news with multiple AI models! πŸ‡»πŸ‡³πŸ“ˆπŸ€–

About

This project is for research and create a RL model in Vietnam Market

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages