Skip to content

Sink Health Monitoring: Test Notifications, Alert Cleanup & Polish #11

Description

@joejohnson123

Overview

Final polish tasks: test notification functionality, alert history cleanup, and overall integration testing.

Depends on: #3 (Worker), #4 (Dispatcher), #7 (Anomaly), #8 (UI)

Test Notifications

Feature

Allow users to send a test alert through all configured channels to verify their setup works before waiting for a real alert.

Implementation

  • Add Sequin.Monitors.send_test_notification(monitor) function
  • Creates a fake alert payload with:
    • status: "test"
    • message: "This is a test notification from monitor '{name}'"
    • metric_value: 0
    • Current timestamp
  • Dispatches through all configured channels
  • Returns per-channel success/failure results
  • Does NOT create an alert record in the DB

UI Integration

  • "Test" button on monitor form (after saving) and monitor list cards
  • Shows toast with results: "✅ Slack: sent, ❌ Discord: HTTP 403"

API Integration

Alert History Cleanup

Oban Worker: Sequin.Monitors.PruneAlertsWorker

  • Runs daily (Oban cron)
  • Deletes resolved alerts older than 30 days
  • Deletes triggered (stale/abandoned) alerts older than 90 days
  • Logs count of pruned records

Metric Sample Cleanup

  • Already handled by AnomalyDetector pruning (2-hour window)
  • Add a safety net: PruneAlertsWorker also prunes samples older than 4 hours

Integration Testing

End-to-End Test Scenarios

  1. Threshold monitor lifecycle:

    • Create sink with monitor (threshold: failed_messages > 10, window: 0s for testing)
    • Simulate failed messages accumulating
    • Verify alert fires and webhook receives notification
    • Simulate messages clearing
    • Verify alert resolves and recovery notification sent
  2. Anomaly monitor lifecycle:

    • Create sink with anomaly monitor
    • Seed 30+ metric samples at low values
    • Inject a spike value
    • Verify anomaly detected and alert fires
    • Return to normal values
    • Verify alert resolves
  3. Cooldown behavior:

    • Trigger an alert
    • Verify second evaluation within cooldown doesn't re-alert
    • Wait past cooldown
    • Verify next evaluation does fire
  4. Multi-channel dispatch:

    • Configure monitor with 2+ channels
    • Verify all channels receive notification
    • Simulate one channel failing
    • Verify other channels still receive notification
  5. Monitor disable/enable:

    • Disable a monitor
    • Verify it's skipped during evaluation
    • Re-enable
    • Verify it resumes evaluation

Files to Create

  • lib/sequin/monitors/prune_alerts_worker.ex
  • test/sequin/monitors/integration_test.exs

Files to Modify

  • lib/sequin/monitors/monitors.ex — add send_test_notification/1
  • lib/sequin/monitors/notification_dispatcher.ex — handle test alert type
  • lib/sequin/application.ex — register PruneAlertsWorker Oban cron

Acceptance Criteria

  • Test notification sends to all channels and reports results
  • Test notification clearly marked as test (not confused with real alerts)
  • Alert pruning runs daily and cleans up old records
  • Sample pruning safety net works
  • All 5 integration test scenarios pass
  • No extra logs generated by tests
  • No Process.sleep in tests

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions