Open datasets built from the ThreatCluster corpus: ransomware leak-site activity, CVE exploitation signals, and deduplicated security incidents. 100,711 rows in total, refreshed periodically.
Also on the Hugging Face Hub and Kaggle. What each one covers and where it falls short: threatcluster.io/datasets.
| Dataset | Contents | Rows | Mirrors |
|---|---|---|---|
ransomware-leak-site-victims |
Ransomware leak-site victims (live) | 20,627 | Hugging Face · Kaggle |
cve-exploitation-signals |
CVE exploitation signals (live) | 60,879 | Hugging Face · Kaggle |
threat-incident-clusters |
Threat incident clusters (live) | 19,205 | Hugging Face · Kaggle |
Each directory holds newline-delimited JSON (data.jsonl, one object per line)
and a dataset card describing every field and its caveats. Read the card before
using the data — each set has limits that matter.
import json
rows = [json.loads(l) for l in open("cve-exploitation-signals/data.jsonl")]
# or straight from the Hub
from datasets import load_dataset
ds = load_dataset("threatcluster/cve-exploitation-signals")Victim listings collected first-hand from ransomware and extortion leak sites: which group named which organisation, when, in what sector and country.
ransomware-leak-site-victims/data.jsonl — 20,627 rows. Fields and caveats: ransomware-leak-site-victims/README.md.
One row per CVE joining reference data (CVSS, CWE, affected products) with exploitation evidence: CISA KEV, public exploit availability and ransomware use.
cve-exploitation-signals/data.jsonl — 60,879 rows. Fields and caveats: cve-exploitation-signals/README.md.
Security incidents as deduplicated stories rather than individual articles, with the entities involved and the outlets that reported each one.
threat-incident-clusters/data.jsonl — 19,205 rows. Fields and caveats: threat-incident-clusters/README.md.
Victim screenshots and enrichment, leak-site post URLs, the extortion copy groups write about their victims, and the full text of source articles. The first three carry personal data or belong behind the paid API; the last is third-party copyright. Only ThreatCluster's own summaries, scores and metadata are published, alongside links back to the original reporting.
- Leak-site listings are claims, not confirmed breaches. Groups name organisations that never paid, re-list old victims and occasionally fabricate.
- Exploitation labels are observed, not exhaustive.
in_kevandhas_exploitmean "known to us"; absence is not evidence of absence. Both are time-dependent, so respectpublished_datewhen splitting train and test data. - Incident titles and summaries are model-generated and not human-verified.
threat_scoreis a ranking signal, not a severity scale. Significant incidents routinely score in the 20s.
- Hugging Face — https://huggingface.co/threatcluster (loads with
datasets.load_dataset) - Kaggle — https://www.kaggle.com/datasets/threatcluster/threat-intelligence-datasets (all three as one dataset)
- This repository — raw
data.jsonlper dataset - Live, queryable — the ThreatCluster API, free tier available
Data: CC BY 4.0 — free to use and redistribute, including
commercially, with attribution to ThreatCluster. The GPL-3.0 LICENSE file
covers code, not data.
Published to support defensive security research, measurement and education. The organisations named here are the injured parties in criminal attacks. Do not use this data to target, harass or profile them.
These are periodic snapshots. The live corpus is available through the ThreatCluster API, which has a free tier, and through the public feeds at threatcluster.io/feeds.