Congratulations and Thank You for This Incredible Work!
I just came across RedSage and I have to say — this is one of the most impressive cybersecurity LLM projects I've seen. The combination of the 11.8B-token CyberFineWeb corpus, the agentic multi-turn dialogue pipeline, and the comprehensive RedSage-Bench evaluation suite is genuinely exciting. The +5.59-point improvement over Llama-3.1-8B and Qwen3-8B on cyber-benchmarks speaks for itself. Congratulations on the ICLR 2026 acceptance — well deserved!
Request 1: Further Data Release
I noticed from the release checklist that some datasets are still pending, including:
- RedSage-CFW (CyberFineWeb filtered corpus)
- RedSage-Seed curated sources
- RedSage-Conv multi-turn dialogues
- The cybersecurity-filtering code and agentic data augmentation scripts
Could you share an estimated timeline for releasing these? The agentic pipeline that generates multi-turn "User-Expert" dialogues seems particularly novel and impactful for the community. Having access to the data generation methodology — even if the full corpus can't be released due to licensing — would be incredibly valuable for reproducibility and follow-up research.
Request 2: Evaluation Code & Baseline Results
Similarly, the evaluation suite (RedSage-MCQ, RedSage-OpenQA, and the LLM-as-judge rubric) is central to validating and comparing future work in this space. Could you also prioritize releasing:
- The RedSage-OpenQA data and lighteval implementation (currently pending per the checklist)
- The baseline results table with RedSage variants and common 8B baselines
- The LLM-as-judge rubric details
Having a standardized, reproducible benchmark would greatly help the broader cybersecurity NLP community build on top of RedSage-Bench.
I understand releasing everything takes time and effort, and I deeply appreciate what has already been made available. Just wanted to flag these as high-priority from a community perspective.
Thank you again for this outstanding contribution — looking forward to seeing how RedSage evolves! 🙏
Congratulations and Thank You for This Incredible Work!
I just came across RedSage and I have to say — this is one of the most impressive cybersecurity LLM projects I've seen. The combination of the 11.8B-token CyberFineWeb corpus, the agentic multi-turn dialogue pipeline, and the comprehensive RedSage-Bench evaluation suite is genuinely exciting. The +5.59-point improvement over Llama-3.1-8B and Qwen3-8B on cyber-benchmarks speaks for itself. Congratulations on the ICLR 2026 acceptance — well deserved!
Request 1: Further Data Release
I noticed from the release checklist that some datasets are still pending, including:
Could you share an estimated timeline for releasing these? The agentic pipeline that generates multi-turn "User-Expert" dialogues seems particularly novel and impactful for the community. Having access to the data generation methodology — even if the full corpus can't be released due to licensing — would be incredibly valuable for reproducibility and follow-up research.
Request 2: Evaluation Code & Baseline Results
Similarly, the evaluation suite (RedSage-MCQ, RedSage-OpenQA, and the LLM-as-judge rubric) is central to validating and comparing future work in this space. Could you also prioritize releasing:
Having a standardized, reproducible benchmark would greatly help the broader cybersecurity NLP community build on top of RedSage-Bench.
I understand releasing everything takes time and effort, and I deeply appreciate what has already been made available. Just wanted to flag these as high-priority from a community perspective.
Thank you again for this outstanding contribution — looking forward to seeing how RedSage evolves! 🙏