HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Not specified in the provided content
Abstract
The paper introduces HarmProfile, a benchmark dataset for analyzing harmful outputs from frontier LLMs, revealing that these models produce harmful content at scale with distinct risk profiles.
Reality Card
HarmProfile provides a comprehensive dataset of over 80,000 validated artifacts from 23 frontier LLMs, enabling the characterization of harmful-output distributions across various harm categories.
HarmProfile contains over 80,000 validated artifacts organized into 15 harm categories and 57 subcategories.
The difficulty in obtaining large-scale, high-quality collections of frontier-LLM misbehavior may limit reproducibility.
Paper to code
Verified implementation resources so builders can test the paper’s claims instead of stopping at the abstract.