HATEBENCH

Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns

Conference Paper (2025)
Author(s)

Xinyue Shen (CISPA Helmholtz Center for Information Security)

Yixin Wu (CISPA Helmholtz Center for Information Security)

Yiting Qu (CISPA Helmholtz Center for Information Security)

Michael Backes (CISPA Helmholtz Center for Information Security)

Savvas Zannettou (TU Delft - Technology, Policy and Management)

Yang Zhang (CISPA Helmholtz Center for Information Security)

Research Group
Organisation & Governance
More Info
expand_more
Publication Year
2025
Language
English
Research Group
Organisation & Governance
Pages (from-to)
221-240
Publisher
USENIX Association
ISBN (electronic)
9781939133526
Event
34th USENIX Security Symposium, USENIX Security 2025 (2025-08-13 - 2025-08-15), Seattle, United States
Page Views
27
Reuse Rights

Other than for strictly personal use, it is not permitted to download, forward or distribute the text or part of it, without the consent of the author(s) and/or copyright holder(s), unless the work is under an open content license such as Creative Commons.

Abstract

Large Language Models (LLMs) have raised increasing concerns about their misuse in generating hate speech. Among all the efforts to address this issue, hate speech detectors play a crucial role. However, the effectiveness of different detectors against LLM-generated hate speech remains largely unknown. In this paper, we propose HATEBENCH, a framework for benchmarking hate speech detectors on LLM-generated hate speech. We first construct a hate speech dataset of 7,838 samples generated by six widely-used LLMs covering 34 identity groups, with meticulous annotations by three labelers. We then assess the effectiveness of eight representative hate speech detectors on the LLM-generated dataset. Our results show that while detectors are generally effective in identifying LLM-generated hate speech, their performance degrades with newer versions of LLMs. We also reveal the potential of LLM-driven hate campaigns, a new threat that LLMs bring to the field of hate speech detection. By leveraging advanced techniques like adversarial attacks and model stealing attacks, the adversary can intentionally evade the detector and automate hate campaigns online. The most potent adversarial attack achieves an attack success rate of 0.966, and its attack efficiency can be further improved by 13−21× through model stealing attacks with acceptable attack performance. We hope our study can serve as a call to action for the research community and platform moderators to fortify defenses against these emerging threats.

Files

Usenixsecurity25-shen.pdf
(pdf | 12.3 Mb)
License info not available