The Risk of Racial Bias in Hate Speech Detection
SociolinguisticsDiscourse & deception
The short version
This study examines how dialect-insensitive annotation can introduce racial bias into hate-speech datasets and the models trained on them. It also tests whether making annotators aware of dialect reduces their tendency to label African-American English as offensive.
Notes for the reading desk
- Annotations can carry biases into downstream language models.
- A surface linguistic feature can be mistaken for an indicator of offensiveness.
- Dialect awareness is relevant to the design and interpretation of annotation tasks.
Citation
Sap, M., Card, D., Gabriel, S., Choi, Y., & Smith, N. A. (2019). The Risk of Racial Bias in Hate Speech Detection. Proceedings of ACL, 1668–1678. https://doi.org/10.18653/v1/P19-1163
Summary and reading notes are editorial guides. The linked paper is the original source.