Back to the bookshelf
ACL2019

The Risk of Racial Bias in Hate Speech Detection

Maarten Sap · Dallas Card · Saadia Gabriel · Yejin Choi · Noah A. Smith

SociolinguisticsDiscourse & deception

The short version

This study examines how dialect-insensitive annotation can introduce racial bias into hate-speech datasets and the models trained on them. It also tests whether making annotators aware of dialect reduces their tendency to label African-American English as offensive.

Notes for the reading desk

  1. Annotations can carry biases into downstream language models.
  2. A surface linguistic feature can be mistaken for an indicator of offensiveness.
  3. Dialect awareness is relevant to the design and interpretation of annotation tasks.

Explore the themes

Citation

Sap, M., Card, D., Gabriel, S., Choi, Y., & Smith, N. A. (2019). The Risk of Racial Bias in Hate Speech Detection. Proceedings of ACL, 1668–1678. https://doi.org/10.18653/v1/P19-1163

Summary and reading notes are editorial guides. The linked paper is the original source.