论文标题
哪一种更具毒性?拼图的发现有毒评论的严重程度
Which one is more toxic? Findings from Jigsaw Rate Severity of Toxic Comments
论文作者
论文摘要
在线仇恨言论的扩散需要创建可以检测毒性的算法。过去的大多数研究都集中在这一发现作为分类任务上,但是分配绝对毒性标签通常很棘手。因此,过去很少有作品将相同的任务转变为回归。本文显示了拼图的最近发布的毒性严重程度测量数据集上对不同变压器和传统机器学习模型的比较评估。我们进一步使用解释性分析来证明模型预测的问题。
The proliferation of online hate speech has necessitated the creation of algorithms which can detect toxicity. Most of the past research focuses on this detection as a classification task, but assigning an absolute toxicity label is often tricky. Hence, few of the past works transform the same task into a regression. This paper shows the comparative evaluation of different transformers and traditional machine learning models on a recently released toxicity severity measurement dataset by Jigsaw. We further demonstrate the issues with the model predictions using explainability analysis.