IRCI Article ID: IRCI-AR-0000002623

Detection of Hate Speech Code Mix Involving English and Other Nigerian Languages

Journal: Journal of Information Systems and Informatics

Publication: 2023-12-02 · Vol. 5 No. 4 · pp. 1416–1431

DOI: 10.51519/journalisi.v5i4.595

Cite this article

Citation

Choose a citation style or copy BibTeX for your reference manager.

 
View Original Publication

Abstract

Hate speech is a recurrent event and has become a cause for global concern. The proliferation of hate speech has recently become prevalent, breeding room for violence and discrimination against specific individuals or groups. In Nigeria, message masking (use of language-mix) has become the new normal, especially in disseminating hateful and inciting comments. Hence, there is a need to curb the spread over social media. Therefore, this research focuses on detecting hate speech on social media with a code-mix of English, Pidgin and any of the three major Nigerian languages (Hausa, Igbo and Yoruba). The research used two machine learning algorithms: Support Vector Machine (SVM) and Random Forest (RF). Data were collected from tweets on the EndSARS protest and the 2023 Nigerian elections. The major features were extracted, and the text was converted into vectors using TF-IDF and Bag-of-words (BoW), which were used to train and test the model. The result showed that SVM performed better in classifying hate speech than RF on both TF-IDF and BoW features, averaging 93.43% for accuracy, 93.70% for precision, 93.43% for recall, and 93.57% for F1-score.

0
IRCI Cited By
0
Indexed References

Authors

References

No references were harvested yet.

Cited By (0)

No indexed citing article has been matched by IRCI yet.