Papers › Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for...

Izindaba-Tindzaba: Machine learning news categorisation for Long and Short Text for isiZulu and Siswati

12 Jun 2023arXiv:2306.07426archive 2025-07-28

Andani Madodonga, Vukosi Marivate, Matthew Adendorff

Local/Native South African languages are classified as low-resource languages. As such, it is essential to build the resources for these languages so that they can benefit from advances in the field of natural language processing. In this work, the focus was to create annotated news datasets for the isiZulu and Siswati native languages based on news topic classification tasks and present the findings from these baseline classification models. Due to the shortage of data for these native South African languages, the datasets that were created were augmented and oversampled to increase data size and overcome class classification imbalance. In total, four different classification models were used namely Logistic regression, Naive bayes, XGBoost and LSTM. These models were trained on three different word embeddings namely Bag-Of-Words, TFIDF and Word2vec. The results of this study showed that XGBoost, Logistic Regression and LSTM, trained from Word2vec performed better than the other combinations.

PaperPDFCode

Code

dsfsi/za-isizulu-siswati-news-2022 officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationTopic ClassificationWord Embeddingsregression

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

FocusLSTMLogistic RegressionSigmoid ActivationTanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections