Papers › AdaptText: A Novel Framework for Domain-Independent Automated Sinhala Text Classification

AdaptText: A Novel Framework for Domain-Independent Automated Sinhala Text Classification

13 Nov 202110th International Conference on Information and Automation for Sustainability (ICIAfS) 2021 11archive 2025-07-28

Yathindra Kodithuwakku, Saman Hettiarachchi

Sinhala language is being the widely used language in Sri Lanka. With the advancement of internet usage in Sri Lanka, an incredible amount of Sinhala text data is being added to the internet. In order to manage, analyze and make decisions from the available text data, it requires text classification. Being a low resource and morphologically rich language requires higher expertise and a considerable amount of budget and time to develop an effective task-specific text classifier. This research aims to develop a domain or dataset agnostic and automated solution to improve the quality and address current research gaps of text classification in Sinhala. Based on the solution, a high-level development framework and a user interface are developed. In addition, we perform a cross-domain evaluation with multiple datasets to evaluate the effectiveness of the solution. The proposed framework achieved state-of-the-art results for the Sinhala text classification.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationText Classificationtext-classification

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections