Papers › Clustering Urdu News Using Headlines
Clustering Urdu News Using Headlines
Samia Khaliq, Waheed Iqbal, Faisal Bukhari, Kamran Malik
This paper that proposes and evaluates a new algorithm to automatically cluster Urdu news from different news agencies. The task is challenging because there are no language processing libraries for the Urdu language. The authors' experimental dataset consists of news from famous Pakistani media houses, including Jang, BBC Urdu, Express, UrduPoint, and Voice of America Urdu (VOA). The proposed algorithm only uses headlines to cluster the news. The authors argue that news headlines provide a concise summary of the news, which motivates them to use it instead of using the entire news story. Their experimental evaluation shows micro and macro averages for precision of 0.45 and 0.48 respectively for identifying similar news using headlines.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Datasets
Introduced by this paper, per the archive.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Text Clustering | Urdu News Headlines Dataset | Vector Space Model | Related Headlines | 85 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections