Papers › TFW2V: An Enhanced Document Similarity Method for the Morphologically Rich Finnish Language

TFW2V: An Enhanced Document Similarity Method for the Morphologically Rich Finnish Language

23 Dec 2021NLP4DH (ICON) 2021 12arXiv:2112.12489archive 2025-07-28

Quan Duong, Mika Hämäläinen, Khalid Alnajjar

Measuring the semantic similarity of different texts has many important applications in Digital Humanities research such as information retrieval, document clustering and text summarization. The performance of different methods depends on the length of the text, the domain and the language. This study focuses on experimenting with some of the current approaches to Finnish, which is a morphologically rich language. At the same time, we propose a simple method, TFW2V, which shows high efficiency in handling both long text documents and limited amounts of data. Furthermore, we design an objective evaluation method which can be used as a framework for benchmarking text similarity approaches.

PaperPDFConference PDFCode

Code

ruathudo/tfw2v officialmentioned in papertf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

BenchmarkingClusteringInformation RetrievalRetrievalSemantic SimilaritySemantic Textual SimilarityText Summarizationtext similarity

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections