Papers › KaWAT: A Word Analogy Task Dataset for Indonesian

KaWAT: A Word Analogy Task Dataset for Indonesian

17 Jun 2019arXiv:1906.09912archive 2025-07-28

Kemal Kurniawan

We introduced KaWAT (Kata Word Analogy Task), a new word analogy task dataset for Indonesian. We evaluated on it several existing pretrained Indonesian word embeddings and embeddings trained on Indonesian online news corpus. We also tested them on two downstream tasks and found that pretrained word embeddings helped either by reducing the training epochs or yielding significant performance gains.

PaperPDFCode

Code

kata-ai/id-word2vec officialmentioned in papermentioned on GitHub report
kata-ai/kawat officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Word Embeddings

Datasets

Introduced by this paper, per the archive.

KaWAT

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections