Papers › "A Passage to India": Pre-trained Word Embeddings for Indian Languages

"A Passage to India": Pre-trained Word Embeddings for Indian Languages

27 Dec 2021arXiv:2112.13800archive 2025-07-28

Kumar Saurav, Kumar Saunack, Diptesh Kanojia, Pushpak Bhattacharyya

Dense word vectors or 'word embeddings' which encode semantic properties of words, have now become integral to NLP tasks like Machine Translation (MT), Question Answering (QA), Word Sense Disambiguation (WSD), and Information Retrieval (IR). In this paper, we use various existing approaches to create multiple word embeddings for 14 Indian languages. We place these embeddings for all these languages, viz., Assamese, Bengali, Gujarati, Hindi, Kannada, Konkani, Malayalam, Marathi, Nepali, Odiya, Punjabi, Sanskrit, Tamil, and Telugu in a single repository. Relatively newer approaches that emphasize catering to context (BERT, ELMo, etc.) have shown significant improvements, but require a large amount of resources to generate usable models. We release pre-trained embeddings generated using both contextual and non-contextual approaches. We also use MUSE and XLM to train cross-lingual embeddings for all pairs of the aforementioned languages. To show the efficacy of our embeddings, we evaluate our embedding models on XPOS, UPOS and NER tasks for all these languages. We release a total of 436 models using 8 different approaches. We hope they are useful for the resource-constrained Indian language NLP. The title of this paper refers to the famous novel 'A Passage to India' by E.M. Forster, published initially in 1924.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Information RetrievalMachine TranslationNERQuestion AnsweringRetrievalWord EmbeddingsWord Sense Disambiguation

Datasets

Introduced by this paper, per the archive.

Pre-trained Transliterated Embeddings for Indian Languages

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

AdamAttentionAttention DropoutBPEBiLSTMDense ConnectionsDropoutELMoLSTMLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSigmoid ActivationSoftmaxTanh ActivationXLM

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections