Browse State-of-the-Art › Transliteration
Transliteration
54 papers with code · 0 benchmarks · 5 datasets archive 2025-07-28
Transliteration is a mechanism for converting a word in a source (foreign) language to a target language, and often adopts approaches from machine translation. In machine translation, the objective is to preserve the semantic meaning of the utterance as much as possible while following the syntactic structure in the target language. In Transliteration, the objective is to preserve the original pronunciation of the source word as much as possible while following the phonological structures of the target language.
For example, the city’s name “Manchester” has become well known by people of languages other than English. These new words are often named entities that are important in cross-lingual information retrieval, information extraction, machine translation, and often present out-of-vocabulary challenges to spoken language technologies such as automatic speech recognition, spoken keyword search, and text-to-speech.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 54 papers with code (435 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 May 2020 3 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 2 pointer-only (licence)The transformer has been shown to outperform recurrent neural network-based sequence-to-sequence models in various word-level NLP tasks.
-
6 May 2022 2 repositories listedTransliteration is very important in the Indian language context due to the usage of multiple scripts and the widespread use of romanized inputs.
-
1 Jun 2021 2 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)2) Pronunciation-based SubChar tokenizers can encode Chinese homophones into the same transliteration sequences and produce the same tokenization output, hence being robust to homophone typos.
-
2 Jul 2020 2 repositories listedThis paper describes the Dakshina dataset, a new resource consisting of text in both the Latin and native scripts for 12 South Asian languages.
-
16 Apr 2018 2 repositories listedWe present a treebank of Hindi-English code-switching tweets under Universal Dependencies scheme and propose a neural stacking model for parsing that efficiently leverages part-of-speech tag and syntactic tree…
-
22 Mar 2025 1 repository listedThe study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages.
-
31 Dec 2024 1 repository listedWe propose two methods to address this problem: Our baseline is a rule-based method, which is then compared against our second method where we approach the transliteration problem as a sequence-to-sequence task akin to…
-
13 Dec 2024 1 repository listedIn this work, we present the development of a reverse transliteration model to convert romanized Malayalam to native script using an encoder-decoder framework built with attention-based bidirectional Long Short Term…
-
11 Dec 2024 1 repository listedWe present GR-NLP-TOOLKIT, an open-source natural language processing (NLP) toolkit developed specifically for modern Greek.
-
25 Sep 2024 1 repository listedHowever, we also show that better alignment does not always yield better downstream performance, suggesting that further research is needed to clarify the connection between alignment and performance.
-
1 Jul 2024 1 repository listedSmall Language Models (SLMs) are generally considered more compact versions of large language models (LLMs).
-
28 Jun 2024 1 repository listedHowever, the transfer performance is often hindered when a low-resource target language is written in a different script than the high-resource source language, even though the two languages may be related or share…
-
26 Jun 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThis study identifies the potential vulnerabilities of Large Language Models (LLMs) to 'jailbreak' attacks, specifically focusing on the Arabic language and its various forms.
-
16 May 2024 1 repository listedTransMI can create strong baselines for data that is transliterated into a common script by exploiting an existing mPLM and its tokenizer without any training.
-
15 May 2024 1 repository listedNames are provided for 16.
-
12 Jan 2024 1 repository listedAs a consequence, mPLMs are faced with a script barrier: representations from different scripts are located in different subspaces, which can result in crosslingual transfer involving languages of different scripts…
-
6 Aug 2023 1 repository listedIn this work, we study the task of ``visually'' translating scene text from a source language (e.
-
28 Jun 2023 1 repository listedLarge language models (LLMs) have demonstrated impressive performance on various downstream tasks without requiring fine-tuning, including ChatGPT, a chat-based model built on top of LLMs such as GPT-3.
-
2 Jun 2023 1 repository listedThe end-to-end pipeline achieves a top-5 classification accuracy of 0.
-
25 May 2023 1 repository listedThis work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Turkmen, Uyghur, and Uzbek.
-
19 May 2023 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedWe evaluate commonly used models on the benchmark.
-
26 Jan 2023 1 repository listedThis paper presents an open-source software library that provides a set of finite-state transducer (FST) components and corresponding utilities for manipulating the writing systems of languages that use the Perso-Arabic…
-
1 Jun 2022 1 repository listedQuestions are posted in Amharic, English, or Amharic but in a Latin script.
-
19 May 2022 1 repository listedMachine transliteration, as defined in this paper, is a process of automatically transforming written script of words from a source alphabet into words of another target alphabet within the same language, while…
-
28 Feb 2022 1 repository listedWe demonstrate an application of ParaNames by training a multilingual model for canonical name translation to and from English.
-
29 Jan 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedWe empirically measure the effect of transliteration on MLLMs in this context.
-
15 Nov 2021 1 repository listedThis research paper bestows a tiny contribution to this research in the form of sentiment analysis of code-mixed social media comments in the popular Dravidian languages Kannada, Tamil and Malayalam.
-
1 Nov 2021 1 repository listed
-
22 Sep 2021 1 repository listedWe hypothesize and validate that multilingual fine-tuning of pre-trained language models can yield better performance on downstream NLP applications, compared to models fine-tuned on individual languages.
-
31 Aug 2021 1 repository listedTransliteration is very common on social media, but transliterated text is not adequately handled by modern neural models for various NLP tasks.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections