{"url":"/task/transliteration","name":"Transliteration","slug":"transliteration","description_markdown":"**Transliteration** is a mechanism for converting a word in a source (foreign) language to a target language, and often adopts approaches from machine translation. In machine translation, the objective is to preserve the semantic meaning of the utterance as much as possible while following the syntactic structure in the target language. In Transliteration, the objective is to preserve the original pronunciation of the source word as much as possible while following the phonological structures of the target language.\n\nFor example, the city’s name “Manchester” has become well known by people of languages other than English. These new words are often named entities that are important in cross-lingual information retrieval, information extraction, machine translation, and often present out-of-vocabulary challenges to spoken language technologies such as automatic speech recognition, spoken keyword search, and text-to-speech.\n\n\n<span class=\"description-source\">Source: [Phonology-Augmented Statistical Framework for Machine Transliteration using Limited Linguistic Resources ](https://arxiv.org/abs/1810.03184)</span>","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":435,"papers_with_code":54,"benchmarks":0,"benchmark_tables_in_archive":0,"benchmark_tables_shown":0,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":5,"subtasks":0,"parent_tasks":1},"benchmarks":[],"datasets":[{"url":"/dataset/dakshina","name":"Dakshina","full_name":"Dakshina","num_papers_in_archive":14},{"url":"/dataset/tsac","name":"TSAC","full_name":"Tunisian Sentiment Analysis Corpus","num_papers_in_archive":9},{"url":"/dataset/arzen","name":"ArzEn","full_name":"Corpus of Egyptian Arabic-English Code-switching","num_papers_in_archive":3},{"url":"/dataset/anetac","name":"ANETAC","full_name":"Arabic Named Entity Transliteration and Classification","num_papers_in_archive":2},{"url":"/dataset/bianet","name":"Bianet","full_name":"","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/machine-translation","name":"Machine Translation"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":54,"tagged_in_all":435,"items":[{"url":"/paper/applying-the-transformer-to-character-level","title":"Applying the Transformer to Character-level Transduction","date":"2020-05-20","arxiv_id":"2005.10213","repositories_listed":3,"syntology":{"n":5,"n_ran":2,"n_unverified":3,"n_pointer_only":2}},{"url":"/paper/aksharantar-towards-building-open","title":"Aksharantar: Open Indic-language Transliteration datasets and models for the Next Billion Users","date":"2022-05-06","arxiv_id":"2205.03018","repositories_listed":2,"syntology":null},{"url":"/paper/shuowen-jiezi-linguistically-informed","title":"Sub-Character Tokenization for Chinese Pretrained Language Models","date":"2021-06-01","arxiv_id":"2106.00400","repositories_listed":2,"syntology":{"n":4,"n_ran":3,"n_unverified":1,"n_pointer_only":4}},{"url":"/paper/processing-south-asian-languages-written-in-1","title":"Processing South Asian Languages Written in the Latin Script: the Dakshina Dataset","date":"2020-07-02","arxiv_id":"2007.01176","repositories_listed":2,"syntology":null},{"url":"/paper/universal-dependency-parsing-for-hindi","title":"Universal Dependency Parsing for Hindi-English Code-switching","date":"2018-04-16","arxiv_id":"1804.05868","repositories_listed":2,"syntology":null},{"url":"/paper/parsipy-nlp-toolkit-for-historical-persian","title":"ParsiPy: NLP Toolkit for Historical Persian Texts in Python","date":"2025-03-22","arxiv_id":"2503.17810","repositories_listed":1,"syntology":null},{"url":"/paper/sinhala-transliteration-a-comparative","title":"Sinhala Transliteration: A Comparative Analysis Between Rule-based and Seq2Seq Approaches","date":"2024-12-31","arxiv_id":"2501.00529","repositories_listed":1,"syntology":null},{"url":"/paper/romanized-to-native-malayalam-script","title":"Romanized to Native Malayalam Script Transliteration Using an Encoder-Decoder Framework","date":"2024-12-13","arxiv_id":"2412.09957","repositories_listed":1,"syntology":null},{"url":"/paper/gr-nlp-toolkit-an-open-source-nlp-toolkit-for","title":"GR-NLP-TOOLKIT: An Open-Source NLP Toolkit for Modern Greek","date":"2024-12-11","arxiv_id":"2412.08520","repositories_listed":1,"syntology":null},{"url":"/paper/how-transliterations-improve-crosslingual","title":"How Transliterations Improve Crosslingual Alignment","date":"2024-09-25","arxiv_id":"2409.17326","repositories_listed":1,"syntology":null},{"url":"/paper/can-small-language-models-learn-unlearn-and","title":"Can Small Language Models Learn, Unlearn, and Retain Noise Patterns?","date":"2024-07-01","arxiv_id":"2407.00996","repositories_listed":1,"syntology":null},{"url":"/paper/breaking-the-script-barrier-in-multilingual","title":"Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment","date":"2024-06-28","arxiv_id":"2406.19759","repositories_listed":1,"syntology":null},{"url":"/paper/jailbreaking-llms-with-arabic-transliteration","title":"Jailbreaking LLMs with Arabic Transliteration and Arabizi","date":"2024-06-26","arxiv_id":"2406.18725","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/transmi-a-framework-to-create-strong","title":"TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data","date":"2024-05-16","arxiv_id":"2405.09913","repositories_listed":1,"syntology":null},{"url":"/paper/paranames-1-0-creating-an-entity-name-corpus","title":"ParaNames 1.0: Creating an Entity Name Corpus for 400+ Languages using Wikidata","date":"2024-05-15","arxiv_id":"2405.09496","repositories_listed":1,"syntology":null},{"url":"/paper/translico-a-contrastive-learning-framework-to","title":"TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models","date":"2024-01-12","arxiv_id":"2401.06620","repositories_listed":1,"syntology":null},{"url":"/paper/towards-scene-text-to-scene-text-translation","title":"Show Me the World in My Language: Establishing the First Baseline for Scene-Text to Scene-Text Translation","date":"2023-08-06","arxiv_id":"2308.03024","repositories_listed":1,"syntology":null},{"url":"/paper/taqyim-evaluating-arabic-nlp-tasks-using","title":"Taqyim: Evaluating Arabic NLP Tasks Using ChatGPT Models","date":"2023-06-28","arxiv_id":"2306.16322","repositories_listed":1,"syntology":null},{"url":"/paper/deepscribe-localization-and-classification-of","title":"DeepScribe: Localization and Classification of Elamite Cuneiform Signs Via Deep Learning","date":"2023-06-02","arxiv_id":"2306.01268","repositories_listed":1,"syntology":null},{"url":"/paper/multilingual-text-to-speech-synthesis-for","title":"Multilingual Text-to-Speech Synthesis for Turkic Languages Using Transliteration","date":"2023-05-25","arxiv_id":"2305.15749","repositories_listed":1,"syntology":null},{"url":"/paper/xtreme-up-a-user-centric-scarce-data","title":"XTREME-UP: A User-Centric Scarce-Data Benchmark for Under-Represented Languages","date":"2023-05-19","arxiv_id":"2305.11938","repositories_listed":1,"syntology":{"n":11,"n_ran":0,"n_unverified":11,"n_pointer_only":0}},{"url":"/paper/beyond-arabic-software-for-perso-arabic","title":"Beyond Arabic: Software for Perso-Arabic Script Manipulation","date":"2023-01-26","arxiv_id":"2301.11406","repositories_listed":1,"syntology":null},{"url":"/paper/question-answering-classification-for-amharic","title":"Question Answering Classification for Amharic Social Media Community Based Questions","date":"2022-06-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/a-machine-transliteration-tool-between-uzbek","title":"A machine transliteration tool between Uzbek alphabets","date":"2022-05-19","arxiv_id":"2205.09578","repositories_listed":1,"syntology":null},{"url":"/paper/paranames-a-massively-multilingual-entity","title":"ParaNames: A Massively Multilingual Entity Name Corpus","date":"2022-02-28","arxiv_id":"2202.14035","repositories_listed":1,"syntology":null},{"url":"/paper/does-transliteration-help-multilingual","title":"Does Transliteration Help Multilingual Language Modeling?","date":"2022-01-29","arxiv_id":"2201.12501","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/iiitt-dravidian-codemix-fire2021","title":"IIITT@Dravidian-CodeMix-FIRE2021: Transliterate or translate? Sentiment analysis of code-mixed text in Dravidian languages","date":"2021-11-15","arxiv_id":"2111.07906","repositories_listed":1,"syntology":null},{"url":"/paper/orthographic-transliteration-for-kabyle","title":"Orthographic Transliteration for Kabyle Speech Recognition","date":"2021-11-01","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/role-of-language-relatedness-in-multilingual","title":"Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages","date":"2021-09-22","arxiv_id":"2109.10534","repositories_listed":1,"syntology":null},{"url":"/paper/cross-lingual-text-classification-of","title":"Cross-Lingual Text Classification of Transliterated Hindi and Malayalam","date":"2021-08-31","arxiv_id":"2108.13620","repositories_listed":1,"syntology":null}],"syntology_records":5,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}