Browse State-of-the-Art › Word Alignment
Word Alignment
92 papers with code · 7 benchmarks · 4 datasets archive 2025-07-28
Word Alignment is the task of finding the correspondence between source and target words in a pair of sentences that are translations of each other.
Source: Neural Network-based Word Alignment through Score Aggregation
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| en-es (2 rows) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
| en-fr (2 rows) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
| es-en (2 rows) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
| fr-en (2 rows) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
| en-it (1 row) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
| MUSE en-de (1 row) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
| MUSE en-pt (1 row) | Barycenter Alignment | Unsupervised Multilingual Alignment using Wasserstein Barycenter | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 92 papers with code (551 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
11 Oct 2017 20 repositories listed Syntology ran 6 of 8 samples · 2 unverified · 8 pointer-only (licence)We finally describe experiments on the English-Esperanto low-resource language pair, on which there only exists a limited amount of parallel data, to show the potential impact of our method in fully unsupervised machine…
-
30 Oct 2014 5 repositories listedOur experiments on the WMT14 English to French translation task show that this method provides a substantial improvement of up to 2.
-
30 Sep 2021 4 repositories listed Syntology ran 6 of 12 samples · 6 unverified · 1 pointer-only (licence)Non-autoregressive text-to-speech (NAR-TTS) models such as FastSpeech 2 and Glow-TTS can synthesize high-quality speech from the given text in parallel.
-
14 Apr 2021 3 repositories listed Syntology ran 3 of 5 samples · 2 unverifiedRecently, it has been argued that encoder-decoder models can be made more interpretable by replacing the softmax function in the attention with its sparse variants.
-
20 Jan 2021 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn addition, we demonstrate that we are able to train multilingual word aligners that can obtain robust performance on different language pairs.
-
29 Apr 2020 3 repositories listedWe introduce MultiATIS++, a new multilingual NLU corpus that extends the Multilingual ATIS corpus to nine languages across four language families, and evaluate our method using the corpus.
-
18 Apr 2020 3 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedWe find that alignments created from embeddings are superior for four and comparable for two language pairs compared to those produced by traditional statistical aligners, even with abundant parallel data; e.
-
4 Dec 2023 2 repositories listed Syntology ran 2 of 14 samples · 12 unverifiedHowever, predominant paradigms, driven by casting instance-level tasks as an object-word alignment, bring heavy cross-modality interaction, which is not effective in prompting object detection and visual grounding.
-
9 Jun 2023 2 repositories listedMost existing word alignment methods rely on manual alignment datasets or parallel corpora, which limits their usefulness.
-
19 May 2023 2 repositories listedAttention is the core mechanism of today's most used architectures for natural language processing and has been analyzed from many perspectives, including its effectiveness for machine translation-related tasks.
-
28 Nov 2022 2 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedTranslating training data into many languages has emerged as a practical solution for improving cross-lingual transfer.
-
15 Mar 2022 2 repositories listedRecent dominant methods for video-language pre-training (VLP) learn transferable representations from the raw pixels in an end-to-end manner to achieve advanced performance on downstream video-language retrieval.
-
30 Jul 2021 2 repositories listedThe quantitative evaluation demonstrates that our backbone translation models achieve state-of-the-art translation performance and our quality estimation well correlates with both BLEU and human judgment.
-
15 May 2025 1 repository listedIn this work, we revisit WAs for label projection, systematically investigating the effects of low-level design decisions on token-level XLT: (i) the algorithm for projecting labels between (multi-)token spans, (ii)…
-
10 Mar 2025 1 repository listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)As large language models (LLMs) increasingly become central to various applications and interact with diverse user populations, ensuring their reliable and consistent performance is becoming more important.
-
16 Jul 2024 1 repository listedReal world deployments of word alignment are almost certain to cover both high and low resource languages.
-
5 Apr 2024 1 repository listedThis paper addresses text-supervised semantic segmentation, aiming to learn a model capable of segmenting arbitrary visual concepts within images by using only image-text pairs without dense annotations.
-
25 Mar 2024 1 repository listed Syntology ran 6 of 11 samples · 5 unverified · 11 pointer-only (licence)In our work, we propose a new task formulation of dense retrieval, cross-lingual contextualized phrase retrieval, which aims to augment cross-lingual applications by addressing polysemy using context information.
-
21 Feb 2024 1 repository listedWe introduce a Translated dataset for Multilingual Coreference Resolution (TransMuCoRes) in 31 South Asian languages using off-the-shelf tools for translation and word-alignment.
-
15 Feb 2024 1 repository listed Syntology ran 7 of 10 samples · 3 unverifiedRecent work has shown that, while large language models (LLMs) demonstrate strong word translation or bilingual lexicon induction (BLI) capabilities in few-shot setups, they still cannot match the performance of…
-
5 Feb 2024 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedTherefore, it is common to exploit translation and label projection to further improve the performance by (1) translating training data that is available in a high-resource language (e.
-
Salute the Classic: Revisiting Challenges of Machine Translation in the Age of Large Language Models16 Jan 2024 1 repository listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)This study revisits these challenges, offering insights into their ongoing relevance in the context of advanced Large Language Models (LLMs): domain mismatch, amount of parallel data, rare word prediction, translation…
-
1 Jan 2024 1 repository listedWe propose a novel Linguistic-Aware Patch Slimming (LAPS) framework for fine-grained alignment which explicitly identifies redundant visual patches with language supervision and rectifies their semantic and spatial…
-
25 Oct 2023 1 repository listed Syntology ran 6 of 8 samples · 2 unverified · 8 pointer-only (licence)CoDet then leverages visual similarities to discover the co-occurring objects and align them with the shared concept.
-
21 Oct 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedBilingual Lexicon Induction (BLI) is a core task in multilingual NLP that still, to a large extent, relies on calculating cross-lingual word representations.
-
24 Aug 2023 1 repository listed Syntology ran 8 of 9 samples · 1 unverified · 9 pointer-only (licence)The experimental results demonstrate significant improvements in translation performance with SWIE based on BLOOMZ-3b, particularly in zero-shot and long text translations due to reduced instruction forgetting risk.
-
10 Jun 2023 1 repository listedIn contrast, Arabic DL-based multimodal sentiment analysis (MSA) is still in its infantile stage due, mainly, to the lack of standard datasets.
-
7 Jun 2023 1 repository listed Syntology ran 0 of 14 samples · 14 unverifiedMonolingual word alignment is crucial to model semantic interactions between sentences.
-
26 May 2023 1 repository listedOn the task of Machine Translation (MT), multiple works have investigated few-shot prompting mechanisms to elicit better translations from LLMs.
-
22 May 2023 1 repository listedAutomatically highlighting words that cause semantic differences between two documents could be useful for a wide range of applications.
Syntology lines on 16 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections