Browse State-of-the-Art › Multilingual NLP
Multilingual NLP
42 papers with code · 0 benchmarks · 8 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 42 papers with code (96 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Nov 2019 35 repositories listed Syntology ran 25 of 59 samples · 34 unverified · 52 pointer-only (licence)We also present a detailed empirical analysis of the key factors that are required to achieve these gains, including the trade-offs between (1) positive transfer and capacity dilution and (2) the performance of high and…
-
9 Nov 2022 7 repositories listed Syntology ran 2 of 11 samples · 9 unverifiedLarge language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions.
-
3 Jul 2020 6 repositories listedWhile BERT is an effective method for learning monolingual sentence embeddings for semantic similarity and embedding based transfer learning (Reimers and Gurevych, 2019), BERT based cross-lingual sentence embeddings…
-
2 May 2024 3 repositories listedThis paper introduces UQA, a novel dataset for question answering and text comprehension in Urdu, a low-resource language with over 70 million native speakers.
-
6 Feb 2024 2 repositories listedWe recommend future work to include an operationalization of 'typological diversity' that empirically justifies the diversity of language samples.
-
22 May 2023 2 repositories listedColexNet's nodes are concepts and its edges are colexifications.
-
6 May 2021 2 repositories listedThe introduction of pretrained cross-lingual language models brought decisive improvements to multilingual NLP tasks.
-
27 Jan 2020 2 repositories listedParallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications.
-
3 Oct 2017 2 repositories listedMultilinguality is gradually becoming ubiquitous in the sense that more and more researchers have successfully shown that using additional languages help improve the results in many Natural Language Processing tasks.
-
18 Jun 2024 1 repository listedRapidly growing numbers of multilingual news consumers pose an increasing challenge to news recommender systems in terms of providing customized recommendations.
-
13 Jun 2024 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedPerformance prediction is a method to estimate the performance of Language Models (LMs) on various Natural Language Processing (NLP) tasks, mitigating computational costs associated with model capacity and data for…
-
7 May 2024 1 repository listedHowever, Fairlearn's metrics suggest that the SVM approaches equitable levels with a demographic parity ratio of 0.
-
29 Apr 2024 1 repository listedThis study investigates the factors influencing the performance of multilingual large language models (MLLMs) across diverse languages.
-
6 Mar 2024 1 repository listedTypologically diverse benchmarks are increasingly created to track the progress achieved in multilingual NLP.
-
15 Feb 2024 1 repository listed Syntology ran 7 of 10 samples · 3 unverifiedRecent work has shown that, while large language models (LLMs) demonstrate strong word translation or bilingual lexicon induction (BLI) capabilities in few-shot setups, they still cannot match the performance of…
-
21 Oct 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedBilingual Lexicon Induction (BLI) is a core task in multilingual NLP that still, to a large extent, relies on calculating cross-lingual word representations.
-
13 Oct 2023 1 repository listedThis study presents a large multi-modal Bangla YouTube clickbait dataset consisting of 253, 070 data points collected through an automated process using the YouTube API and Python web automation frameworks.
-
12 Oct 2023 1 repository listedBy comparing the performance of CapsNet to that of other architectures, such as DistilBERT, Vanilla Neural Networks (VNN), and Convolutional Neural Networks (CNN), we were able to achieve an accuracy of 90.
-
16 Sep 2023 1 repository listedFoundational large language models (LLMs) can be instruction-tuned to perform open-domain question answering, facilitating applications like chat assistants.
-
19 May 2023 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedWe evaluate commonly used models on the benchmark.
-
15 May 2023 1 repository listedThis paper introduces PMIndiaSum, a multilingual and massively parallel summarization corpus focused on languages in India.
-
11 May 2023 1 repository listedWhile a large body of work leveraged MMTs to mine parallel data and induce bilingual document embeddings, much less effort has been devoted to training general-purpose (massively) multilingual document encoder that can…
-
3 Apr 2023 1 repository listedMultilingual pre-training on monolingual data ignores the availability of parallel data in many language pairs.
-
30 Oct 2022 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedThis crucial step is done via 1) creating a word similarity dataset, comprising positive word pairs (i.
-
10 Oct 2022 1 repository listedTimely and effective response to humanitarian crises requires quick and accurate analysis of large amounts of text data - a process that can highly benefit from expert-assisted NLP systems trained on validated and…
-
24 Jun 2022 1 repository listedOur model sets the new state of the art performance of 67.
-
1 Jun 2022 1 repository listedWe present the TeDDi sample, a diversity sample of text data for language comparison and multilingual Natural Language Processing.
-
15 Mar 2022 1 repository listed Syntology ran 5 of 6 samples · 1 unverified · 4 pointer-only (licence)At Stage C1, we propose to refine standard cross-lingual linear maps between static word embeddings (WEs) via a contrastive learning objective; we also show how to integrate it into the self-learning procedure for even…
-
16 Nov 2021 1 repository listedAs Stage C1, we propose to refine standard cross-lingual linear maps between static word embeddings (WEs) via a contrastive learning objective; we also show how to integrate it into the self-learning procedure for even…
-
1 Nov 2021 1 repository listedMultilingual Named Entity Recognition (NER) is a key intermediate task which is needed in many areas of NLP.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections