Browse State-of-the-Art › Multi-task Language Understanding

Multi-task Language Understanding

44 papers with code · 5 benchmarks · 5 datasets archive 2025-07-28

Methodology

The test covers 57 tasks including elementary mathematics, US history, computer science, law, and more. https://arxiv.org/pdf/2009.03300.pdf

Description from the archive archive 2025-07-28.

Benchmarks archive 2025-07-28

5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
MML (44 rows) GPT-4 o1(300b) GPT-4o as the Gold Standard: A Scalable and General Purpose... — — Compare
BBH-nlp (15 rows) Qwen2.5-72B — — — Compare
MGSM (12 rows) PaLM 2 (few-shot, k=8, SC) PaLM 2 Technical Report code — Compare
BBH-alg (7 rows) code-davinci-002 175B (CoT) Evaluating Large Language Models Trained on Code code Syntology ran 6 of 39 samples · 33 unverified Compare
MMLU (5-Shot) (1 row) Sakalti/ultiima-78B MERGE: Fast Private Text Generation code — Compare

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

5 datasets whose archive record lists this task, ordered by the archive's paper count.

Subtasks archive 2025-07-28

No subtask under this task in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 44 papers with code (57 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 28 May 2020 67 repositories listed Syntology ran 15 of 65 samples · 50 unverified · 4 pointer-only (licence)
    By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do.
  • 26 Jul 2019 67 repositories listed Syntology ran 22 of 48 samples · 26 unverified · 23 pointer-only (licence)
    Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging.
  • 27 Feb 2023 57 repositories listed Syntology ran 26 of 58 samples · 32 unverified · 4 pointer-only (licence)
    We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters.
  • 26 Sep 2019 48 repositories listed Syntology ran 46 of 126 samples · 80 unverified · 22 pointer-only (licence)
    Increasing model size when pretraining natural language representations often results in improved performance on downstream tasks.
  • 14 Feb 2019 21 repositories listed
    Natural language processing tasks, such as question answering, machine translation, reading comprehension, and summarization, are typically approached with supervised learning on taskspecific datasets.
  • 18 Jul 2023 19 repositories listed Syntology ran 31 of 52 samples · 21 unverified · 16 pointer-only (licence)
    In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters.
  • 7 Sep 2020 18 repositories listed Syntology ran 5 of 26 samples · 21 unverified · 1 pointer-only (licence)
    By comprehensively evaluating the breadth and depth of a model's academic and professional understanding, our test can be used to analyze models across many tasks and to identify important shortcomings.
  • 7 Jul 2021 13 repositories listed Syntology ran 6 of 39 samples · 33 unverified · 2 pointer-only (licence)
    We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities.
  • 15 Mar 2023 11 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)
    We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.
  • 14 Apr 2022 11 repositories listed Syntology ran 0 of 3 samples · 3 unverified
    We introduce GPT-NeoX-20B, a 20 billion parameter autoregressive language model trained on the Pile, whose weights will be made freely and openly available to the public through a permissive license.
  • 20 Oct 2022 9 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 2 pointer-only (licence)
    We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks…
  • 5 Oct 2022 9 repositories listed Syntology ran 5 of 21 samples · 16 unverified
    We introduce GLM-130B, a bilingual (English and Chinese) pre-trained language model with 130 billion parameters.
  • 5 Apr 2022 7 repositories listed Syntology ran 30 of 37 samples · 7 unverified
    To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
  • 8 Jan 2024 6 repositories listed Syntology ran 5 of 5 samples · 0 unverified
    In particular, Mixtral vastly outperforms Llama 2 70B on mathematics, code generation, and multilingual benchmarks.
  • 10 Oct 2023 6 repositories listed Syntology ran 9 of 11 samples · 2 unverified · 1 pointer-only (licence)
    We introduce Mistral 7B v0.
  • 31 Jul 2024 5 repositories listed Syntology ran 2 of 9 samples · 7 unverified
    This paper presents a new set of foundation models, called Llama 3.
  • 22 Jan 2025 4 repositories listed
    We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1.
  • 30 Jan 2023 3 repositories listed Syntology ran 0 of 13 samples · 13 unverified · 13 pointer-only (licence)
    We introduce REPLUG, a retrieval-augmented language modeling framework that treats the language model (LM) as a black box and augments it with a tuneable retrieval model.
  • 8 Dec 2021 3 repositories listed
    Language modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
  • 3 Jun 2024 2 repositories listed Syntology ran 9 of 12 samples · 3 unverified
    In the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning…
  • 5 Jan 2024 2 repositories listed
    Using PESC during instruction tuning, our best sparse model outperforms other sparse and dense models and exhibits superior general capabilities compared to GPT-3.
  • 30 Mar 2023 2 repositories listed
    The use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering.
  • 5 Aug 2022 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)
    Retrieval augmented models are known to excel at knowledge intensive tasks without the need for as many parameters, but it is unclear whether they work in few-shot settings.
  • 10 May 2022 2 repositories listed Syntology ran 0 of 16 samples · 16 unverified
    Our model also achieve strong results at in-context learning, outperforming 175B GPT-3 on zero-shot SuperGLUE and tripling the performance of T5-XXL on one-shot summarization.
  • 29 Mar 2022 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 4 pointer-only (licence)
    We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget.
  • 18 Nov 2021 2 repositories listed
    Computing a simple average of the models' parameters therefore corresponds to making an isotropic Gaussian approximation to their posteriors.
  • 2 May 2020 2 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 1 pointer-only (licence)
    As evidence, we use the latest advances in language modeling to build a single pre-trained QA model, UnifiedQA, that performs surprisingly well across 17 QA datasets spanning 4 diverse formats.
  • 16 Feb 2025 1 repository listed
    We propose two benchmarks for Turkic language MMLU: TUMLU is a comprehensive, multilingual, and natively developed language understanding benchmark specifically designed for Turkic languages.
  • 19 Dec 2024 1 repository listed
    This benchmark reassesses LLMs' understanding of world knowledge by averting both unintentional and malicious data leakage.
  • 13 Dec 2024 1 repository listed
    Scaling large language models (LLMs) significantly improves performance but comes with prohibitive computational costs.

Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections