Browse State-of-the-Art › Multiple-choice
Multiple-choice
483 papers with code · 0 benchmarks · 12 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 483 papers with code (1,107 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
3 May 2015 21 repositories listed Syntology ran 6 of 7 samples · 1 unverified · 6 pointer-only (licence)Given an image and a natural language question about the image, the task is to provide an accurate natural language answer.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
16 Nov 2023 6 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 1 pointer-only (licence)In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.
-
29 Dec 2022 5 repositories listedNearly all jurisdictions in the United States require a professional license exam, commonly referred to as "the Bar Exam," as a precondition for law practice.
-
29 Apr 2022 5 repositories listed Syntology ran 18 of 24 samples · 6 unverified · 7 pointer-only (licence)Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research.
-
5 Jun 2024 4 repositories listedIn this study, we introduce a novel approach to generate quizzes from Turkish educational texts, marking a pioneering endeavor in educational technology specifically tailored to the Turkish educational context.
-
22 Feb 2024 4 repositories listed Syntology ran 10 of 39 samples · 29 unverified · 17 pointer-only (licence)The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities.
-
9 Dec 2023 4 repositories listed Syntology ran 14 of 22 samples · 8 unverified · 9 pointer-only (licence)We introduce Contrastive Activation Addition (CAA), an innovative method for steering language models by modifying their activations during forward passes.
-
3 Nov 2023 4 repositories listedOne of the major limitations of the basic score is that it treats all words as equally important.
-
27 Nov 2018 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedWhile this task is easy for humans, it is tremendously difficult for today's vision systems, requiring higher-order cognition and commonsense reasoning about the world.
-
2 Nov 2018 4 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedTo investigate question answering with prior knowledge, we present CommonsenseQA: a challenging new dataset for commonsense question answering.
-
31 Dec 2024 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)To bridge this gap, we introduce MapEval, a benchmark designed to assess diverse and complex map-based user queries with geo-spatial reasoning.
-
11 Jun 2024 3 repositories listed Syntology ran 7 of 17 samples · 10 unverifiedIn this paper, we present the VideoLLaMA 2, a set of Video Large Language Models (Video-LLMs) designed to enhance spatial-temporal modeling and audio understanding in video and audio-oriented tasks.
-
28 Nov 2023 3 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedWith the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models.
-
20 Nov 2023 3 repositories listed Syntology ran 7 of 8 samples · 1 unverifiedWe present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry.
-
30 Jul 2023 3 repositories listedBased on powerful Large Language Models (LLMs), recent generative Multimodal Large Language Models (MLLMs) have gained prominence as a pivotal research area, exhibiting remarkable capability for both comprehension and…
-
12 Jul 2023 3 repositories listedIn response to these challenges, we propose MMBench, a bilingual benchmark for assessing the multi-modal capabilities of VLMs.
-
27 May 2023 3 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedFine-tuning language models (LMs) has yielded success on diverse downstream tasks, but as LMs grow in size, backpropagation requires a prohibitively large amount of memory.
-
16 Dec 2021 3 repositories listedTo enable building and testing models on long-document comprehension, we introduce QuALITY, a multiple-choice QA dataset with context passages in English that have an average length of about 5, 000 tokens, much longer…
-
28 Sep 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedOpen domain question answering (OpenQA) tasks have been recently attracting more and more attention from the natural language processing (NLP) community.
-
28 Aug 2018 3 repositories listedIn this paper we propose a retriever-reader model that learns to attend on essential terms during the question answering process.
-
27 Jun 2016 3 repositories listedVisual question answering (VQA) is an interesting learning setting for evaluating the abilities and shortcomings of current systems for image understanding.
-
20 May 2025 2 repositories listed Syntology ran 5 of 14 samples · 9 unverifiedLarge multimodal models (LMMs) have recently emerged as a powerful tool for long video understanding (LVU), prompting the development of standardized LVU benchmarks to evaluate their performance.
-
28 Feb 2025 2 repositories listed Syntology ran 0 of 9 samples · 9 unverifiedLarge Language Models (LLMs) and LLM-based agents show great promise in accelerating scientific research.
-
17 Nov 2024 2 repositories listedThe advancement of Multimodal Large Language Models (MLLMs) has enabled significant progress in multimodal understanding, expanding their capacity to analyze video content.
-
6 Aug 2024 2 repositories listedWe present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series.
-
3 Aug 2024 2 repositories listed Syntology ran 9 of 14 samples · 5 unverifiedThe recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone.
-
2 Aug 2024 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedMotivated by this, we introduce MuChoMusic, a benchmark for evaluating music understanding in multimodal language models focused on audio.
-
15 Jul 2024 2 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)We demonstrate that an attacker can embed a backdoor in LLMs, which, when activated by a specific trigger in the input, manipulates the model's uncertainty without affecting the final output.
-
14 Jun 2024 2 repositories listedEnhancing Language Models' (LMs) ability to understand purchase intentions in E-commerce scenarios is crucial for their effective assistance in various downstream tasks.
Syntology lines on 19 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections