Browse State-of-the-Art › Multiple-choice › Papers, page 8
Multiple-choice
Papers archive 2025-07-28
archive papers tagged: 1,107 · with a code link: 483 · where Syntology ran a sample: 161 (124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (161 of 1,107 tagged: 124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument)
Page 8 of 12: papers 701 to 800 of 1,107, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
10 Oct 2024 0 repositories listed
-
Answering Questions in Stages: Prompt Chaining for Contract QA9 Oct 2024 0 repositories listed
-
ACPBench: Reasoning about Action, Change, and Planning8 Oct 2024 0 repositories listed
-
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition8 Oct 2024 0 repositories listed
-
Listening to the Wise Few: Select-and-Copy Attention Heads for Multiple-Choice QA3 Oct 2024 0 repositories listed
-
3 Oct 2024 0 repositories listed
-
Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model1 Oct 2024 0 repositories listed
-
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs30 Sep 2024 0 repositories listed
-
Task-Adaptive Pretrained Language Models via Clustered-Importance Sampling30 Sep 2024 0 repositories listed
-
Mitigating Selection Bias with Node Pruning and Auxiliary Options27 Sep 2024 0 repositories listed
-
DARE: Diverse Visual Question Answering with Robustness Evaluation26 Sep 2024 0 repositories listed
-
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ25 Sep 2024 0 repositories listed
-
RISCORE: Enhancing In-Context Riddle Solving in Language Models through Context-Reconstructed Example Augmentation24 Sep 2024 0 repositories listed
-
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation23 Sep 2024 0 repositories listed
-
Evaluating the Performance and Robustness of LLMs in Materials Science Q&A and Property Predictions22 Sep 2024 0 repositories listed
-
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge20 Sep 2024 0 repositories listed
-
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination19 Sep 2024 0 repositories listed
-
Edu-Values: Towards Evaluating the Chinese Education Values of Large Language Models19 Sep 2024 0 repositories listed
-
Efficient Knowledge Distillation: Empowering Small Language Models with Teacher Model Insights19 Sep 2024 0 repositories listed
-
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do17 Sep 2024 0 repositories listed
-
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia13 Sep 2024 0 repositories listed
-
Exploring syntactic information in sentence embeddings through multilingual subject-verb agreement10 Sep 2024 0 repositories listed
-
Towards Democratizing Multilingual Large Language Models For Medicine Through A Two-Stage Instruction Fine-tuning Approach9 Sep 2024 0 repositories listed
-
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models5 Sep 2024 0 repositories listed
-
The Role of Large Language Models in Musicology: Are We Ready to Trust the Machines?3 Sep 2024 0 repositories listed
-
Novel-WD: Exploring acquisition of Novel World Knowledge in LLMs Using Prefix-Tuning30 Aug 2024 0 repositories listed
-
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations22 Aug 2024 0 repositories listed
-
How Susceptible are LLMs to Influence in Prompts?17 Aug 2024 0 repositories listed
-
Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil9 Aug 2024 0 repositories listed
-
Winning Amazon KDD Cup'245 Aug 2024 0 repositories listed
-
Recent Advances in Multi-Choice Machine Reading Comprehension: A Survey on Methods and Datasets4 Aug 2024 0 repositories listed
-
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models23 Jul 2024 0 repositories listed
-
Improved Few-Shot Image Classification Through Multiple-Choice Questions23 Jul 2024 0 repositories listed
-
MIBench: Evaluating Multimodal Large Language Models over Multiple Images21 Jul 2024 0 repositories listed
-
Adversarial Databases Improve Success in Retrieval-based Large Language Models19 Jul 2024 0 repositories listed
-
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models16 Jul 2024 0 repositories listed
-
AstroMLab 1: Who Wins Astronomy Jeopardy!?15 Jul 2024 0 repositories listed
-
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models15 Jul 2024 0 repositories listed
-
LAB-Bench: Measuring Capabilities of Language Models for Biology Research14 Jul 2024 0 repositories listed
-
Evaluating Nuanced Bias in Large Language Model Free Response Answers11 Jul 2024 0 repositories listed
-
CFinBench: A Comprehensive Chinese Financial Benchmark for Large Language Models2 Jul 2024 0 repositories listed
-
Changing Answer Order Can Decrease MMLU Accuracy27 Jun 2024 0 repositories listed
-
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration24 Jun 2024 0 repositories listed
-
SynDARin: Synthesising Datasets for Automated Reasoning in Low-Resource Languages20 Jun 2024 0 repositories listed
-
Enhancing Distractor Generation for Multiple-Choice Questions with Retrieval Augmented Pretraining and Knowledge Graph Integration19 Jun 2024 0 repositories listed
-
QRMeM: Unleash the Length Limitation through Question then Reflection Memory Mechanism19 Jun 2024 0 repositories listed
-
Aqulia-Med LLM: Pioneering Full-Process Open-Source Medical Language Models18 Jun 2024 0 repositories listed
-
On the Principles behind Opinion Dynamics in Multi-Agent Systems of Large Language Models18 Jun 2024 0 repositories listed
-
QOG:Question and Options Generation based on Language Model18 Jun 2024 0 repositories listed
-
VELOCITI: Benchmarking Video-Language Compositional Reasoning with Strict Entailment16 Jun 2024 0 repositories listed
-
VCEval: Rethinking What is a Good Educational Video and How to Automatically Evaluate It15 Jun 2024 0 repositories listed
-
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models13 Jun 2024 0 repositories listed
-
Bayesian Statistical Modeling with Predictors from LLMs13 Jun 2024 0 repositories listed
-
OLMES: A Standard for Language Model Evaluations12 Jun 2024 0 repositories listed
-
Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context10 Jun 2024 0 repositories listed
-
Towards a Personal Health Large Language Model10 Jun 2024 0 repositories listed
-
Do LLMs Recognize me, When I is not me: Assessment of LLMs Understanding of Turkish Indexical Pronouns in Indexical Shift Contexts8 Jun 2024 0 repositories listed
-
Investigating and Addressing Hallucinations of LLMs in Tasks Involving Negation8 Jun 2024 0 repositories listed
-
CRiskEval: A Chinese Multi-Level Risk Evaluation Benchmark Dataset for Large Language Models7 Jun 2024 0 repositories listed
-
Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?6 Jun 2024 0 repositories listed
-
Explore then Determine: A GNN-LLM Synergy Framework for Reasoning over Knowledge Graph3 Jun 2024 0 repositories listed
-
DGRC: An Effective Fine-tuning Framework for Distractor Generation in Chinese Multi-choice Reading Comprehension29 May 2024 0 repositories listed
-
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints28 May 2024 0 repositories listed
-
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer27 May 2024 0 repositories listed
-
Imagery as Inquiry: Exploring A Multimodal Dataset for Conversational Recommendation23 May 2024 0 repositories listed
-
Robust portfolio optimization model for electronic coupon allocation21 May 2024 0 repositories listed
-
Exploring the Capabilities of Prompted Large Language Models in Educational and Assessment Applications19 May 2024 0 repositories listed
-
COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain17 May 2024 0 repositories listed
-
From Generalist to Specialist: Improving Large Language Models for Medical Physics Using ARCoT17 May 2024 0 repositories listed
-
AmazUtah_NLP at SemEval-2024 Task 9: A MultiChoice Question Answering System for Commonsense Defying Reasoning16 May 2024 0 repositories listed
-
14 May 2024 0 repositories listed
-
MCS-SQL: Leveraging Multiple Prompts and Multiple-Choice Selection For Text-to-SQL Generation13 May 2024 0 repositories listed
-
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning6 May 2024 0 repositories listed
-
Math Multiple Choice Question Generation via Human-Large Language Model Collaboration1 May 2024 0 repositories listed
-
FoundaBench: Evaluating Chinese Fundamental Knowledge Capabilities of Large Language Models29 Apr 2024 0 repositories listed
-
AI and Machine Learning for Next Generation Science Assessments23 Apr 2024 0 repositories listed
-
Improving Automated Distractor Generation for Math Multiple-choice Questions with Overgenerate-and-rank19 Apr 2024 0 repositories listed
-
18 Apr 2024 0 repositories listed
-
Is There No Such Thing as a Bad Question? H4R: HalluciBot For Ratiocination, Rewriting, Ranking, and Routing18 Apr 2024 0 repositories listed
-
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models17 Apr 2024 0 repositories listed
-
Question Difficulty Ranking for Multiple-Choice Reading Comprehension16 Apr 2024 0 repositories listed
-
9 Apr 2024 0 repositories listed
-
Cleared for Takeoff? Compositional & Conditional Reasoning may be the Achilles Heel to (Flight-Booking) Language Agents5 Apr 2024 0 repositories listed
-
LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language Models19 Mar 2024 0 repositories listed
-
Enhancing Event Causality Identification with Rationale and Structure-Aware Causal Question Answering17 Mar 2024 0 repositories listed
-
Few-Shot Image Classification and Segmentation as Visual Question Answering Using Vision-Language Models15 Mar 2024 0 repositories listed
-
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic14 Mar 2024 0 repositories listed
-
Exploring the Comprehension of ChatGPT in Traditional Chinese Medicine Knowledge14 Mar 2024 0 repositories listed
-
Rethinking Generative Large Language Model Evaluation for Semantic Comprehension12 Mar 2024 0 repositories listed
-
MedKP: Medical Dialogue with Knowledge Enhancement and Clinical Pathway Encoding11 Mar 2024 0 repositories listed
-
An Improved Traditional Chinese Evaluation Suite for Foundation Model4 Mar 2024 0 repositories listed
-
Automated Generation of Multiple-Choice Cloze Questions for Assessing English Vocabulary Using GPT-turbo 3.54 Mar 2024 0 repositories listed
-
Controlling Cloze-test Question Item Difficulty with PLM-based Surrogate Models for IRT Assessment3 Mar 2024 0 repositories listed
-
KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations3 Mar 2024 0 repositories listed
-
Predictions from language models for multiple-choice tasks are not robust under variation of scoring methods1 Mar 2024 0 repositories listed
-
Unsupervised multiple choices question answering via universal corpus27 Feb 2024 0 repositories listed
-
Identifying Multiple Personalities in Large Language Models with External Evaluation22 Feb 2024 0 repositories listed
-
Beyond Probabilities: Unveiling the Misalignment in Evaluating Large Language Models21 Feb 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.