Browse State-of-the-Art › Multiple-choice › Papers, page 9
Multiple-choice
Papers archive 2025-07-28
archive papers tagged: 1,107 · with a code link: 483 · where Syntology ran a sample: 161 (124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (161 of 1,107 tagged: 124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument)
Page 9 of 12: papers 801 to 900 of 1,107, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
KorNAT: LLM Alignment Benchmark for Korean Social Values and Common Knowledge21 Feb 2024 0 repositories listed
-
Ranking Large Language Models without Ground Truth21 Feb 2024 0 repositories listed
-
Digital Comprehensibility Assessment of Simplified Texts among Persons with Intellectual Disabilities20 Feb 2024 0 repositories listed
-
Stick to your Role! Stability of Personal Values Expressed in Large Language Models19 Feb 2024 0 repositories listed
-
KMMLU: Measuring Massive Multitask Language Understanding in Korean18 Feb 2024 0 repositories listed
-
Prompting Implicit Discourse Relation Annotation7 Feb 2024 0 repositories listed
-
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification6 Feb 2024 0 repositories listed
-
SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark6 Feb 2024 0 repositories listed
-
LLMs May Perform MCQA by Selecting the Least Incorrect Option2 Feb 2024 0 repositories listed
-
Distractor Generation in Multiple-Choice Tasks: A Survey of Methods, Datasets, and Evaluation2 Feb 2024 0 repositories listed
-
Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis28 Jan 2024 0 repositories listed
-
Towards Collective Superintelligence: Amplifying Group IQ using Conversational Swarms25 Jan 2024 0 repositories listed
-
Instruction Fine-Tuning: Does Prompt Loss Matter?24 Jan 2024 0 repositories listed
-
What Large Language Models Know and What People Think They Know24 Jan 2024 0 repositories listed
-
Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding13 Jan 2024 0 repositories listed
-
Automated Answer Validation using Text Similarity13 Jan 2024 0 repositories listed
-
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs' Pragmatics Capabilities13 Jan 2024 0 repositories listed
-
A Joint-Reasoning based Disease Q&A System6 Jan 2024 0 repositories listed
-
The Earth is Flat? Unveiling Factual Errors in Large Language Models1 Jan 2024 0 repositories listed
-
FusionMind -- Improving question and answering with external context fusion31 Dec 2023 0 repositories listed
-
BloomVQA: Assessing Hierarchical Multi-modal Comprehension20 Dec 2023 0 repositories listed
-
Perception Test 2023: A Summary of the First Challenge And Outcome20 Dec 2023 0 repositories listed
-
Self-Evaluation Improves Selective Generation in Large Language Models14 Dec 2023 0 repositories listed
-
A Foundational Multimodal Vision Language AI Assistant for Human Pathology13 Dec 2023 0 repositories listed
-
A Comparative Study of AI-Generated (GPT-4) and Human-crafted MCQs in Programming Education5 Dec 2023 0 repositories listed
-
Unleashing the Potential of Large Language Model: Zero-shot VQA for Flood Disaster Scenario4 Dec 2023 0 repositories listed
-
Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading Comprehension30 Nov 2023 0 repositories listed
-
Investigating Data Contamination in Modern Benchmarks for Large Language Models16 Nov 2023 0 repositories listed
-
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology16 Nov 2023 0 repositories listed
-
Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset14 Nov 2023 0 repositories listed
-
Characterizing Large Language Models as Rationalizers of Knowledge-intensive Tasks9 Nov 2023 0 repositories listed
-
Assessing Distractors in Multiple-Choice Tests8 Nov 2023 0 repositories listed
-
Evaluating multiple large language models in pediatric ophthalmology7 Nov 2023 0 repositories listed
-
Evaluating the Potential of Leading Large Language Models in Reasoning Biology Questions5 Nov 2023 0 repositories listed
-
More Robots are Coming: Large Multimodal Models (ChatGPT) can Solve Visually Diverse Images of Parsons Problems3 Nov 2023 0 repositories listed
-
DeSIQ: Towards an Unbiased, Challenging Benchmark for Social Intelligence Understanding24 Oct 2023 0 repositories listed
-
Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond23 Oct 2023 0 repositories listed
-
Evaluating the Symbol Binding Ability of Large Language Models for Multiple-Choice Questions in Vietnamese General Education18 Oct 2023 0 repositories listed
-
Field-testing items using artificial intelligence: Natural language processing with transformers18 Oct 2023 0 repositories listed
-
Investigating Uncertainty Calibration of Aligned Language Models under the Multiple-Choice Setting18 Oct 2023 0 repositories listed
-
Mitigating Bias for Question Answering Models by Tracking Bias Influence13 Oct 2023 0 repositories listed
-
On the Performance of Multimodal Language Models4 Oct 2023 0 repositories listed
-
Automating question generation from educational text26 Sep 2023 0 repositories listed
-
HANS, are you clever? Clever Hans Effect Analysis of Neural Systems21 Sep 2023 0 repositories listed
-
Benchmarks for Pirá 2.0, a Reading Comprehension Dataset about the Ocean, the Brazilian Coast, and Climate Change19 Sep 2023 0 repositories listed
-
Language models are susceptible to incorrect patient self-diagnosis in medical applications17 Sep 2023 0 repositories listed
-
Self-Assessment Tests are Unreliable Measures of LLM Personality15 Sep 2023 0 repositories listed
-
Performance of ChatGPT-3.5 and GPT-4 on the United States Medical Licensing Examination With and Without Distractions12 Sep 2023 0 repositories listed
-
Use neural networks to recognize students' handwritten letters and incorrect symbols12 Sep 2023 0 repositories listed
-
An Automatic Evaluation Framework for Multi-turn Medical Consultations Capabilities of Large Language Models5 Sep 2023 0 repositories listed
-
Generalised Winograd Schema and its Contextuality31 Aug 2023 0 repositories listed
-
Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions22 Aug 2023 0 repositories listed
-
A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology9 Aug 2023 0 repositories listed
-
Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla18 Jul 2023 0 repositories listed
-
Analyzing Multiple-Choice Reading and Listening Comprehension Tests3 Jul 2023 0 repositories listed
-
Analysis of the Cambridge Multiple-Choice Questions Reading Dataset with a Focus on Candidate Response Distribution22 Jun 2023 0 repositories listed
-
RECAP-KG: Mining Knowledge Graphs from Raw GP Notes for Remote COVID-19 Assessment in Primary Care17 Jun 2023 0 repositories listed
-
Can ChatGPT pass the Vietnamese National High School Graduation Examination?15 Jun 2023 0 repositories listed
-
Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses15 Jun 2023 0 repositories listed
-
Investigating the Effectiveness of ChatGPT in Mathematical Reasoning and Problem Solving: Evidence from the Vietnamese National High School Graduation Examination10 Jun 2023 0 repositories listed
-
Network-based Representations and Dynamic Discrete Choice Models for Multiple Discrete Choice Analysis7 Jun 2023 0 repositories listed
-
BUCA: A Binary Classification Approach to Unsupervised Commonsense Question Answering25 May 2023 0 repositories listed
-
Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs24 May 2023 0 repositories listed
-
Make a Choice! Knowledge Base Question Answering with In-Context Learning23 May 2023 0 repositories listed
-
Query Rewriting for Retrieval-Augmented Large Language Models23 May 2023 0 repositories listed
-
Contextual Response Interpretation for Automated Structured Interviews: A Case Study in Market Research30 Apr 2023 0 repositories listed
-
Who's the Best Detective? LLMs vs. MLs in Detecting Incoherent Fourth Grade Math Answers21 Apr 2023 0 repositories listed
-
Analyzing the Performance of ChatGPT in Cardiology and Vascular Pathologies15 Apr 2023 0 repositories listed
-
Prompt Engineering and Calibration for Zero-Shot Commonsense Reasoning14 Apr 2023 0 repositories listed
-
DISTO: Evaluating Textual Distractors for Multi-Choice Questions using Negative Sampling based Approach10 Apr 2023 0 repositories listed
-
Bridging the Language Gap: Knowledge Injected Multilingual Question Answering6 Apr 2023 0 repositories listed
-
GPT-4 to GPT-3.5: 'Hold My Scalpel' -- A Look at the Competency of OpenAI's GPT on the Plastic Surgery In-Service Training Exam4 Apr 2023 0 repositories listed
-
Automatic Generation of Multiple-Choice Questions25 Mar 2023 0 repositories listed
-
A Graph-Guided Reasoning Approach for Open-ended Commonsense Question Answering18 Mar 2023 0 repositories listed
-
Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses?16 Mar 2023 0 repositories listed
-
Generating multiple-choice questions for medical question answering with distractors and cue-masking13 Mar 2023 0 repositories listed
-
10 Mar 2023 0 repositories listed
-
Large Language Models (GPT) Struggle to Answer Multiple-Choice Questions about Code9 Mar 2023 0 repositories listed
-
Empowering Sentence Encoders with Prompting and Label Retrieval for Zero-shot Text Classification20 Dec 2022 0 repositories listed
-
True Detective: A Deep Abductive Reasoning Benchmark Undoable for GPT-3 and Challenging for GPT-420 Dec 2022 0 repositories listed
-
Question-type Identification for Academic Questions in Online Learning Platform24 Nov 2022 0 repositories listed
-
AGReE: A system for generating Automated Grammar Reading Exercises28 Oct 2022 0 repositories listed
-
AI-based Arabic Language and Speech Tutor22 Oct 2022 0 repositories listed
-
Two-Turn Debate Doesn't Help Humans Answer Hard Reading Comprehension Questions19 Oct 2022 0 repositories listed
-
Understanding Prior Bias and Choice Paralysis in Transformer-based Language Representation Models through Four Experimental Probes3 Oct 2022 0 repositories listed
-
A Weak Supervision Approach for Predicting Difficulty of Technical Interview Questions1 Oct 2022 0 repositories listed
-
Document-level Event Factuality Identification via Machine Reading Comprehension Frameworks with Transfer Learning1 Oct 2022 0 repositories listed
-
Using contradictions improves question answering systems28 Sep 2022 0 repositories listed
-
Treatment Effects with Multidimensional Unobserved Heterogeneity: Identification of the Marginal Treatment Effect23 Sep 2022 0 repositories listed
-
Multiple-Choice Question Generation: Towards an Automated Assessment Framework23 Sep 2022 0 repositories listed
-
Scheduling Algorithms for Federated Learning with Minimal Energy Consumption13 Sep 2022 0 repositories listed
-
Zero-shot Event Causality Identification with Question Answering1 Sep 2022 0 repositories listed
-
From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams11 Jun 2022 0 repositories listed
-
CroaTPAS: A Survey-based Evaluation1 Jun 2022 0 repositories listed
-
HRCA+: Advanced Multiple-choice Machine Reading Comprehension Method1 Jun 2022 0 repositories listed
-
PADDLe: a Platform to Identify Complex Words for Learners of French as a Foreign Language (FFL)1 Jun 2022 0 repositories listed
-
1 May 2022 0 repositories listed
-
Automatic Generation of Distractors for Fill-in-the-Blank Exercises with Round-Trip Neural Machine Translation1 May 2022 0 repositories listed
-
Clozer”:" Adaptable Data Augmentation for Cloze-style Reading Comprehension1 May 2022 0 repositories listed
-
Unsupervised multiple-choice question generation for out-of-domain Q&A fine-tuning1 May 2022 0 repositories listed