Browse State-of-the-Art › Multiple-choice › Papers, page 6
Multiple-choice
Papers archive 2025-07-28
archive papers tagged: 1,107 · with a code link: 483 · where Syntology ran a sample: 161 (124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (161 of 1,107 tagged: 124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument)
Page 6 of 12: papers 501 to 600 of 1,107, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms5 Jun 2025 0 repositories listed
-
Multiple-Choice Question Generation Using Large Language Models: Methodology and Educator Insights5 Jun 2025 0 repositories listed
-
Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales4 Jun 2025 0 repositories listed
-
Performance of leading large language models in May 2025 in Membership of the Royal College of General Practitioners-style examination questions: a cross-sectional analysis3 Jun 2025 0 repositories listed
-
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation2 Jun 2025 0 repositories listed
-
Beyond Multiple Choice: Evaluating Steering Vectors for Adaptive Free-Form Summarization30 May 2025 0 repositories listed
-
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases30 May 2025 0 repositories listed
-
PersianMedQA: Language-Centric Evaluation of LLMs in the Persian Medical Domain30 May 2025 0 repositories listed
-
VUDG: A Dataset for Video Understanding Domain Generalization30 May 2025 0 repositories listed
-
Image Aesthetic Reasoning: A New Benchmark for Medical Image Screening with MLLMs29 May 2025 0 repositories listed
-
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence29 May 2025 0 repositories listed
-
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine29 May 2025 0 repositories listed
-
Large Language Models Often Know When They Are Being Evaluated28 May 2025 0 repositories listed
-
SOSBENCH: Benchmarking Safety Alignment on Scientific Knowledge27 May 2025 0 repositories listed
-
CP-Router: An Uncertainty-Aware Router Between LLM and LRM26 May 2025 0 repositories listed
-
DFIR-Metric: A Benchmark Dataset for Evaluating Large Language Models in Digital Forensics and Incident Response26 May 2025 0 repositories listed
-
Genome-Bench: A Scientific Reasoning Benchmark from Real-World Expert Discussions26 May 2025 0 repositories listed
-
My Answer Is NOT 'Fair': Mitigating Social Bias in Vision-Language Models via Fair and Biased Residuals26 May 2025 0 repositories listed
-
Enhancing LLMs' Reasoning-Intensive Multimedia Search Capabilities through Fine-Tuning and Reinforcement Learning24 May 2025 0 repositories listed
-
AutoMCQ -- Automatically Generate Code Comprehension Questions using GenAI22 May 2025 0 repositories listed
-
Collaboration among Multiple Large Language Models for Medical Question Answering22 May 2025 0 repositories listed
-
KoBALT: Korean Benchmark For Advanced Linguistic Tasks22 May 2025 0 repositories listed
-
Improving LLM First-Token Predictions in Multiple-Choice Question Answering via Prefilling Attack21 May 2025 0 repositories listed
-
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets21 May 2025 0 repositories listed
-
Set-LLM: A Permutation-Invariant LLM21 May 2025 0 repositories listed
-
Uncovering Cultural Representation Disparities in Vision-Language Models20 May 2025 0 repositories listed
-
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications20 May 2025 0 repositories listed
-
LEXam: Benchmarking Legal Reasoning on 340 Law Exams19 May 2025 0 repositories listed
-
MR. Judge: Multimodal Reasoner as a Judge19 May 2025 0 repositories listed
-
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models16 May 2025 0 repositories listed
-
ZeroTuning: Unlocking the Initial Token's Power to Enhance Large Language Models Without Training16 May 2025 0 repositories listed
-
Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation15 May 2025 0 repositories listed
-
The CoT Encyclopedia: Analyzing, Predicting, and Controlling how a Reasoning Model will Think15 May 2025 0 repositories listed
-
KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive Reasoning14 May 2025 0 repositories listed
-
SafePath: Conformal Prediction for Safe LLM-Based Autonomous Navigation14 May 2025 0 repositories listed
-
How well do LLMs reason over tabular data, really?12 May 2025 0 repositories listed
-
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information9 May 2025 0 repositories listed
-
Tell Me Who Your Students Are: GPT Can Generate Valid Multiple-Choice Questions When Students' (Mis)Understanding Is Hinted9 May 2025 0 repositories listed
-
Developing A Framework to Support Human Evaluation of Bias in Generated Free Response Text5 May 2025 0 repositories listed
-
5 May 2025 0 repositories listed Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LLM-based Text Simplification and its Effect on User Comprehension and Cognitive Load4 May 2025 0 repositories listed
-
LookAlike: Consistent Distractor Generation in Math MCQs3 May 2025 0 repositories listed
-
Adaptive Wizard for Removing Cross-Tier Misconfigurations in Active Directory2 May 2025 0 repositories listed
-
SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning22 Apr 2025 0 repositories listed
-
LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception21 Apr 2025 0 repositories listed
-
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models20 Apr 2025 0 repositories listed
-
Assessing AI-Generated Questions' Alignment with Cognitive Frameworks in Educational Assessment19 Apr 2025 0 repositories listed
-
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model18 Apr 2025 0 repositories listed
-
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain18 Apr 2025 0 repositories listed
-
Benchmarking Next-Generation Reasoning-Focused Large Language Models in Ophthalmology: A Head-to-Head Evaluation on 5,888 Items15 Apr 2025 0 repositories listed
-
AgMMU: A Comprehensive Agricultural Multimodal Understanding and Reasoning Benchmark14 Apr 2025 0 repositories listed
-
Large Language Models Could Be Rote Learners11 Apr 2025 0 repositories listed
-
InstructionBench: An Instructional Video Understanding Benchmark7 Apr 2025 0 repositories listed
-
Can AI Master Construction Management (CM)? Benchmarking State-of-the-Art Large Language Models on CM Certification Exams4 Apr 2025 0 repositories listed
-
From ChatGPT to DeepSeek AI: A Comprehensive Analysis of Evolution, Deviation, and Future Implications in AI-Language Models4 Apr 2025 0 repositories listed
-
ACPBench Hard: Unrestrained Reasoning about Action, Change, and Planning31 Mar 2025 0 repositories listed
-
Order Independence With Finetuning30 Mar 2025 0 repositories listed
-
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering23 Mar 2025 0 repositories listed
-
Evaluating Clinical Competencies of Large Language Models with a General Practice Benchmark22 Mar 2025 0 repositories listed
-
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia21 Mar 2025 0 repositories listed
-
AutoDrive-QA- Automated Generation of Multiple-Choice Questions for Autonomous Driving Datasets Using Large Vision-Language Models20 Mar 2025 0 repositories listed
-
CodeReviewQA: The Code Review Comprehension Assessment for Large Language Models20 Mar 2025 0 repositories listed
-
FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding19 Mar 2025 0 repositories listed
-
VisNumBench: Evaluating Number Sense of Multimodal Large Language Models19 Mar 2025 0 repositories listed
-
Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data13 Mar 2025 0 repositories listed
-
It is Too Many Options: Pitfalls of Multiple-Choice Questions in Generative AI and Medical Education13 Mar 2025 0 repositories listed
-
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory13 Mar 2025 0 repositories listed
-
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words10 Mar 2025 0 repositories listed
-
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations10 Mar 2025 0 repositories listed
-
Towards Conversational AI for Disease Management8 Mar 2025 0 repositories listed
-
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces8 Mar 2025 0 repositories listed
-
Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework7 Mar 2025 0 repositories listed
-
Structured Outputs Enable General-Purpose LLMs to be Medical Experts5 Mar 2025 0 repositories listed
-
The impact of AI and peer feedback on research writing skills: a study using the CGScholar platform among Kazakhstani scholars5 Mar 2025 0 repositories listed
-
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering3 Mar 2025 0 repositories listed
-
MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts28 Feb 2025 0 repositories listed
-
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning27 Feb 2025 0 repositories listed
-
ANPMI: Assessing the True Comprehension Capabilities of LLMs for Multiple Choice Questions26 Feb 2025 0 repositories listed
-
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning25 Feb 2025 0 repositories listed
-
Reversal Blessing: Thinking Backward May Outpace Thinking Forward in Multi-choice Questions25 Feb 2025 0 repositories listed
-
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models25 Feb 2025 0 repositories listed
-
The Lazy Student's Dream: ChatGPT Passing an Engineering Course on Its Own23 Feb 2025 0 repositories listed
-
LegalBench.PT: A Benchmark for Portuguese Law22 Feb 2025 0 repositories listed
-
Do LLMs Make Mistakes Like Students? Exploring Natural Alignment between Language Models and Human Error Patterns21 Feb 2025 0 repositories listed
-
MHQA: A Diverse, Knowledge Intensive Mental Health Question Answering Challenge for Language Models21 Feb 2025 0 repositories listed
-
Fundamental Limitations in Defending LLM Finetuning APIs20 Feb 2025 0 repositories listed
-
MCQA-Eval: Efficient Confidence Evaluation in NLG with Gold-Standard Correctness Labels20 Feb 2025 0 repositories listed
-
Unveiling Cultural Blind Spots: Analyzing the Limitations of mLLMs in Procedural Text Comprehension20 Feb 2025 0 repositories listed
-
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh19 Feb 2025 0 repositories listed
-
Is This Collection Worth My LLM's Time? Automatically Measuring Information Potential in Text Corpora19 Feb 2025 0 repositories listed
-
Towards Geo-Culturally Grounded LLM Generations19 Feb 2025 0 repositories listed
-
VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare19 Feb 2025 0 repositories listed
-
Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above19 Feb 2025 0 repositories listed
-
Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs18 Feb 2025 0 repositories listed
-
None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks18 Feb 2025 0 repositories listed
-
OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities18 Feb 2025 0 repositories listed
-
Multi-Modal Retrieval Augmentation for Open-Ended and Knowledge-Intensive Video Question Answering17 Feb 2025 0 repositories listed
-
LogiDynamics: Unraveling the Dynamics of Logical Inference in Large Language Model Reasoning16 Feb 2025 0 repositories listed
-
14 Feb 2025 0 repositories listed
-
Objective quantification of mood states using large language models13 Feb 2025 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.