Browse State-of-the-Art › Multiple-choice › Papers, page 7
Multiple-choice
Papers archive 2025-07-28
archive papers tagged: 1,107 · with a code link: 483 · where Syntology ran a sample: 161 (124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (161 of 1,107 tagged: 124 with a run with no instrument failure, 37 where every run was a failure of Syntology's instrument)
Page 7 of 12: papers 601 to 700 of 1,107, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
A Semantic Parsing Algorithm to Solve Linear Ordering Problems12 Feb 2025 0 repositories listed
-
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs12 Feb 2025 0 repositories listed
-
SB-Bench: Stereotype Bias Benchmark for Large Multimodal Models12 Feb 2025 0 repositories listed
-
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian11 Feb 2025 0 repositories listed
-
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark10 Feb 2025 0 repositories listed
-
LLMs to Support a Domain Specific Knowledge Assistant6 Feb 2025 0 repositories listed
-
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs6 Feb 2025 0 repositories listed
-
Evalita-LLM: Benchmarking Large Language Models on Italian4 Feb 2025 0 repositories listed
-
The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts3 Feb 2025 0 repositories listed
-
CoddLLM: Empowering Large Language Models for Data Analytics1 Feb 2025 0 repositories listed
-
InnerThoughts: Disentangling Representations and Predictions in Large Language Models29 Jan 2025 0 repositories listed
-
Attribution analysis of legal language as used by LLM28 Jan 2025 0 repositories listed
-
Inferring from Logits: Exploring Best Practices for Decoding-Free Generative Candidate Selection28 Jan 2025 0 repositories listed
-
Town Hall Debate Prompting: Enhancing Logical Reasoning in LLMs through Multi-Persona Interaction28 Jan 2025 0 repositories listed
-
Options-Aware Dense Retrieval for Multiple-Choice query Answering27 Jan 2025 0 repositories listed
-
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI26 Jan 2025 0 repositories listed
-
LLM Evaluation Based on Aerospace Manufacturing Expertise: Automated Generation and Multi-Model Question Answering25 Jan 2025 0 repositories listed
-
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion25 Jan 2025 0 repositories listed
-
Humanity's Last Exam24 Jan 2025 0 repositories listed
-
Auto-Evaluation: A Critical Measure in Driving Improvements in Quality and Safety of AI-Generated Lesson Resources23 Jan 2025 0 repositories listed
-
On the Reasoning Capacity of AI Models and How to Quantify It23 Jan 2025 0 repositories listed
-
The AI Penalization Effect: People Reduce Compensation for Workers Who Use AI22 Jan 2025 0 repositories listed
-
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction21 Jan 2025 0 repositories listed
-
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!18 Jan 2025 0 repositories listed
-
Empowering Large Language Models in Wireless Communication: A Novel Dataset and Fine-Tuning Framework16 Jan 2025 0 repositories listed
-
Multiple Choice Questions: Reasoning Makes Large Language Models (LLMs) More Self-Confident Even When They Are Wrong16 Jan 2025 0 repositories listed
-
Vision-Language Models Do Not Understand Negation16 Jan 2025 0 repositories listed
-
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History15 Jan 2025 0 repositories listed
-
Rethinking AI Cultural Alignment13 Jan 2025 0 repositories listed
-
Hierarchical Divide-and-Conquer for Fine-Grained Alignment in LLM-Based Medical Evaluation12 Jan 2025 0 repositories listed
-
First Token Probability Guided RAG for Telecom Question Answering11 Jan 2025 0 repositories listed
-
DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests8 Jan 2025 0 repositories listed
-
Knowledge Retrieval Based on Generative AI8 Jan 2025 0 repositories listed
-
Localizing AI: Evaluating Open-Weight Language Models for Languages of Baltic States7 Jan 2025 0 repositories listed
-
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering2 Jan 2025 0 repositories listed
-
Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation1 Jan 2025 0 repositories listed
-
FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding1 Jan 2025 0 repositories listed
-
IllusionBench: A Large-scale and Comprehensive Benchmark for Visual Illusion Understanding in Vision-Language Models1 Jan 2025 0 repositories listed
-
Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs1 Jan 2025 0 repositories listed
-
Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answering1 Jan 2025 0 repositories listed
-
A review of faithfulness metrics for hallucination assessment in Large Language Models31 Dec 2024 0 repositories listed
-
AraSTEM: A Native Arabic Multiple Choice Question Benchmark for Evaluating LLMs Knowledge In STEM Subjects31 Dec 2024 0 repositories listed
-
EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta31 Dec 2024 0 repositories listed
-
Monty Hall and Optimized Conformal Prediction to Improve Decision-Making with LLMs31 Dec 2024 0 repositories listed
-
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation31 Dec 2024 0 repositories listed
-
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity30 Dec 2024 0 repositories listed
-
HindiLLM: Large Language Model for Hindi29 Dec 2024 0 repositories listed
-
Using Large Language Models for Automated Grading of Student Writing about Science25 Dec 2024 0 repositories listed
-
In Case You Missed It: ARC 'Challenge' Is Not That Challenging23 Dec 2024 0 repositories listed
-
Are You Doubtful? Oh, It Might Be Difficult Then! Exploring the Use of Model Uncertainty for Question Difficulty Estimation16 Dec 2024 0 repositories listed
-
Auto-bidding in real-time auctions via Oracle Imitation Learning (OIL)16 Dec 2024 0 repositories listed
-
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding16 Dec 2024 0 repositories listed
-
Seeing the Forest and the Trees: Solving Visual Graph and Tree Based Data Structure Problems using Large Multimodal Models15 Dec 2024 0 repositories listed
-
A recent evaluation on the performance of LLMs on radiation oncology physics using questions of randomly shuffled options14 Dec 2024 0 repositories listed
-
Do LLMs Act as Repositories of Causal Knowledge?14 Dec 2024 0 repositories listed
-
Superhuman performance of a large language model on the reasoning tasks of a physician14 Dec 2024 0 repositories listed
-
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing13 Dec 2024 0 repositories listed
-
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering13 Dec 2024 0 repositories listed
-
ACQ: A Unified Framework for Automated Programmatic Creativity in Online Advertising9 Dec 2024 0 repositories listed
-
MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects6 Dec 2024 0 repositories listed
-
Establishing Task Scaling Laws via Compute-Efficient Model Ladders5 Dec 2024 0 repositories listed
-
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering5 Dec 2024 0 repositories listed
-
The use of large language models to enhance cancer clinical trial educational materials2 Dec 2024 0 repositories listed
-
Unlocking Video-LLM via Agent-of-Thoughts Distillation2 Dec 2024 0 repositories listed
-
Uhura: A Benchmark for Evaluating Scientific Question Answering and Truthfulness in Low-Resource African Languages1 Dec 2024 0 repositories listed
-
Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments30 Nov 2024 0 repositories listed
-
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark29 Nov 2024 0 repositories listed
-
Applying IRT to Distinguish Between Human and Generative AI Responses to Multiple-Choice Assessments28 Nov 2024 0 repositories listed
-
Sparse Attention Vectors: Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers28 Nov 2024 0 repositories listed
-
Multiple Choice Learning for Efficient Speech Separation with Many Speakers27 Nov 2024 0 repositories listed
-
NEMO: Can Multimodal LLMs Identify Attribute-Modified Objects?26 Nov 2024 0 repositories listed
-
GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis25 Nov 2024 0 repositories listed
-
SAGEval: The frontiers of Satisfactory Agent based NLG Evaluation for reference-free open-ended text25 Nov 2024 0 repositories listed
-
AfriMed-QA: A Pan-African, Multi-Specialty, Medical Question-Answering Benchmark Dataset23 Nov 2024 0 repositories listed
-
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation20 Nov 2024 0 repositories listed
-
Testing Uncertainty of Large Language Models for Physics Knowledge and Reasoning18 Nov 2024 0 repositories listed
-
A Benchmark for Long-Form Medical Question Answering14 Nov 2024 0 repositories listed
-
SHARP: Unlocking Interactive Hallucination via Stance Transfer in Role-Playing Agents12 Nov 2024 0 repositories listed
-
Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability10 Nov 2024 0 repositories listed
-
Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators8 Nov 2024 0 repositories listed
-
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding7 Nov 2024 0 repositories listed
-
FactTest: Factuality Testing in Large Language Models with Finite-Sample and Distribution-Free Guarantees4 Nov 2024 0 repositories listed
-
Enhancing LLM Evaluations: The Garbling Trick3 Nov 2024 0 repositories listed
-
Benchmarking Bias in Large Language Models during Role-Playing1 Nov 2024 0 repositories listed
-
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest27 Oct 2024 0 repositories listed
-
25 Oct 2024 0 repositories listed
-
Beyond Multiple-Choice Accuracy: Real-World Challenges of Implementing Large Language Models in Healthcare24 Oct 2024 0 repositories listed
-
Large Language Models Still Exhibit Bias in Long Text23 Oct 2024 0 repositories listed
-
GeoCode-GPT: A Large Language Model for Geospatial Code Generation Tasks22 Oct 2024 0 repositories listed
-
Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the U.S21 Oct 2024 0 repositories listed
-
Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models18 Oct 2024 0 repositories listed
-
LabSafety Bench: Benchmarking LLMs on Safety Issues in Scientific Labs18 Oct 2024 0 repositories listed
-
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy17 Oct 2024 0 repositories listed
-
LAR-ECHR: A New Legal Argument Reasoning Task and Dataset for Cases of the European Court of Human Rights17 Oct 2024 0 repositories listed
-
Not All Options Are Created Equal: Textual Option Weighting for Token-Efficient LLM-Based Knowledge Tracing14 Oct 2024 0 repositories listed
-
Personalised Feedback Framework for Online Education Programmes Using Generative AI14 Oct 2024 0 repositories listed
-
LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models13 Oct 2024 0 repositories listed
-
The Future of Learning in the Age of Generative AI: Automated Question Generation and Assessment with Large Language Models12 Oct 2024 0 repositories listed
-
MRAG-Bench: Vision-Centric Evaluation for Retrieval-Augmented Multimodal Models10 Oct 2024 0 repositories listed
-
Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models10 Oct 2024 0 repositories listed