Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 14
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 14 of 29: papers 1,301 to 1,400 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Humor Mechanics: Advancing Humor Generation with Multistep Reasoning 12 May 2024 · 1 repository · arXiv:2405.07280
-
L(u)PIN: LLM-based Political Ideology Nowcasting 12 May 2024 · 0 repositories · arXiv:2405.07320
-
Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis 12 May 2024 · 1 repository · arXiv:2405.07248Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MedConceptsQA: Open Source Medical Concepts QA Benchmark 12 May 2024 · 1 repository · arXiv:2405.07348Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Automating Thematic Analysis: How LLMs Analyse Controversial Topics 11 May 2024 · 0 repositories · arXiv:2405.06919
-
Identifying Key Terms in Prompts for Relevance Evaluation with GPT Models 11 May 2024 · 0 repositories · arXiv:2405.06931
-
TacoERE: Cluster-aware Compression for Event Relation Extraction 11 May 2024 · 0 repositories · arXiv:2405.06890
-
CANAL -- Cyber Activity News Alerting Language Model: Empirical Approach vs. Expensive LLM 10 May 2024 · 0 repositories · arXiv:2405.06772
-
Large Language Model in Financial Regulatory Interpretation 10 May 2024 · 0 repositories · arXiv:2405.06808
-
Multimodal LLMs Struggle with Basic Visual Network Analysis: a VNA Benchmark 10 May 2024 · 1 repository · arXiv:2405.06634
-
Artificial Intelligence as the New Hacker: Developing Agents for Offensive Security 9 May 2024 · 0 repositories · arXiv:2406.07561
-
Can large language models understand uncommon meanings of common words? 9 May 2024 · 0 repositories · arXiv:2405.05741
-
Digital Diagnostics: The Potential Of Large Language Models In Recognizing Symptoms Of Common Illnesses 9 May 2024 · 0 repositories · arXiv:2405.06712
-
Letter to the Editor: What are the legal and ethical considerations of submitting radiology reports to ChatGPT? 9 May 2024 · 0 repositories · arXiv:2405.05647
-
People cannot distinguish GPT-4 from a human in a Turing test 9 May 2024 · 0 repositories · arXiv:2405.08007
-
Smurfs: Leveraging Multiple Proficiency Agents with Context-Efficiency for Tool Planning 9 May 2024 · 1 repository · arXiv:2405.05955Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ACORN: Aspect-wise Commonsense Reasoning Explanation Evaluation 8 May 2024 · 1 repository · arXiv:2405.04818
-
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge 8 May 2024 · 1 repository · arXiv:2405.05253
-
Seeds of Stereotypes: A Large-Scale Textual Analysis of Race and Gender Associations with Diseases in Online Sources 8 May 2024 · 0 repositories · arXiv:2405.05049
-
Enhancing the Efficiency and Accuracy of Underlying Asset Reviews in Structured Finance: The Application of Multi-agent Framework 7 May 2024 · 1 repository · arXiv:2405.04294
-
NaturalCodeBench: Examining Coding Performance Mismatch on HumanEval and Natural User Prompts 7 May 2024 · 1 repository · arXiv:2405.04520
-
AlphaMath Almost Zero: Process Supervision without Process 6 May 2024 · 1 repository · arXiv:2405.03553
-
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions 6 May 2024 · 1 repository · arXiv:2405.03205
-
GREEN: Generative Radiology Report Evaluation and Error Notation 6 May 2024 · 0 repositories · arXiv:2405.03595
-
MAmmoTH2: Scaling Instructions from the Web 6 May 2024 · 0 repositories · arXiv:2405.03548
-
Can Large Language Models Make the Grade? An Empirical Study Evaluating LLMs Ability to Mark Short Answer Questions in K-12 Education 5 May 2024 · 0 repositories · arXiv:2405.02985
-
Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation 5 May 2024 · 0 repositories · arXiv:2405.06681
-
NegativePrompt: Leveraging Psychology for Large Language Models Enhancement via Negative Emotional Stimuli 5 May 2024 · 1 repository · arXiv:2405.02814Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Open-SQL Framework: Enhancing Text-to-SQL on Open-source Large Language Models 4 May 2024 · 0 repositories · arXiv:2405.06674
-
PropertyGPT: LLM-driven Formal Verification of Smart Contracts through Retrieval-Augmented Property Generation 4 May 2024 · 1 repository · arXiv:2405.02580
-
Automating the Enterprise with Foundation Models 3 May 2024 · 1 repository · arXiv:2405.03710Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Comparative Analysis of Retrieval Systems in the Real World 3 May 2024 · 0 repositories · arXiv:2405.02048
-
Attribution in Scientific Literature: New Benchmark and Methods 3 May 2024 · 0 repositories · arXiv:2405.02228
-
Single and Multi-Hop Question-Answering Datasets for Reticular Chemistry with GPT-4-Turbo 3 May 2024 · 1 repository · arXiv:2405.02128
-
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law 2 May 2024 · 1 repository · arXiv:2405.01769
-
How Can I Get It Right? Using GPT to Rephrase Incorrect Trainee Responses 2 May 2024 · 0 repositories · arXiv:2405.00970
-
MiniGPT-3D: Efficiently Aligning 3D Point Clouds with Large Language Models using 2D Priors 2 May 2024 · 1 repository · arXiv:2405.01413
-
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models 2 May 2024 · 1 repository · arXiv:2405.01535
-
WildChat: 1M ChatGPT Interaction Logs in the Wild 2 May 2024 · 0 repositories · arXiv:2405.01470
-
CourseAssist: Pedagogically Appropriate AI Tutor for Computer Science Education 1 May 2024 · 0 repositories · arXiv:2407.10246
-
Self-Play Preference Optimization for Language Model Alignment 1 May 2024 · 1 repository · arXiv:2405.00675Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
A Framework for Leveraging Human Computation Gaming to Enhance Knowledge Graphs for Accuracy Critical Generative AI Applications 30 Apr 2024 · 0 repositories · arXiv:2404.19729
-
Constrained Decoding for Secure Code Generation 30 Apr 2024 · 2 repositories · arXiv:2405.00218Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
Do Large Language Models Understand Conversational Implicature -- A case study with a chinese sitcom 30 Apr 2024 · 1 repository · arXiv:2404.19509Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Extending Llama-3's Context Ten-Fold Overnight 30 Apr 2024 · 1 repository · arXiv:2404.19553
-
Harmonic LLMs are Trustworthy 30 Apr 2024 · 0 repositories · arXiv:2404.19708
-
Octopus v4: Graph of language models 30 Apr 2024 · 0 repositories · arXiv:2404.19296
-
RepEval: Effective Text Evaluation with LLM Representation 30 Apr 2024 · 1 repository · arXiv:2404.19563Syntology official: no sample here; runs from other or unrecorded repositories · 18 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 2 honoured, 0 violated, 13 with no contract checked; 3 where Syntology's instrument failed) · 9 unverified (of 27 harvested samples) · 6 pointer-only (licence)
-
Automated Construction of Theme-specific Knowledge Graphs 29 Apr 2024 · 0 repositories · arXiv:2404.19146
-
Can GPT-4 do L2 analytic assessment? 29 Apr 2024 · 0 repositories · arXiv:2404.18557
-
Capabilities of Gemini Models in Medicine 29 Apr 2024 · 0 repositories · arXiv:2404.18416
-
How secure is AI-generated Code: A Large-Scale Comparison of Large Language Models 29 Apr 2024 · 1 repository · arXiv:2404.18353
-
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in 29 Apr 2024 · 0 repositories · arXiv:2404.18460
-
GPT-4 passes most of the 297 written Polish Board Certification Examinations 29 Apr 2024 · 0 repositories · arXiv:2405.01589
-
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report 29 Apr 2024 · 1 repository · arXiv:2405.00732Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
ComposerX: Multi-Agent Symbolic Music Composition with LLMs 28 Apr 2024 · 1 repository · arXiv:2404.18081Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
PatentGPT: A Large Language Model for Intellectual Property 28 Apr 2024 · 0 repositories · arXiv:2404.18255
-
Advancing Healthcare Automation: Multi-Agent System for Medical Necessity Justification 27 Apr 2024 · 0 repositories · arXiv:2404.17977
-
Automating Customer Needs Analysis: A Comparative Study of Large Language Models in the Travel Industry 27 Apr 2024 · 0 repositories · arXiv:2404.17975
-
Detection of Conspiracy Theories Beyond Keyword Bias in German-Language Telegram Using Large Language Models 27 Apr 2024 · 0 repositories · arXiv:2404.17985
-
Enhancing Pre-Trained Generative Language Models with Question Attended Span Extraction on Machine Reading Comprehension 27 Apr 2024 · 0 repositories · arXiv:2404.17991
-
Evaluation of Few-Shot Learning for Classification Tasks in the Polish Language 27 Apr 2024 · 0 repositories · arXiv:2404.17832
-
Using LLMs in Software Requirements Specifications: An Empirical Evaluation 27 Apr 2024 · 0 repositories · arXiv:2404.17842
-
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation 25 Apr 2024 · 1 repository · arXiv:2405.00715
-
AI Coders Are Among Us: Rethinking Programming Language Grammar Towards Efficient Code Generation 25 Apr 2024 · 1 repository · arXiv:2404.16333
-
Global Concept Explanations for Graphs by Contrastive Learning 25 Apr 2024 · 2 repositories · arXiv:2404.16532
-
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages 25 Apr 2024 · 1 repository · arXiv:2404.16816
-
Influence of Solution Efficiency and Valence of Instruction on Additive and Subtractive Solution Strategies in Humans and GPT-4 25 Apr 2024 · 0 repositories · arXiv:2404.16692
-
LLM-Based Section Identifiers Excel on Open Source but Stumble in Real World Applications 25 Apr 2024 · 1 repository · arXiv:2404.16294
-
Player-Driven Emergence in LLM-Driven Game Narrative 25 Apr 2024 · 0 repositories · arXiv:2404.17027
-
The GPT Surprise: Offering Large Language Model Chat in a Massive Coding Class Reduced Engagement but Increased Adopters Exam Performances 25 Apr 2024 · 0 repositories · arXiv:2407.09975
-
Utilizing Large Language Models to Identify Reddit Users Considering Vaping Cessation for Digital Interventions 25 Apr 2024 · 0 repositories · arXiv:2404.17607
-
Assessing The Potential Of Mid-Sized Language Models For Clinical QA 24 Apr 2024 · 0 repositories · arXiv:2404.15894
-
Can Foundational Large Language Models Assist with Conducting Pharmaceuticals Manufacturing Investigations? 24 Apr 2024 · 0 repositories · arXiv:2404.15578
-
Prompt Leakage effect and defense strategies for multi-turn LLM interactions 24 Apr 2024 · 0 repositories · arXiv:2404.16251
-
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering 24 Apr 2024 · 1 repository · arXiv:2404.15655Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
The Promise and Challenges of Using LLMs to Accelerate the Screening Process of Systematic Reviews 24 Apr 2024 · 0 repositories · arXiv:2404.15667
-
Towards a Holistic Evaluation of LLMs on Factual Knowledge Recall 24 Apr 2024 · 0 repositories · arXiv:2404.16164
-
Aligning LLM Agents by Learning Latent Preference from User Edits 23 Apr 2024 · 1 repository · arXiv:2404.15269
-
ClinicalAgent: Clinical Trial Multi-Agent System with Large Language Model-based Reasoning 23 Apr 2024 · 0 repositories · arXiv:2404.14777
-
DesignProbe: A Graphic Design Benchmark for Multimodal Large Language Models 23 Apr 2024 · 0 repositories · arXiv:2404.14801
-
PRISM: Patient Records Interpretation for Semantic Clinical Trial Matching using Large Language Models 23 Apr 2024 · 0 repositories · arXiv:2404.15549
-
From Complexity to Clarity: How AI Enhances Perceptions of Scientists and the Public's Understanding of Science 23 Apr 2024 · 0 repositories · arXiv:2405.00706
-
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models 23 Apr 2024 · 1 repository · arXiv:2404.15522
-
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO 22 Apr 2024 · 1 repository · arXiv:2404.13957Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Information Re-Organization Improves Reasoning in Large Language Models 22 Apr 2024 · 0 repositories · arXiv:2404.13985
-
Navigating the Path of Writing: Outline-guided Text Generation with Large Language Models 22 Apr 2024 · 0 repositories · arXiv:2404.13919
-
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations 21 Apr 2024 · 1 repository · arXiv:2404.14459
-
SVGEditBench: A Benchmark Dataset for Quantitative Assessment of LLM's SVG Editing Capabilities 21 Apr 2024 · 1 repository · arXiv:2404.13710Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Beyond Accuracy: Investigating Error Types in GPT-4 Responses to USMLE Questions 20 Apr 2024 · 1 repository · arXiv:2404.13307
-
Large Language Models as Test Case Generators: Performance Evaluation and Enhancement 20 Apr 2024 · 1 repository · arXiv:2404.13340Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data 19 Apr 2024 · 1 repository · arXiv:2404.12933
-
CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models 19 Apr 2024 · 1 repository · arXiv:2404.13161Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Dubo-SQL: Diverse Retrieval-Augmented Generation and Fine Tuning for Text-to-SQL 19 Apr 2024 · 1 repository · arXiv:2404.12560
-
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence 18 Apr 2024 · 1 repository · arXiv:2404.11826Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models 18 Apr 2024 · 0 repositories · arXiv:2404.12494
-
CAUS: A Dataset for Question Generation based on Human Cognition Leveraging Large Language Models 18 Apr 2024 · 1 repository · arXiv:2404.11835
-
Concept Induction using LLMs: a user experiment for assessment 18 Apr 2024 · 0 repositories · arXiv:2404.11875
-
OpenBezoar: Small, Cost-Effective and Open Models Trained on Mixes of Instruction Data 18 Apr 2024 · 1 repository · arXiv:2404.12195
-
Uncovering Safety Risks of Large Language Models through Concept Activation Vector 18 Apr 2024 · 1 repository · arXiv:2404.12038Syntology official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 5 pointer-only (licence)