Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 5
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 5 of 29: papers 401 to 500 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Rethinking Emotion Annotations in the Era of Large Language Models 10 Dec 2024 · 0 repositories · arXiv:2412.07906
-
Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models 10 Dec 2024 · 2 repositories · arXiv:2412.07214
-
Anchoring Bias in Large Language Models: An Experimental Study 9 Dec 2024 · 0 repositories · arXiv:2412.06593
-
Exploring Memorization and Copyright Violation in Frontier LLMs: A Study of the New York Times v. OpenAI 2023 Lawsuit 9 Dec 2024 · 0 repositories · arXiv:2412.06370
-
Optimizing Multi-Task Learning for Enhanced Performance in Large Language Models 9 Dec 2024 · 0 repositories · arXiv:2412.06249
-
Evaluating Robustness of LLMs on Crisis-Related Microblogs across Events, Information Types, and Linguistic Features 8 Dec 2024 · 0 repositories · arXiv:2412.10413
-
Fully Open Source Moxin-7B Technical Report 8 Dec 2024 · 1 repository · arXiv:2412.06845
-
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor 8 Dec 2024 · 1 repository · arXiv:2412.07801
-
M³-20M: A Large-Scale Multi-Modal Molecule Dataset for AI-driven Drug Design and Discovery 8 Dec 2024 · 1 repository · arXiv:2412.06847
-
Can the Rookies Cut the Tough Cookie? Exploring the Use of LLMs for SQL Equivalence Checking 7 Dec 2024 · 0 repositories · arXiv:2412.05561
-
Innovative Sentiment Analysis and Prediction of Stock Price Using FinBERT, GPT-4 and Logistic Regression: A Data-Driven Approach 7 Dec 2024 · 0 repositories · arXiv:2412.06837
-
Shifting NER into High Gear: The Auto-AdvER Approach 7 Dec 2024 · 0 repositories · arXiv:2412.05655
-
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning 7 Dec 2024 · 2 repositories · arXiv:2412.05586Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
100% Elimination of Hallucinations on RAGTruth for GPT-4 and GPT-3.5 Turbo 6 Dec 2024 · 0 repositories · arXiv:2412.05223
-
Are Frontier Large Language Models Suitable for Q&A in Science Centres? 6 Dec 2024 · 0 repositories · arXiv:2412.05200
-
Enhancing LLMs for Impression Generation in Radiology Reports through a Multi-Agent System 6 Dec 2024 · 0 repositories · arXiv:2412.06828
-
Addressing Hallucinations with RAG and NMISS in Italian Healthcare LLM Chatbots 5 Dec 2024 · 0 repositories · arXiv:2412.04235
-
How Good is ChatGPT in Giving Adaptive Guidance Using Knowledge Graphs in E-Learning Environments? 5 Dec 2024 · 0 repositories · arXiv:2412.03856
-
A Water Efficiency Dataset for African Data Centers 4 Dec 2024 · 0 repositories · arXiv:2412.03716
-
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts? 4 Dec 2024 · 0 repositories · arXiv:2412.03235
-
Leveraging Large Language Models for Comparative Literature Summarization with Reflective Incremental Mechanisms 3 Dec 2024 · 0 repositories · arXiv:2412.02149
-
Patent-CR: A Dataset for Patent Claim Revision 3 Dec 2024 · 0 repositories · arXiv:2412.02549
-
RARE: Retrieval-Augmented Reasoning Enhancement for Large Language Models 3 Dec 2024 · 1 repository · arXiv:2412.02830
-
Automated Extraction of Acronym-Expansion Pairs from Scientific Papers 2 Dec 2024 · 0 repositories · arXiv:2412.01093
-
NYT-Connections: A Deceptively Simple Text Classification Task that Stumps System-1 Thinkers 2 Dec 2024 · 0 repositories · arXiv:2412.01621
-
PKRD-CoT: A Unified Chain-of-thought Prompting for Multi-Modal Large Language Models in Autonomous Driving 2 Dec 2024 · 0 repositories · arXiv:2412.02025
-
R-Bot: An LLM-based Query Rewrite System 2 Dec 2024 · 0 repositories · arXiv:2412.01661
-
The Promise and Peril of Generative AI: Evidence from GPT-4 as Sell-Side Analysts 2 Dec 2024 · 0 repositories · arXiv:2412.01069
-
Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments 30 Nov 2024 · 0 repositories · arXiv:2412.00323
-
On Domain-Specific Post-Training for Multimodal Large Language Models 29 Nov 2024 · 0 repositories · arXiv:2411.19930
-
Training Agents with Weakly Supervised Feedback from Large Language Models 29 Nov 2024 · 0 repositories · arXiv:2411.19547
-
A Lean Dataset for International Math Olympiad: Small Steps towards Writing Math Proofs for Hard Problems 28 Nov 2024 · 0 repositories · arXiv:2411.18872
-
MAG-V: A Multi-Agent Framework for Synthetic Data Generation and Verification 28 Nov 2024 · 0 repositories · arXiv:2412.04494
-
MATATA: Weakly Supervised End-to-End MAthematical Tool-Augmented Reasoning for Tabular Applications 28 Nov 2024 · 0 repositories · arXiv:2411.18915
-
SmartLLMSentry: A Comprehensive LLM Based Smart Contract Vulnerability Detection Framework 28 Nov 2024 · 0 repositories · arXiv:2411.19234
-
The Impact of Example Selection in Few-Shot Prompting on Automated Essay Scoring Using GPT Models 28 Nov 2024 · 0 repositories · arXiv:2411.18924
-
Dspy-based Neural-Symbolic Pipeline to Enhance Spatial Reasoning in LLMs 27 Nov 2024 · 0 repositories · arXiv:2411.18564
-
Aligning Knowledge Concepts to Whole Slide Images for Precise Histopathology Image Analysis 27 Nov 2024 · 1 repository · arXiv:2411.18101
-
The importance of visual modelling languages in generative software engineering 27 Nov 2024 · 1 repository · arXiv:2411.17976
-
Training and Evaluating Language Models with Template-based Data Generation 27 Nov 2024 · 1 repository · arXiv:2411.18104Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Can artificial intelligence predict clinical trial outcomes? 26 Nov 2024 · 0 repositories · arXiv:2411.17595
-
ER2Score: LLM-based Explainable and Customizable Metric for Assessing Radiology Reports with Reward-Control Loss 26 Nov 2024 · 0 repositories · arXiv:2411.17301
-
"Give me the code" -- Log Analysis of First-Year CS Students' Interactions With GPT 26 Nov 2024 · 0 repositories · arXiv:2411.17855
-
Leveraging Large Language Models and Topic Modeling for Toxicity Classification 26 Nov 2024 · 1 repository · arXiv:2411.17876
-
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation 26 Nov 2024 · 2 repositories · arXiv:2411.17945Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
Push the Limit of Multi-modal Emotion Recognition by Prompting LLMs with Receptive-Field-Aware Attention Weighting 26 Nov 2024 · 0 repositories · arXiv:2411.17674
-
Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring 25 Nov 2024 · 0 repositories · arXiv:2411.16337
-
CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning 25 Nov 2024 · 0 repositories · arXiv:2411.16313
-
Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models 25 Nov 2024 · 0 repositories · arXiv:2411.16797
-
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown 24 Nov 2024 · 0 repositories · arXiv:2411.15993
-
Don't Mesh with Me: Generating Constructive Solid Geometry Instead of Meshes by Fine-Tuning a Code-Generation LLM 22 Nov 2024 · 0 repositories · arXiv:2411.15279
-
Purrfessor: A Fine-tuned Multimodal LLaVA Diet Health Chatbot 22 Nov 2024 · 0 repositories · arXiv:2411.14925
-
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data 22 Nov 2024 · 1 repository · arXiv:2411.15004Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels 21 Nov 2024 · 1 repository · arXiv:2411.13775Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Explaining GPT-4's Schema of Depression Using Machine Behavior Analysis 21 Nov 2024 · 0 repositories · arXiv:2411.13800
-
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI 21 Nov 2024 · 1 repository · arXiv:2411.14522Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly 21 Nov 2024 · 0 repositories · arXiv:2411.14121
-
Understanding World or Predicting Future? A Comprehensive Survey of World Models 21 Nov 2024 · 0 repositories · arXiv:2411.14499
-
BIPro: Zero-shot Chinese Poem Generation via Block Inverse Prompting Constrained Generation Framework 20 Nov 2024 · 0 repositories · arXiv:2411.13237
-
Exploring Large Language Models for Climate Forecasting 20 Nov 2024 · 0 repositories · arXiv:2411.13724
-
The Impossible Test: A 2024 Unsolvable Dataset and A Chance for an AGI Quiz 20 Nov 2024 · 0 repositories · arXiv:2411.14486
-
Evaluating Tokenizer Performance of Large Language Models Across Official Indian Languages 19 Nov 2024 · 0 repositories · arXiv:2411.12240
-
Chapter 7 Review of Data-Driven Generative AI Models for Knowledge Extraction from Scientific Literature in Healthcare 18 Nov 2024 · 0 repositories · arXiv:2411.11635
-
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback 18 Nov 2024 · 1 repository · arXiv:2412.03578Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Popular LLMs Amplify Race and Gender Disparities in Human Mobility 18 Nov 2024 · 0 repositories · arXiv:2411.14469
-
IntentGPT: Few-shot Intent Discovery with Large Language Models 16 Nov 2024 · 0 repositories · arXiv:2411.10670
-
Does Prompt Formatting Have Any Impact on LLM Performance? 15 Nov 2024 · 0 repositories · arXiv:2411.10541
-
LoRA-LiteE: A Computationally Efficient Framework for Chatbot Preference-Tuning 15 Nov 2024 · 0 repositories · arXiv:2411.09947
-
Adopting RAG for LLM-Aided Future Vehicle Design 14 Nov 2024 · 0 repositories · arXiv:2411.09590
-
Automating Autograding: Large Language Models as Test Suite Generators for Introductory Programming 14 Nov 2024 · 0 repositories · arXiv:2411.09261
-
Evaluating Gender Bias in Large Language Models 14 Nov 2024 · 0 repositories · arXiv:2411.09826
-
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs 14 Nov 2024 · 1 repository · arXiv:2411.09492
-
Continuous GNN-based Anomaly Detection on Edge using Efficient Adaptive Knowledge Graph Learning 13 Nov 2024 · 0 repositories · arXiv:2411.09072
-
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks 12 Nov 2024 · 0 repositories · arXiv:2411.07464
-
Evaluating ChatGPT-3.5 Efficiency in Solving Coding Problems of Different Complexity Levels: An Empirical Analysis 12 Nov 2024 · 1 repository · arXiv:2411.07529
-
Improving Grapheme-to-Phoneme Conversion through In-Context Knowledge Retrieval with Large Language Models 12 Nov 2024 · 0 repositories · arXiv:2411.07563
-
Large Language Models Can Self-Improve in Long-context Reasoning 12 Nov 2024 · 1 repository · arXiv:2411.08147Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Verbosity ≠ Veracity: Demystify Verbosity Compensation Behavior of Large Language Models 12 Nov 2024 · 1 repository · arXiv:2411.07858
-
StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification 11 Nov 2024 · 1 repository · arXiv:2411.07076
-
AI's Spatial Intelligence: Evaluating AI's Understanding of Spatial Transformations in PSVT:R and Augmented Reality 9 Nov 2024 · 0 repositories · arXiv:2411.06269
-
Using Language Models to Disambiguate Lexical Choices in Translation 8 Nov 2024 · 1 repository · arXiv:2411.05781
-
HourVideo: 1-Hour Video-Language Understanding 7 Nov 2024 · 1 repository · arXiv:2411.04998Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
LLM2CLIP: Powerful Language Model Unlocks Richer Visual Representation 7 Nov 2024 · 1 repository · arXiv:2411.04997
-
Measuring short-form factuality in large language models 7 Nov 2024 · 1 repository · arXiv:2411.04368
-
A Comparative Study of Recent Large Language Models on Generating Hospital Discharge Summaries for Lung Cancer Patients 6 Nov 2024 · 0 repositories · arXiv:2411.03805
-
Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning 6 Nov 2024 · 1 repository · arXiv:2411.03978Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Diversity Helps Jailbreak Large Language Models 6 Nov 2024 · 0 repositories · arXiv:2411.04223
-
From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond 6 Nov 2024 · 0 repositories · arXiv:2411.03590
-
Uncertainty Quantification for Clinical Outcome Predictions with (Large) Language Models 5 Nov 2024 · 0 repositories · arXiv:2411.03497
-
Receiver-Centric Generative Semantic Communications 5 Nov 2024 · 0 repositories · arXiv:2411.03127
-
VERITAS: A Unified Approach to Reliability Evaluation 5 Nov 2024 · 0 repositories · arXiv:2411.03300
-
Disrupting Test Development with AI Assistants 4 Nov 2024 · 0 repositories · arXiv:2411.02328
-
Seq-VCR: Preventing Collapse in Intermediate Transformer Representations for Enhanced Reasoning 4 Nov 2024 · 1 repository · arXiv:2411.02344
-
Towards Leveraging News Media to Support Impact Assessment of AI Technologies 4 Nov 2024 · 0 repositories · arXiv:2411.02536
-
A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why? 3 Nov 2024 · 0 repositories · arXiv:2411.01414
-
High-performance automated abstract screening with large language model ensembles 3 Nov 2024 · 0 repositories · arXiv:2411.02451
-
Integration of Large Vision Language Models for Efficient Post-disaster Damage Assessment and Reporting 3 Nov 2024 · 0 repositories · arXiv:2411.01511
-
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models 3 Nov 2024 · 0 repositories · arXiv:2411.01703
-
Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems 2 Nov 2024 · 0 repositories · arXiv:2411.01173
-
Evaluating the Impact of Lab Test Results on Large Language Models Generated Differential Diagnoses from Clinical Case Vignettes 1 Nov 2024 · 0 repositories · arXiv:2411.02523