Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 2
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 2 of 29: papers 101 to 200 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Utilizing LLMs to Investigate the Disputed Role of Evidence in Electronic Cigarette Health Policy Formation in Australia and the UK 10 May 2025 · 0 repositories · arXiv:2505.06782
-
Camera Control at the Edge with Language Models for Scene Understanding 9 May 2025 · 0 repositories · arXiv:2505.06402
-
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information 9 May 2025 · 0 repositories · arXiv:2505.06046
-
AI Approaches to Qualitative and Quantitative News Analytics on NATO Unity 8 May 2025 · 0 repositories · arXiv:2505.06313
-
Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended Action Environments 8 May 2025 · 1 repository · arXiv:2505.05540
-
Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization 8 May 2025 · 0 repositories · arXiv:2505.05070
-
HiPerRAG: High-Performance Retrieval Augmented Generation for Scientific Insights 7 May 2025 · 0 repositories · arXiv:2505.04846
-
Red Teaming the Mind of the Machine: A Systematic Evaluation of Prompt Injection and Jailbreak Vulnerabilities in LLMs 7 May 2025 · 0 repositories · arXiv:2505.04806
-
LLM-OptiRA: LLM-Driven Optimization of Resource Allocation for Non-Convex Problems in Wireless Communications 4 May 2025 · 1 repository · arXiv:2505.02091
-
SEval-Ex: A Statement-Level Framework for Explainable Summarization Evaluation 4 May 2025 · 0 repositories · arXiv:2505.02235
-
Semantic Intelligence: Integrating GPT-4 with A Planning in Low-Cost Robotics 3 May 2025 · 0 repositories · arXiv:2505.01931
-
Enhancing SPARQL Query Rewriting for Complex Ontology Alignments 2 May 2025 · 0 repositories · arXiv:2505.01309
-
Good News for Script Kiddies? Evaluating Large Language Models for Automated Exploit Generation 2 May 2025 · 0 repositories · arXiv:2505.01065
-
Zero-Shot Document-Level Biomedical Relation Extraction via Scenario-based Prompt Design in Two-Stage with LLM 2 May 2025 · 0 repositories · arXiv:2505.01077
-
Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems 30 Apr 2025 · 0 repositories · arXiv:2505.00061
-
Geolocating Earth Imagery from ISS: Integrating Machine Learning with Astronaut Photography for Enhanced Geographic Mapping 29 Apr 2025 · 1 repository · arXiv:2504.21194
-
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry 29 Apr 2025 · 0 repositories · arXiv:2504.20849
-
Multimodal Large Language Models for Medicine: A Comprehensive Survey 29 Apr 2025 · 0 repositories · arXiv:2504.21051
-
YoChameleon: Personalized Vision and Language Generation 29 Apr 2025 · 0 repositories · arXiv:2504.20998
-
Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi 28 Apr 2025 · 0 repositories · arXiv:2504.20276
-
m-KAILIN: Knowledge-Driven Agentic Scientific Corpus Distillation Framework for Biomedical Large Language Models Training 28 Apr 2025 · 0 repositories · arXiv:2504.19565
-
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering 27 Apr 2025 · 1 repository · arXiv:2504.19384
-
Why you shouldn't fully trust ChatGPT: A synthesis of this AI tool's error rates across disciplines and the software engineering lifecycle 26 Apr 2025 · 0 repositories · arXiv:2504.18858
-
Advanced Chest X-Ray Analysis via Transformer-Based Image Descriptors and Cross-Model Attention Mechanism 23 Apr 2025 · 0 repositories · arXiv:2504.16774
-
Amplified Vulnerabilities: Structured Jailbreak Attacks on LLM-based Multi-Agent Debate 23 Apr 2025 · 0 repositories · arXiv:2504.16489
-
Emo Pillars: Knowledge Distillation to Support Fine-Grained Context-Aware and Context-Less Emotion Classification 23 Apr 2025 · 0 repositories · arXiv:2504.16856
-
Leveraging LLMs as Meta-Judges: A Multi-Agent Framework for Evaluating LLM Judgments 23 Apr 2025 · 0 repositories · arXiv:2504.17087
-
A Large-scale Class-level Benchmark Dataset for Code Generation with LLMs 22 Apr 2025 · 0 repositories · arXiv:2504.15564
-
LLMs as Data Annotators: How Close Are We to Human Performance 21 Apr 2025 · 0 repositories · arXiv:2504.15022
-
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents 20 Apr 2025 · 0 repositories · arXiv:2504.14650
-
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale 19 Apr 2025 · 1 repository · arXiv:2504.14225Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
SimplifyMyText: An LLM-Based System for Inclusive Plain Language Text Simplification 19 Apr 2025 · 0 repositories · arXiv:2504.14223
-
LLM Sensitivity Evaluation Framework for Clinical Diagnosis 18 Apr 2025 · 0 repositories · arXiv:2504.13475
-
Accuracy is Not Agreement: Expert-Aligned Evaluation of Crash Narrative Classification Models 17 Apr 2025 · 0 repositories · arXiv:2504.13068
-
Exploring Expert Failures Improves LLM Agent Tuning 17 Apr 2025 · 0 repositories · arXiv:2504.13145
-
Mitigating LLM Hallucinations with Knowledge Graphs: A Case Study 16 Apr 2025 · 0 repositories · arXiv:2504.12422
-
Exploring the Role of Knowledge Graph-Based RAG in Japanese Medical Question Answering with Small-Scale LLMs 15 Apr 2025 · 0 repositories · arXiv:2504.10982
-
Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models 14 Apr 2025 · 0 repositories · arXiv:2504.10615
-
Can LLMs handle WebShell detection? Overcoming Detection Challenges with Behavioral Function-Aware Framework 14 Apr 2025 · 0 repositories · arXiv:2504.13811
-
EMAFusion: A Self-Optimizing System for Seamless LLM Selection and Integration 14 Apr 2025 · 0 repositories · arXiv:2504.10681
-
Keyword Extraction, and Aspect Classification in Sinhala, English, and Code-Mixed Content 14 Apr 2025 · 0 repositories · arXiv:2504.10679
-
ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model 13 Apr 2025 · 1 repository · arXiv:2504.09421
-
Integrating Large Language Models for Automated Structural Analysis 13 Apr 2025 · 0 repositories · arXiv:2504.09754
-
Iterative Self-Training for Code Generation via Reinforced Re-Ranking 13 Apr 2025 · 0 repositories · arXiv:2504.09643
-
LLMTaxo: Leveraging Large Language Models for Constructing Taxonomy of Factual Claims from Social Media 11 Apr 2025 · 0 repositories · arXiv:2504.12325
-
RTLRepoCoder: Repository-Level RTL Code Completion through the Combination of Fine-Tuning and Retrieval Augmentation 11 Apr 2025 · 0 repositories · arXiv:2504.08862
-
AttentionDefense: Leveraging System Prompt Attention for Explainable Defense Against Novel Jailbreaks 10 Apr 2025 · 0 repositories · arXiv:2504.12321
-
Has the Creativity of Large-Language Models peaked? An analysis of inter- and intra-LLM variability 10 Apr 2025 · 0 repositories · arXiv:2504.12320
-
Revisiting Prompt Optimization with Large Reasoning Models-A Case Study on Event Extraction 10 Apr 2025 · 0 repositories · arXiv:2504.07357
-
Synthetic Fluency: Hallucinations, Confabulations, and the Creation of Irish Words in LLM-Generated Translations 10 Apr 2025 · 0 repositories · arXiv:2504.07680
-
Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation 9 Apr 2025 · 0 repositories · arXiv:2504.08806
-
Text-to-Image Models and Their Representation of People from Different Nationalities Engaging in Activities 8 Apr 2025 · 0 repositories · arXiv:2504.06313
-
Assessing how hyperparameters impact Large Language Models' sarcasm detection performance 8 Apr 2025 · 0 repositories · arXiv:2504.06166
-
InstructionBench: An Instructional Video Understanding Benchmark 7 Apr 2025 · 0 repositories · arXiv:2504.05040
-
Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent 7 Apr 2025 · 0 repositories · arXiv:2504.04702
-
CoLa -- Learning to Interactively Collaborate with Large LMs 3 Apr 2025 · 0 repositories · arXiv:2504.02965
-
LearNAT: Learning NL2SQL with AST-guided Task Decomposition for Large Language Models 3 Apr 2025 · 0 repositories · arXiv:2504.02327
-
Task as Context Prompting for Accurate Medical Symptom Coding Using Large Language Models 3 Apr 2025 · 1 repository · arXiv:2504.03051
-
PiCo: Jailbreaking Multimodal Large Language Models via Pictorial Code Contextualization 2 Apr 2025 · 0 repositories · arXiv:2504.01444
-
Automated Factual Benchmarking for In-Car Conversational Systems using Large Language Models 1 Apr 2025 · 0 repositories · arXiv:2504.01248
-
Collaborative LLM Numerical Reasoning with Local Data Protection 1 Apr 2025 · 0 repositories · arXiv:2504.00299
-
SRLCG: Self-Rectified Large-Scale Code Generation with Multidimensional Chain-of-Thought and Dynamic Backtracking 1 Apr 2025 · 0 repositories · arXiv:2504.00532
-
JudgeLRM: Large Reasoning Models as a Judge 31 Mar 2025 · 0 repositories · arXiv:2504.00050
-
Large Language Models Pass the Turing Test 31 Mar 2025 · 0 repositories · arXiv:2503.23674
-
LLM4FS: Leveraging Large Language Models for Feature Selection and How to Improve It 31 Mar 2025 · 0 repositories · arXiv:2503.24157
-
Exploring GPT-4 for Robotic Agent Strategy with Real-Time State Feedback and a Reactive Behaviour Framework 30 Mar 2025 · 0 repositories · arXiv:2503.23601
-
FeRG-LLM : Feature Engineering by Reason Generation Large Language Models 30 Mar 2025 · 0 repositories · arXiv:2503.23371
-
RARE: Retrieval-Augmented Reasoning Modeling 30 Mar 2025 · 1 repository · arXiv:2503.23513Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples)
-
A Training-free LLM Framework with Interaction between Contextually Related Subtasks in Solving Complex Tasks 29 Mar 2025 · 0 repositories · arXiv:2503.23053
-
Multimodal machine learning with large language embedding model for polymer property prediction 29 Mar 2025 · 1 repository · arXiv:2503.22962
-
How Well Can Vison-Language Models Understand Humans' Intention? An Open-ended Theory of Mind Question Evaluation Benchmark 28 Mar 2025 · 0 repositories · arXiv:2503.22093
-
Integrating Artificial Intelligence with Human Expertise: An In-depth Analysis of ChatGPT's Capabilities in Generating Metamorphic Relations 28 Mar 2025 · 0 repositories · arXiv:2503.22141
-
Understanding Inequality of LLM Fact-Checking over Geographic Regions with Agent and Retrieval models 28 Mar 2025 · 0 repositories · arXiv:2503.22877
-
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment 27 Mar 2025 · 0 repositories · arXiv:2503.21720
-
Using large language models to produce literature reviews: Usages and systematic biases of microphysics parametrizations in 2699 publications 27 Mar 2025 · 0 repositories · arXiv:2503.21352
-
Vision Language Models versus Machine Learning Models Performance on Polyp Detection and Classification in Colonoscopy Images 27 Mar 2025 · 1 repository · arXiv:2503.21840
-
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations 26 Mar 2025 · 0 repositories · arXiv:2503.20126
-
Iterative Prompting with Persuasion Skills in Jailbreaking Large Language Models 26 Mar 2025 · 0 repositories · arXiv:2503.20320
-
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications 25 Mar 2025 · 0 repositories · arXiv:2503.19276
-
Enabling Rapid Shared Human-AI Mental Model Alignment via the After-Action Review 25 Mar 2025 · 1 repository · arXiv:2503.19607
-
Fundamental Limits of Perfect Concept Erasure 25 Mar 2025 · 1 repository · arXiv:2503.20098
-
SCI-IDEA: Context-Aware Scientific Ideation Using Token and Sentence Embeddings 25 Mar 2025 · 0 repositories · arXiv:2503.19257
-
Taxonomy Inference for Tabular Data Using Large Language Models 25 Mar 2025 · 0 repositories · arXiv:2503.21810
-
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation 24 Mar 2025 · 1 repository · arXiv:2503.18476Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
TDRI: Two-Phase Dialogue Refinement and Co-Adaptation for Interactive Image Generation 22 Mar 2025 · 0 repositories · arXiv:2503.17669
-
Assessing the Reliability and Validity of GPT-4 in Annotating Emotion Appraisal Ratings 21 Mar 2025 · 0 repositories · arXiv:2503.16883
-
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization 21 Mar 2025 · 0 repositories · arXiv:2503.17136
-
SaudiCulture: A Benchmark for Evaluating Large Language Models Cultural Competence within Saudi Arabia 21 Mar 2025 · 0 repositories · arXiv:2503.17485
-
When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making 21 Mar 2025 · 0 repositories · arXiv:2503.16965Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement 20 Mar 2025 · 0 repositories · arXiv:2503.16024
-
ELTEX: A Framework for Domain-Driven Synthetic Data Generation 19 Mar 2025 · 1 repository · arXiv:2503.15055
-
TruthLens:A Training-Free Paradigm for DeepFake Detection 19 Mar 2025 · 0 repositories · arXiv:2503.15342
-
Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study 19 Mar 2025 · 1 repository · arXiv:2503.15579
-
Gricean Norms as a Basis for Effective Collaboration 18 Mar 2025 · 1 repository · arXiv:2503.14484
-
Large Language Models for Virtual Human Gesture Selection 18 Mar 2025 · 0 repositories · arXiv:2503.14408
-
PENCIL: Long Thoughts with Short Memory 18 Mar 2025 · 1 repository · arXiv:2503.14337Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Fragile Mastery: Are Domain-Specific Trade-Offs Undermining On-Device Language Models? 16 Mar 2025 · 0 repositories · arXiv:2503.22698
-
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks 15 Mar 2025 · 1 repository · arXiv:2504.03665
-
Maritime Mission Planning for Unmanned Surface Vessel using Large Language Model 15 Mar 2025 · 0 repositories · arXiv:2503.12065
-
Prompt Sentiment: The Catalyst for LLM Change 14 Mar 2025 · 0 repositories · arXiv:2503.13510