Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 10
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 10 of 29: papers 901 to 1,000 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter? 20 Jul 2024 · 1 repository · arXiv:2407.14790
-
TraveLLM: Could you plan my new public transit route in face of a network disruption? 20 Jul 2024 · 0 repositories · arXiv:2407.14926
-
Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field 19 Jul 2024 · 0 repositories · arXiv:2407.14076
-
HeCiX: Integrating Knowledge Graphs and Large Language Models for Biomedical Research 19 Jul 2024 · 0 repositories · arXiv:2407.14030
-
Improving Retrieval in Sponsored Search by Leveraging Query Context Signals 19 Jul 2024 · 0 repositories · arXiv:2407.14346
-
LAPIS: Language Model-Augmented Police Investigation System 19 Jul 2024 · 0 repositories · arXiv:2407.20248
-
LLMs left, right, and center: Assessing GPT's capabilities to label political bias from web domains 19 Jul 2024 · 0 repositories · arXiv:2407.14344
-
SQLfuse: Enhancing Text-to-SQL Performance through Comprehensive LLM Synergy 19 Jul 2024 · 0 repositories · arXiv:2407.14568
-
Can Open-Source LLMs Compete with Commercial Models? Exploring the Few-Shot Performance of Current GPT Models in Biomedical Tasks 18 Jul 2024 · 1 repository · arXiv:2407.13511
-
Autonomous self-evolving research on biomedical data: the DREAM paradigm 18 Jul 2024 · 0 repositories · arXiv:2407.13637
-
Is Sarcasm Detection A Step-by-Step Reasoning Process in Large Language Models? 17 Jul 2024 · 0 repositories · arXiv:2407.12725
-
Does Refusal Training in LLMs Generalize to the Past Tense? 16 Jul 2024 · 1 repository · arXiv:2407.11969Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ECoh: Turn-level Coherence Evaluation for Multilingual Dialogues 16 Jul 2024 · 1 repository · arXiv:2407.11660
-
Educational Personalized Learning Path Planning with Large Language Models 16 Jul 2024 · 0 repositories · arXiv:2407.11773
-
GPT Assisted Annotation of Rhetorical and Linguistic Features for Interpretable Propaganda Technique Detection in News Text 16 Jul 2024 · 0 repositories · arXiv:2407.11827
-
Large Language Models as Misleading Assistants in Conversation 16 Jul 2024 · 0 repositories · arXiv:2407.11789
-
LoFTI: Localization and Factuality Transfer to Indian Locales 16 Jul 2024 · 1 repository · arXiv:2407.11833
-
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models 16 Jul 2024 · 0 repositories · arXiv:2407.12877
-
Beyond Generative Artificial Intelligence: Roadmap for Natural Language Generation 15 Jul 2024 · 0 repositories · arXiv:2407.10554
-
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses 15 Jul 2024 · 0 repositories · arXiv:2407.10725
-
CodeV: Empowering LLMs with HDL Generation through Multi-Level Summarization 15 Jul 2024 · 0 repositories · arXiv:2407.10424
-
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation 15 Jul 2024 · 0 repositories · arXiv:2407.10817
-
Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis 15 Jul 2024 · 0 repositories · arXiv:2407.10899
-
Sibyl: Simple yet Effective Agent Framework for Complex Real-world Reasoning 15 Jul 2024 · 1 repository · arXiv:2407.10718Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ChatLogic: Integrating Logic Programming with Large Language Models for Multi-Step Reasoning 14 Jul 2024 · 1 repository · arXiv:2407.10162
-
OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling 13 Jul 2024 · 1 repository · arXiv:2407.09887Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Causality extraction from medical text using Large Language Models (LLMs) 13 Jul 2024 · 0 repositories · arXiv:2407.10020
-
Cohesive Conversations: Enhancing Authenticity in Multi-Agent Simulated Dialogues 13 Jul 2024 · 0 repositories · arXiv:2407.09897
-
LLM-Collaboration on Automatic Science Journalism for the General Audience 13 Jul 2024 · 1 repository · arXiv:2407.09756
-
Benchmarking Language Model Creativity: A Case Study on Code Generation 12 Jul 2024 · 1 repository · arXiv:2407.09007Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Leveraging large language models for nano synthesis mechanism explanation: solid foundations or mere conjectures? 12 Jul 2024 · 1 repository · arXiv:2407.08922
-
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training 12 Jul 2024 · 2 repositories · arXiv:2407.09121Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Self-Evolving GPT: A Lifelong Autonomous Experiential Learner 12 Jul 2024 · 0 repositories · arXiv:2407.08937
-
Self-Prompt Tuning: Enable Autonomous Role-Playing in LLMs 12 Jul 2024 · 0 repositories · arXiv:2407.08995
-
Show, Don't Tell: Evaluating Large Language Models Beyond Textual Understanding with ChildPlay 12 Jul 2024 · 1 repository · arXiv:2407.11068
-
TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models 12 Jul 2024 · 0 repositories · arXiv:2407.09424
-
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs 12 Jul 2024 · 0 repositories · arXiv:2407.09152
-
Converging Paradigms: The Synergy of Symbolic and Connectionist AI in LLM-Empowered Autonomous Agents 11 Jul 2024 · 0 repositories · arXiv:2407.08516
-
Fault Diagnosis in Power Grids with Large Language Model 11 Jul 2024 · 0 repositories · arXiv:2407.08836
-
GPT-4 is judged more human than humans in displaced and inverted Turing tests 11 Jul 2024 · 0 repositories · arXiv:2407.08853
-
GTA: A Benchmark for General Tool Agents 11 Jul 2024 · 1 repository · arXiv:2407.08713
-
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On 11 Jul 2024 · 0 repositories · arXiv:2407.08348
-
Arabic Automatic Story Generation with Large Language Models 10 Jul 2024 · 1 repository · arXiv:2407.07551
-
Evaluating Large Language Models with Grid-Based Game Competitions: An Extensible LLM Benchmark and Leaderboard 10 Jul 2024 · 1 repository · arXiv:2407.07796
-
LitSearch: A Retrieval Benchmark for Scientific Literature Search 10 Jul 2024 · 1 repository · arXiv:2407.18940Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
A Guide To Effectively Leveraging LLMs for Low-Resource Text Summarization: Data Augmentation and Semi-supervised Approaches 10 Jul 2024 · 0 repositories · arXiv:2407.07341
-
Probability of Differentiation Reveals Brittleness of Homogeneity Bias in GPT-4 10 Jul 2024 · 0 repositories · arXiv:2407.07329
-
Examining Long-Context Large Language Models for Environmental Review Document Comprehension 10 Jul 2024 · 0 repositories · arXiv:2407.07321
-
Teaching Transformers Causal Reasoning through Axiomatic Training 10 Jul 2024 · 0 repositories · arXiv:2407.07612
-
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment 10 Jul 2024 · 0 repositories · arXiv:2407.07778
-
Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis 9 Jul 2024 · 1 repository · arXiv:2407.12857Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples)
-
ConvNLP: Image-based AI Text Detection 9 Jul 2024 · 0 repositories · arXiv:2407.07225
-
PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods 9 Jul 2024 · 1 repository · arXiv:2407.06985
-
Prompting Techniques for Secure Code Generation: A Systematic Investigation 9 Jul 2024 · 0 repositories · arXiv:2407.07064
-
Source Code Summarization in the Era of Large Language Models 9 Jul 2024 · 1 repository · arXiv:2407.07959
-
Using Large Language Models for Generating Smart Contracts for Health Insurance from Textual Policies 9 Jul 2024 · 0 repositories · arXiv:2407.07019
-
CodeUpdateArena: Benchmarking Knowledge Editing on API Updates 8 Jul 2024 · 1 repository · arXiv:2407.06249
-
Generative Debunking of Climate Misinformation 8 Jul 2024 · 0 repositories · arXiv:2407.05599
-
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct 8 Jul 2024 · 1 repository · arXiv:2407.05700Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Controllable and Reliable Knowledge-Intensive Task-Oriented Conversational Agents with Declarative Genie Worksheets 8 Jul 2024 · 1 repository · arXiv:2407.05674
-
Meme Analysis using LLM-based Contextual Information and U-net Encapsulated Transformer 8 Jul 2024 · 1 repository
-
Potential of Multimodal Large Language Models for Data Mining of Medical Images and Free-text Reports 8 Jul 2024 · 0 repositories · arXiv:2407.05758
-
Surprising gender biases in GPT 8 Jul 2024 · 0 repositories · arXiv:2407.06003
-
T2VSafetyBench: Evaluating the Safety of Text-to-Video Generative Models 8 Jul 2024 · 0 repositories · arXiv:2407.05965
-
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop 8 Jul 2024 · 0 repositories · arXiv:2407.05925
-
Enhancing Computer Programming Education with LLMs: A Study on Effective Prompt Engineering for Python Code Generation 7 Jul 2024 · 0 repositories · arXiv:2407.05437
-
Large Language Model as an Assignment Evaluator: Insights, Feedback, and Challenges in a 1000+ Student Course 7 Jul 2024 · 0 repositories · arXiv:2407.05216
-
MINDECHO: Role-Playing Language Agents for Key Opinion Leaders 7 Jul 2024 · 0 repositories · arXiv:2407.05305
-
EVA-Score: Evaluating Abstractive Long-form Summarization on Informativeness through Extraction and Validation 6 Jul 2024 · 0 repositories · arXiv:2407.04969
-
How do you know that? Teaching Generative Language Models to Reference Answers to Biomedical Questions 6 Jul 2024 · 1 repository · arXiv:2407.05015
-
Solving for X and Beyond: Can Large Language Models Solve Complex Math Problems with More-Than-Two Unknowns? 6 Jul 2024 · 1 repository · arXiv:2407.05134
-
Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments 5 Jul 2024 · 0 repositories · arXiv:2407.12847
-
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models 5 Jul 2024 · 1 repository · arXiv:2407.04693Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Using LLMs to label medical papers according to the CIViC evidence model 5 Jul 2024 · 1 repository · arXiv:2407.04466
-
Diverse and Fine-Grained Instruction-Following Ability Exploration with Synthetic Data 4 Jul 2024 · 0 repositories · arXiv:2407.03942
-
Evaluating Language Model Context Windows: A "Working Memory" Test and Inference-time Correction 4 Jul 2024 · 1 repository · arXiv:2407.03651
-
GPT-4 vs. Human Translators: A Comprehensive Evaluation of Translation Quality Across Languages, Domains, and Expertise Levels 4 Jul 2024 · 0 repositories · arXiv:2407.03658
-
QET: Enhancing Quantized LLM Parameters and KV cache Compression through Element Substitution and Residual Clustering 4 Jul 2024 · 0 repositories · arXiv:2407.03637
-
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation 4 Jul 2024 · 0 repositories · arXiv:2407.03841
-
Query-Guided Self-Supervised Summarization of Nursing Notes 4 Jul 2024 · 0 repositories · arXiv:2407.04125
-
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems 4 Jul 2024 · 0 repositories · arXiv:2407.03956
-
Towards Automating Text Annotation: A Case Study on Semantic Proximity Annotation using GPT-4 4 Jul 2024 · 0 repositories · arXiv:2407.04130
-
LANE: Logic Alignment of Non-tuning Large Language Models and Online Recommendation Systems for Explainable Reason Generation 3 Jul 2024 · 0 repositories · arXiv:2407.02833
-
Large Language Models as Evaluators for Scientific Synthesis 3 Jul 2024 · 0 repositories · arXiv:2407.02977
-
Learning to Reduce: Towards Improving Performance of Large Language Models on Structured Data 3 Jul 2024 · 0 repositories · arXiv:2407.02750
-
On Large Language Models in National Security Applications 3 Jul 2024 · 1 repository · arXiv:2407.03453
-
SemioLLM: Assessing Large Language Models for Semiological Analysis in Epilepsy Research 3 Jul 2024 · 0 repositories · arXiv:2407.03004
-
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts 3 Jul 2024 · 1 repository · arXiv:2407.03203Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
Assessing the Code Clone Detection Capability of Large Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02402
-
Beyond Numeric Awards: In-Context Dueling Bandits with LLM Agents 2 Jul 2024 · 0 repositories · arXiv:2407.01887
-
Fake News Detection and Manipulation Reasoning via Large Vision-Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02042
-
Improving Visual Storytelling with Multimodal Large Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02586
-
Integrate the Essence and Eliminate the Dross: Fine-Grained Self-Consistency for Free-Form Language Generation 2 Jul 2024 · 1 repository · arXiv:2407.02056Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
LLM-Select: Feature Selection with Large Language Models 2 Jul 2024 · 0 repositories · arXiv:2407.02694
-
Open foundation models for Azerbaijani language 2 Jul 2024 · 0 repositories · arXiv:2407.02337
-
RankRAG: Unifying Context Ranking with Retrieval-Augmented Generation in LLMs 2 Jul 2024 · 0 repositories · arXiv:2407.02485
-
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters 2 Jul 2024 · 1 repository · arXiv:2407.01902Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Art of Saying No: Contextual Noncompliance in Language Models 2 Jul 2024 · 1 repository · arXiv:2407.12043Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning 1 Jul 2024 · 1 repository · arXiv:2407.01687Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Image-to-Text Logic Jailbreak: Your Imagination can Help You Do Anything 1 Jul 2024 · 0 repositories · arXiv:2407.02534