Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 20
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 20 of 29: papers 1,901 to 2,000 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Training microrobots to swim by a large language model 21 Jan 2024 · 0 repositories · arXiv:2402.00044
-
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models 20 Jan 2024 · 1 repository · arXiv:2401.12242Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Enhancing Large Language Models for Clinical Decision Support by Incorporating Clinical Practice Guidelines 20 Jan 2024 · 0 repositories · arXiv:2401.11120
-
Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA 20 Jan 2024 · 0 repositories · arXiv:2401.12998
-
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images 20 Jan 2024 · 1 repository · arXiv:2401.11170Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences 19 Jan 2024 · 1 repository · arXiv:2401.10529
-
Mining experimental data from Materials Science literature with Large Language Models: an evaluation study 19 Jan 2024 · 1 repository · arXiv:2401.11052
-
Beyond Traditional Benchmarks: Analyzing Behaviors of Open LLMs on Data-to-Text Generation 18 Jan 2024 · 0 repositories · arXiv:2401.10186
-
ChatQA: Surpassing GPT-4 on Conversational QA and RAG 18 Jan 2024 · 0 repositories · arXiv:2401.10225
-
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents 18 Jan 2024 · 1 repository · arXiv:2401.10019Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Self-Rewarding Language Models 18 Jan 2024 · 3 repositories · arXiv:2401.10020Syntology 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models 17 Jan 2024 · 0 repositories · arXiv:2401.09002
-
Augmenting Math Word Problems via Iterative Question Composing 17 Jan 2024 · 1 repository · arXiv:2401.09003Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Bridging Research and Readers: A Multi-Modal Automated Academic Papers Interpretation System 17 Jan 2024 · 1 repository · arXiv:2401.09150
-
COCO is "ALL'' You Need for Visual Instruction Fine-tuning 17 Jan 2024 · 0 repositories · arXiv:2401.08968
-
Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text 17 Jan 2024 · 1 repository · arXiv:2401.09407
-
From User Surveys to Telemetry-Driven AI Agents: Exploring the Potential of Personalized Productivity Solutions 17 Jan 2024 · 0 repositories · arXiv:2401.08960
-
Impact of Large Language Model Assistance on Patients Reading Clinical Notes: A Mixed-Methods Study 17 Jan 2024 · 0 repositories · arXiv:2401.09637
-
Evaluating LLMs' Mathematical and Coding Competency through Ontology-guided Interventions 17 Jan 2024 · 1 repository · arXiv:2401.09395Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering 16 Jan 2024 · 2 repositories · arXiv:2401.08500Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation 16 Jan 2024 · 1 repository · arXiv:2401.08417Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis 16 Jan 2024 · 1 repository · arXiv:2401.08508Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Forging Vision Foundation Models for Autonomous Driving: Challenges, Methodologies, and Opportunities 16 Jan 2024 · 1 repository · arXiv:2401.08045
-
MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline 16 Jan 2024 · 2 repositories · arXiv:2401.08190Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MMToM-QA: Multimodal Theory of Mind Question Answering 16 Jan 2024 · 1 repository · arXiv:2401.08743
-
RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture 16 Jan 2024 · 0 repositories · arXiv:2401.08406
-
RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning 16 Jan 2024 · 1 repository · arXiv:2401.08326Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 1 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples)
-
A Novel Approach for Automatic Program Repair using Round-Trip Translation with Large Language Models 15 Jan 2024 · 1 repository · arXiv:2401.07994
-
Consolidating Trees of Robotic Plans Generated Using Large Language Models to Improve Reliability 15 Jan 2024 · 0 repositories · arXiv:2401.07868
-
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding 15 Jan 2024 · 0 repositories · arXiv:2401.07572
-
MapGPT: Map-Guided Prompting with Adaptive Path Planning for Vision-and-Language Navigation 14 Jan 2024 · 0 repositories · arXiv:2401.07314
-
Streamlining the Selection Phase of Systematic Literature Reviews (SLRs) Using AI-Enabled GPT-4 Assistant API 14 Jan 2024 · 0 repositories · arXiv:2402.18582
-
A Novel Multi-Stage Prompting Approach for Language Agnostic MCQ Generation using GPT 13 Jan 2024 · 1 repository · arXiv:2401.07098
-
Assessing Large Language Models in Mechanical Engineering Education: A Study on Mechanics-Focused Conceptual Understanding 13 Jan 2024 · 0 repositories · arXiv:2401.12983
-
Knowledge Distillation of Black-Box Large Language Models 13 Jan 2024 · 0 repositories · arXiv:2401.07013
-
A Survey on the Applications of Frontier AI, Foundation Models, and Large Language Models to Intelligent Transportation Systems 12 Jan 2024 · 0 repositories · arXiv:2401.06831
-
Adapting Large Language Models for Document-Level Machine Translation 12 Jan 2024 · 0 repositories · arXiv:2401.06468
-
Comparing GPT-4 and Open-Source Language Models in Misinformation Mitigation 12 Jan 2024 · 0 repositories · arXiv:2401.06920
-
Few-Shot Detection of Machine-Generated Text using Style Representations 12 Jan 2024 · 1 repository · arXiv:2401.06712Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Fine-grained Hallucination Detection and Editing for Language Models 12 Jan 2024 · 0 repositories · arXiv:2401.06855
-
Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs 12 Jan 2024 · 1 repository · arXiv:2401.06431
-
Health-LLM: Large Language Models for Health Prediction via Wearable Sensor Data 12 Jan 2024 · 1 repository · arXiv:2401.06866Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs 12 Jan 2024 · 2 repositories · arXiv:2401.06373Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
PizzaCommonSense: Learning to Model Commonsense Reasoning about Intermediate Steps in Cooking Recipes 12 Jan 2024 · 1 repository · arXiv:2401.06930
-
Autocompletion of Chief Complaints in the Electronic Health Records using Large Language Models 11 Jan 2024 · 0 repositories · arXiv:2401.06088
-
Evidence to Generate (E2G): A Single-agent Two-step Prompting for Context Grounded and Retrieval Augmented Reasoning 11 Jan 2024 · 0 repositories · arXiv:2401.05787
-
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs 11 Jan 2024 · 0 repositories · arXiv:2401.05940
-
The Benefits of a Concise Chain of Thought on Problem-Solving in Large Language Models 11 Jan 2024 · 1 repository · arXiv:2401.05618Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
CADgpt: Harnessing Natural Language Processing for 3D Modelling to Enhance Computer-Aided Design Workflows 10 Jan 2024 · 0 repositories · arXiv:2401.05476
-
Can AI Write Classical Chinese Poetry like Humans? An Empirical Study Inspired by Turing Test 10 Jan 2024 · 0 repositories · arXiv:2401.04952
-
Knowledge Sharing in Manufacturing using Large Language Models: User Evaluation and Model Benchmarking 10 Jan 2024 · 0 repositories · arXiv:2401.05200
-
Leveraging Print Debugging to Improve Code Generation in Large Language Models 10 Jan 2024 · 0 repositories · arXiv:2401.05319
-
Reinforcement Learning for Optimizing RAG for Domain Chatbots 10 Jan 2024 · 0 repositories · arXiv:2401.06800
-
Arabic Text Diacritization In The Age Of Transfer Learning: Token Classification Is All You Need 9 Jan 2024 · 0 repositories · arXiv:2401.04848
-
DebugBench: Evaluating Debugging Capability of Large Language Models 9 Jan 2024 · 1 repository · arXiv:2401.04621Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples)
-
Informed AI Regulation: Comparing the Ethical Frameworks of Leading LLM Chatbots Using an Ethics-Based Audit to Assess Moral Reasoning and Normative Values 9 Jan 2024 · 1 repository · arXiv:2402.01651
-
A Philosophical Introduction to Language Models -- Part I: Continuity With Classic Debates 8 Jan 2024 · 0 repositories · arXiv:2401.03910
-
Can Large Language Models Beat Wall Street? Unveiling the Potential of AI in Stock Selection 8 Jan 2024 · 0 repositories · arXiv:2401.03737
-
Distortions in Judged Spatial Relations in Large Language Models 8 Jan 2024 · 0 repositories · arXiv:2401.04218
-
LLM4PLC: Harnessing Large Language Models for Verifiable Programming of PLCs in Industrial Control Systems 8 Jan 2024 · 1 repository · arXiv:2401.05443
-
MARG: Multi-Agent Review Generation for Scientific Papers 8 Jan 2024 · 1 repository · arXiv:2401.04259Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Why Solving Multi-agent Path Finding with Large Language Model has not Succeeded Yet 8 Jan 2024 · 0 repositories · arXiv:2401.03630
-
Can generative AI and ChatGPT outperform humans on cognitive-demanding problem-solving tasks in science? 7 Jan 2024 · 0 repositories · arXiv:2401.15081
-
Escalation Risks from Language Models in Military and Diplomatic Decision-Making 7 Jan 2024 · 1 repository · arXiv:2401.03408Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
InFoBench: Evaluating Instruction Following Ability in Large Language Models 7 Jan 2024 · 1 repository · arXiv:2401.03601Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On Leveraging Large Language Models for Enhancing Entity Resolution: A Cost-efficient Approach 7 Jan 2024 · 0 repositories · arXiv:2401.03426
-
CharPoet: A Chinese Classical Poetry Generation System Based on Token-free LLM 7 Jan 2024 · 0 repositories · arXiv:2401.03512
-
Using Large Language Models to Assess Tutors' Performance in Reacting to Students Making Math Errors 6 Jan 2024 · 0 repositories · arXiv:2401.03238
-
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution 5 Jan 2024 · 1 repository · arXiv:2401.03065
-
From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models 5 Jan 2024 · 0 repositories · arXiv:2401.02777
-
Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language Models 5 Jan 2024 · 0 repositories · arXiv:2401.02851
-
PeFoMed: Parameter Efficient Fine-tuning of Multimodal Large Language Models for Medical Imaging 5 Jan 2024 · 1 repository · arXiv:2401.02797
-
Blar-SQL: Faster, Stronger, Smaller NL2SQL 4 Jan 2024 · 0 repositories · arXiv:2401.02997
-
Exploring Boundary of GPT-4V on Marine Analysis: A Preliminary Case Study 4 Jan 2024 · 0 repositories · arXiv:2401.02147
-
AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI 3 Jan 2024 · 2 repositories · arXiv:2401.01651Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
AstroLLaMA-Chat: Scaling AstroLLaMA with Conversational and Diverse Datasets 3 Jan 2024 · 0 repositories · arXiv:2401.01916
-
GPT-4V(ision) is a Generalist Web Agent, if Grounded 3 Jan 2024 · 1 repository · arXiv:2401.01614Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication 3 Jan 2024 · 1 repository · arXiv:2401.01620
-
Team IELAB at TREC Clinical Trial Track 2023: Enhancing Clinical Trial Retrieval with Neural Rankers and Large Language Models 3 Jan 2024 · 0 repositories · arXiv:2401.01566
-
CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation 2 Jan 2024 · 1 repository · arXiv:2401.01275Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Evaluating Large Language Models on the GMAT: Implications for the Future of Business Education 2 Jan 2024 · 0 repositories · arXiv:2401.02985
-
Identification of Regulatory Requirements Relevant to Business Processes: A Comparative Study on Generative AI, Embedding-based Ranking, Crowd and Expert-driven Methods 2 Jan 2024 · 0 repositories · arXiv:2401.02986
-
Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models 2 Jan 2024 · 2 repositories · arXiv:2401.01335Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Uncertainty Resolution in Misinformation Detection 2 Jan 2024 · 0 repositories · arXiv:2401.01197
-
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models 1 Jan 2024 · 1 repository · arXiv:2401.00757Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning 1 Jan 2024 · 0 repositories
-
Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education 1 Jan 2024 · 0 repositories · arXiv:2401.00832
-
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models 31 Dec 2023 · 3 repositories · arXiv:2401.00396Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Red Teaming for Large Language Models At Scale: Tackling Hallucinations on Mathematics Tasks 30 Dec 2023 · 1 repository · arXiv:2401.00290
-
Efficacy of Utilizing Large Language Models to Detect Public Threat Posted Online 29 Dec 2023 · 0 repositories · arXiv:2401.02974
-
Enhancing Quantitative Reasoning Skills of Large Language Models through Dimension Perception 29 Dec 2023 · 0 repositories · arXiv:2312.17532
-
MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation 28 Dec 2023 · 2 repositories · arXiv:2312.17080
-
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub 28 Dec 2023 · 1 repository · arXiv:2312.17294
-
RoleEval: A Bilingual Role Evaluation Benchmark for Large Language Models 26 Dec 2023 · 1 repository · arXiv:2312.16132
-
SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security 26 Dec 2023 · 1 repository · arXiv:2312.15838Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
ESGReveal: An LLM-based approach for extracting structured data from ESG reports 25 Dec 2023 · 0 repositories · arXiv:2312.17264
-
IQAGPT: Image Quality Assessment with Vision-language and ChatGPT Models 25 Dec 2023 · 0 repositories · arXiv:2312.15663
-
DEAP: Design Space Exploration for DNN Accelerator Parallelism 24 Dec 2023 · 0 repositories · arXiv:2312.15388
-
Do LLM Agents Exhibit Social Behavior? 23 Dec 2023 · 0 repositories · arXiv:2312.15198
-
Personalized Large Language Model Assistant with Evolving Conditional Memory 22 Dec 2023 · 0 repositories · arXiv:2312.17257