Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 19
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 19 of 29: papers 1,801 to 1,900 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LLMs Among Us: Generative AI Participating in Digital Discourse 8 Feb 2024 · 0 repositories · arXiv:2402.07940
-
Noise Contrastive Alignment of Language Models with Explicit Rewards 8 Feb 2024 · 3 repositories · arXiv:2402.05369Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation 8 Feb 2024 · 0 repositories · arXiv:2402.05699
-
TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation 8 Feb 2024 · 0 repositories · arXiv:2402.05733
-
Zero-Shot Chain-of-Thought Reasoning Guided by Evolutionary Algorithms in Large Language Models 8 Feb 2024 · 0 repositories · arXiv:2402.05376
-
Can Large Language Model Agents Simulate Human Trust Behavior? 7 Feb 2024 · 1 repository · arXiv:2402.04559Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Conversational Assistants in Knowledge-Intensive Contexts: An Evaluation of LLM- versus Intent-based Systems 7 Feb 2024 · 0 repositories · arXiv:2402.04955
-
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach 7 Feb 2024 · 0 repositories · arXiv:2402.04609
-
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning 7 Feb 2024 · 1 repository · arXiv:2402.04833Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Navigating the Knowledge Sea: Planet-scale answer retrieval using LLMs 7 Feb 2024 · 0 repositories · arXiv:2402.05318
-
Opening the AI black box: program synthesis via mechanistic interpretability 7 Feb 2024 · 1 repository · arXiv:2402.05110Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
TransLLaMa: LLM-based Simultaneous Translation System 7 Feb 2024 · 1 repository · arXiv:2402.04636
-
AnyTool: Self-Reflective, Hierarchical Agents for Large-Scale API Calls 6 Feb 2024 · 1 repository · arXiv:2402.04253Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Behind the Screen: Investigating ChatGPT's Dark Personality Traits and Conspiracy Beliefs 6 Feb 2024 · 0 repositories · arXiv:2402.04110
-
Comparing Abstraction in Humans and Large Language Models Using Multimodal Serial Reproduction 6 Feb 2024 · 0 repositories · arXiv:2402.03618
-
Identifying Reasons for Contraceptive Switching from Real-World Data Using Large Language Models 6 Feb 2024 · 1 repository · arXiv:2402.03597
-
Iterative Prompt Refinement for Radiation Oncology Symptom Extraction Using Teacher-Student Large Language Models 6 Feb 2024 · 0 repositories · arXiv:2402.04075
-
Large Language Models As MOOCs Graders 6 Feb 2024 · 0 repositories · arXiv:2402.03776
-
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs 6 Feb 2024 · 0 repositories · arXiv:2402.03927
-
LLM Agents can Autonomously Hack Websites 6 Feb 2024 · 0 repositories · arXiv:2402.06664
-
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification 6 Feb 2024 · 0 repositories · arXiv:2402.03686
-
Professional Agents -- Evolving Large Language Models into Autonomous Experts with Human-Level Competencies 6 Feb 2024 · 0 repositories · arXiv:2402.03628
-
Self-Discover: Large Language Models Self-Compose Reasoning Structures 6 Feb 2024 · 3 repositories · arXiv:2402.03620Syntology 0 ran · 3 unverified (of 3 harvested samples)
-
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models 5 Feb 2024 · 1 repository · arXiv:2402.02987
-
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 Feb 2024 · 5 repositories · arXiv:2402.03300Syntology official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (of 24 harvested samples) · 3 pointer-only (licence)
-
Graph-enhanced Large Language Models in Asynchronous Plan Reasoning 5 Feb 2024 · 1 repository · arXiv:2402.02805Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Harnessing PubMed User Query Logs for Post Hoc Explanations of Recommended Similar Articles 5 Feb 2024 · 0 repositories · arXiv:2402.03484
-
Is Mamba Capable of In-Context Learning? 5 Feb 2024 · 1 repository · arXiv:2402.03170Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
SWAG: Storytelling With Action Guidance 5 Feb 2024 · 1 repository · arXiv:2402.03483
-
Aligner: Efficient Alignment by Learning to Correct 4 Feb 2024 · 0 repositories · arXiv:2402.02416
-
Evaluating Large Language Models in Analysing Classroom Dialogue 4 Feb 2024 · 0 repositories · arXiv:2402.02380
-
Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation 4 Feb 2024 · 0 repositories · arXiv:2402.14594
-
BetterV: Controlled Verilog Generation with Discriminative Guidance 3 Feb 2024 · 0 repositories · arXiv:2402.03375
-
Do Moral Judgment and Reasoning Capability of LLMs Change with Language? A Study using the Multilingual Defining Issues Test 3 Feb 2024 · 0 repositories · arXiv:2402.02135
-
EffiBench: Benchmarking the Efficiency of Automatically Generated Code 3 Feb 2024 · 1 repository · arXiv:2402.02037Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
How well do LLMs cite relevant medical references? An evaluation framework and analyses 3 Feb 2024 · 1 repository · arXiv:2402.02008
-
How Can Generative AI Enhance the Well-being of Blind? 2 Feb 2024 · 0 repositories · arXiv:2402.07919
-
Integrating Large Language Models in Causal Discovery: A Statistical Causal Approach 2 Feb 2024 · 2 repositories · arXiv:2402.01454
-
TravelPlanner: A Benchmark for Real-World Planning with Language Agents 2 Feb 2024 · 2 repositories · arXiv:2402.01622Syntology official (archive's flag): 13 ran · 17 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 17 harvested samples) · 4 pointer-only (licence)
-
Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language Model 1 Feb 2024 · 0 repositories · arXiv:2402.01051
-
Hierarchical Multi-Label Classification of Online Vaccine Concerns 1 Feb 2024 · 0 repositories · arXiv:2402.01783
-
Ocassionally Secure: A Comparative Analysis of Code Generation Assistants 1 Feb 2024 · 0 repositories · arXiv:2402.00689
-
On the Psychology of GPT-4: Moderately anxious, slightly masculine, honest, and humble 1 Feb 2024 · 0 repositories · arXiv:2402.01777
-
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM 31 Jan 2024 · 0 repositories · arXiv:2402.00097
-
Global-Liar: Factuality of LLMs over Time and Geographic Regions 31 Jan 2024 · 0 repositories · arXiv:2401.17839
-
LLM Voting: Human Choices and AI Collective Decision Making 31 Jan 2024 · 1 repository · arXiv:2402.01766Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval 31 Jan 2024 · 3 repositories · arXiv:2401.18059Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
SCAPE: Searching Conceptual Architecture Prompts using Evolution 31 Jan 2024 · 1 repository · arXiv:2402.00089
-
Scavenging Hyena: Distilling Transformers into Long Convolution Models 31 Jan 2024 · 0 repositories · arXiv:2401.17574
-
Evaluating the Capabilities of LLMs for Supporting Anticipatory Impact Assessment 31 Jan 2024 · 0 repositories · arXiv:2401.18028
-
WSC+: Enhancing The Winograd Schema Challenge Using Tree-of-Experts 31 Jan 2024 · 1 repository · arXiv:2401.17703
-
A Cross-Language Investigation into Jailbreak Attacks in Large Language Models 30 Jan 2024 · 0 repositories · arXiv:2401.16765
-
Conditional and Modal Reasoning in Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.17169
-
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.16745Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Performance Assessment of ChatGPT vs Bard in Detecting Alzheimer's Dementia 30 Jan 2024 · 0 repositories · arXiv:2402.01751
-
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks 30 Jan 2024 · 1 repository · arXiv:2401.17263Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Synthetic Dialogue Dataset Generation using LLM Agents 30 Jan 2024 · 1 repository · arXiv:2401.17461
-
Weaver: Foundation Models for Creative Writing 30 Jan 2024 · 0 repositories · arXiv:2401.17268
-
3DG: A Framework for Using Generative AI for Handling Sparse Learner Performance Data From Intelligent Tutoring Systems 29 Jan 2024 · 1 repository · arXiv:2402.01746
-
Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties 29 Jan 2024 · 0 repositories · arXiv:2402.01741
-
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports 29 Jan 2024 · 0 repositories · arXiv:2401.16578
-
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning 29 Jan 2024 · 0 repositories · arXiv:2401.16185
-
Prompt4Vis: Prompting Large Language Models with Example Mining and Schema Filtering for Tabular Data Visualization 29 Jan 2024 · 0 repositories · arXiv:2402.07909
-
Response Generation for Cognitive Behavioral Therapy with Large Language Models: Comparative Study with Socratic Questioning 29 Jan 2024 · 0 repositories · arXiv:2401.15966
-
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors 29 Jan 2024 · 0 repositories · arXiv:2401.16310
-
Identifying and Improving Disability Bias in GPT-Based Resume Screening 28 Jan 2024 · 0 repositories · arXiv:2402.01732
-
PRE: A Peer Review Based Large Language Model Evaluator 28 Jan 2024 · 0 repositories · arXiv:2401.15641
-
DataFrame QA: A Universal LLM Framework on DataFrame Question Answering Without Data Exposure 27 Jan 2024 · 0 repositories · arXiv:2401.15463
-
Improving Medical Reasoning through Retrieval and Self-Reflection with Retrieval-Augmented Large Language Models 27 Jan 2024 · 1 repository · arXiv:2401.15269
-
MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries 27 Jan 2024 · 2 repositories · arXiv:2401.15391Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Prompting Diverse Ideas: Increasing AI Idea Variance 27 Jan 2024 · 0 repositories · arXiv:2402.01727
-
ChemDFM: A Large Language Foundation Model for Chemistry 26 Jan 2024 · 1 repository · arXiv:2401.14818Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias 26 Jan 2024 · 0 repositories · arXiv:2401.14589
-
Evaluation of LLM Chatbots for OSINT-based Cyber Threat Awareness 26 Jan 2024 · 0 repositories · arXiv:2401.15127
-
From GPT-4 to Gemini and Beyond: Assessing the Landscape of MLLMs on Generalizability, Trustworthiness and Causality through Four Modalities 26 Jan 2024 · 0 repositories · arXiv:2401.15071
-
Health Text Simplification: An Annotated Corpus for Digestive Cancer Education and Novel Strategies for Reinforcement Learning 26 Jan 2024 · 1 repository · arXiv:2401.15043
-
Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks 26 Jan 2024 · 0 repositories · arXiv:2401.15170
-
A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification 25 Jan 2024 · 0 repositories · arXiv:2401.13887
-
Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution 25 Jan 2024 · 0 repositories · arXiv:2401.13996
-
LLM on FHIR -- Demystifying Health Records 25 Jan 2024 · 0 repositories · arXiv:2402.01711
-
Prompting Large Language Models for Zero-Shot Clinical Prediction with Structured Longitudinal Electronic Health Record Data 25 Jan 2024 · 1 repository · arXiv:2402.01713Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation 25 Jan 2024 · 0 repositories · arXiv:2401.13867
-
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models 25 Jan 2024 · 2 repositories · arXiv:2401.13919Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Zero-shot Sequential Neuro-symbolic Reasoning for Automatically Generating Architecture Schematic Designs 25 Jan 2024 · 0 repositories · arXiv:2402.00052
-
Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4 24 Jan 2024 · 0 repositories · arXiv:2401.13810
-
Fine-Grained Stateful Knowledge Exploration: A Novel Paradigm for Integrating Knowledge Graphs with Large Language Models 24 Jan 2024 · 1 repository · arXiv:2401.13444
-
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models 24 Jan 2024 · 1 repository · arXiv:2401.13311Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Evaluation of General Large Language Models in Contextually Assessing Semantic Concepts Extracted from Adult Critical Care Electronic Health Record Notes 24 Jan 2024 · 0 repositories · arXiv:2401.13588
-
How Good is ChatGPT at Face Biometrics? A First Look into Recognition, Soft Biometrics, and Explainability 24 Jan 2024 · 1 repository · arXiv:2401.13641
-
Research about the Ability of LLM in the Tamper-Detection Area 24 Jan 2024 · 0 repositories · arXiv:2401.13504
-
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data 24 Jan 2024 · 0 repositories · arXiv:2401.13223
-
ARGS: Alignment as Reward-Guided Search 23 Jan 2024 · 1 repository · arXiv:2402.01694Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 3 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning 23 Jan 2024 · 0 repositories · arXiv:2401.12863
-
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding 23 Jan 2024 · 1 repository · arXiv:2401.12954Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 4 pointer-only (licence)
-
Quality of Answers of Generative Large Language Models vs Peer Patients for Interpreting Lab Test Results for Lay Patients: Evaluation Study 23 Jan 2024 · 0 repositories · arXiv:2402.01693
-
Investigating Large Language Models for Financial Causality Detection in Multilingual Setup 22 Jan 2024 · 0 repositories
-
Revolutionizing Finance with LLMs: An Overview of Applications and Insights 22 Jan 2024 · 0 repositories · arXiv:2401.11641
-
Speak It Out: Solving Symbol-Related Problems with Symbol-to-Language Conversion for Language Models 22 Jan 2024 · 1 repository · arXiv:2401.11725
-
SuperCLUE-Math6: Graded Multi-Step Math Reasoning Benchmark for LLMs in Chinese 22 Jan 2024 · 1 repository · arXiv:2401.11819
-
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution 21 Jan 2024 · 1 repository · arXiv:2401.11356