Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 26
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 26 of 29: papers 2,501 to 2,600 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
"Guinea Pig Trials" Utilizing GPT: A Novel Smart Agent-Based Modeling Approach for Studying Firm Competition and Collusion 21 Aug 2023 · 2 repositories · arXiv:2308.10974Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Large Language Models on Wikipedia-Style Survey Generation: an Evaluation in NLP Concepts 21 Aug 2023 · 1 repository · arXiv:2308.10410
-
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles 21 Aug 2023 · 1 repository · arXiv:2308.10855Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
On the Adversarial Robustness of Multi-Modal Foundation Models 21 Aug 2023 · 1 repository · arXiv:2308.10741Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models 21 Aug 2023 · 1 repository · arXiv:2308.10755
-
Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis 21 Aug 2023 · 1 repository · arXiv:2308.10783
-
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation 20 Aug 2023 · 1 repository · arXiv:2308.10335Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Can Large Language Models Find And Fix Vulnerable Software? 20 Aug 2023 · 0 repositories · arXiv:2308.10345
-
ChatEDA: A Large Language Model Powered Autonomous Agent for EDA 20 Aug 2023 · 1 repository · arXiv:2308.10204Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
ExpeL: LLM Agents Are Experiential Learners 20 Aug 2023 · 2 repositories · arXiv:2308.10144Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Large Transformers are Better EEG Learners 20 Aug 2023 · 0 repositories · arXiv:2308.11654
-
StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data 20 Aug 2023 · 1 repository · arXiv:2308.10253Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models 19 Aug 2023 · 1 repository · arXiv:2308.09975Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
How susceptible are LLMs to Logical Fallacies? 18 Aug 2023 · 1 repository · arXiv:2308.09853
-
Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment 18 Aug 2023 · 2 repositories · arXiv:2308.09662
-
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct 18 Aug 2023 · 1 repository · arXiv:2308.09583Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
NaijaRC: A Multi-choice Reading Comprehension Dataset for Nigerian Languages 18 Aug 2023 · 1 repository · arXiv:2308.09768
-
Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes 17 Aug 2023 · 2 repositories · arXiv:2308.08769
-
CMB: A Comprehensive Medical Benchmark in Chinese 17 Aug 2023 · 2 repositories · arXiv:2308.08833Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
MaScQA: A Question Answering Dataset for Investigating Materials Science Knowledge of Large Language Models 17 Aug 2023 · 0 repositories · arXiv:2308.09115
-
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models 17 Aug 2023 · 1 repository · arXiv:2308.09729Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Boosting Logical Reasoning in Large Language Models through a New Framework: The Graph of Thought 16 Aug 2023 · 0 repositories · arXiv:2308.08614
-
Self-Deception: Reverse Penetrating the Semantic Firewall of Large Language Models 16 Aug 2023 · 0 repositories · arXiv:2308.11521
-
Time Travel in LLMs: Tracing Data Contamination in Large Language Models 16 Aug 2023 · 1 repository · arXiv:2308.08493Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Domain Adaptation for Code Model-based Unit Test Case Generation 15 Aug 2023 · 0 repositories · arXiv:2308.08033
-
Large Language Models in Introductory Programming Education: ChatGPT's Performance and Implications for Assessments 15 Aug 2023 · 0 repositories · arXiv:2308.08572
-
Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification 15 Aug 2023 · 1 repository · arXiv:2308.07921
-
ChatGPT in Drug Discovery: A Case Study on Anti-Cocaine Addiction Drug Development with Chatbots 14 Aug 2023 · 0 repositories · arXiv:2308.06920
-
Dialogue for Prompting: a Policy-Gradient-Based Discrete Prompt Generation for Few-shot Learning 14 Aug 2023 · 1 repository · arXiv:2308.07272Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Large Language Models for Information Retrieval: A Survey 14 Aug 2023 · 1 repository · arXiv:2308.07107
-
Neural Authorship Attribution: Stylometric Analysis on Large Language Models 14 Aug 2023 · 1 repository · arXiv:2308.07305
-
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher 12 Aug 2023 · 1 repository · arXiv:2308.06463Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use 12 Aug 2023 · 1 repository · arXiv:2308.06595Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Assessing Student Errors in Experimentation Using Artificial Intelligence and Large Language Models: A Comparative Study with Human Raters 11 Aug 2023 · 0 repositories · arXiv:2308.06088
-
Learning Deductive Reasoning from Synthetic Corpus based on Formal Logic 11 Aug 2023 · 3 repositories · arXiv:2308.07336Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Adaptive Low Rank Adaptation of Segment Anything to Salient Object Detection 10 Aug 2023 · 1 repository · arXiv:2308.05426
-
Metacognitive Prompting Improves Understanding in Large Language Models 10 Aug 2023 · 1 repository · arXiv:2308.05342
-
Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems 10 Aug 2023 · 0 repositories · arXiv:2308.05713
-
Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment 10 Aug 2023 · 1 repository · arXiv:2308.05374
-
A Comparative Study of Open-Source Large Language Models, GPT-4 and Claude 2: Multiple-Choice Test Taking in Nephrology 9 Aug 2023 · 0 repositories · arXiv:2308.04709
-
ChatGPT for Arabic Grammatical Error Correction 8 Aug 2023 · 0 repositories · arXiv:2308.04492
-
Gromov-Wasserstein unsupervised alignment reveals structural correspondences between the color similarity structures of humans and large language models 8 Aug 2023 · 0 repositories · arXiv:2308.04381
-
Few-shot medical image classification with simple shape and texture text descriptors using vision-language models 8 Aug 2023 · 0 repositories · arXiv:2308.04005
-
Shepherd: A Critic for Language Model Generation 8 Aug 2023 · 1 repository · arXiv:2308.04592
-
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models 7 Aug 2023 · 2 repositories · arXiv:2308.03825Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench 7 Aug 2023 · 1 repository · arXiv:2308.03656Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
CORAL: Expert-Curated medical Oncology Reports to Advance Language Model Inference 7 Aug 2023 · 1 repository · arXiv:2308.03853
-
SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs 7 Aug 2023 · 1 repository · arXiv:2308.03349Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Exploiting Code Symmetries for Learning Program Semantics 7 Aug 2023 · 0 repositories · arXiv:2308.03312
-
Pre-Trained Large Language Models for Industrial Control 6 Aug 2023 · 0 repositories · arXiv:2308.03028
-
A criterion for Artificial General Intelligence: hypothetic-deductive reasoning, tested on ChatGPT 5 Aug 2023 · 0 repositories · arXiv:2308.02950
-
ChatGPT for GTFS: Benchmarking LLMs on GTFS Understanding and Retrieval 4 Aug 2023 · 1 repository · arXiv:2308.02618
-
Scaling Clinical Trial Matching Using Large Language Models: A Case Study in Oncology 4 Aug 2023 · 0 repositories · arXiv:2308.02180
-
ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation 3 Aug 2023 · 2 repositories · arXiv:2308.01861
-
Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings 3 Aug 2023 · 0 repositories · arXiv:2308.02575
-
Large Language Model Displays Emergent Ability to Interpret Novel Literary Metaphors 3 Aug 2023 · 0 repositories · arXiv:2308.01497
-
Exploring the psychology of LLMs' Moral and Legal Reasoning 2 Aug 2023 · 0 repositories · arXiv:2308.01264
-
Flows: Building Blocks of Reasoning and Collaborating AI 2 Aug 2023 · 2 repositories · arXiv:2308.01285Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
An Effective Data Creation Pipeline to Generate High-quality Financial Instruction Data for Large Language Model 31 Jul 2023 · 0 repositories · arXiv:2308.01415
-
Deception Abilities Emerged in Large Language Models 31 Jul 2023 · 0 repositories · arXiv:2307.16513
-
Evaluating ChatGPT and GPT-4 for Visual Programming 30 Jul 2023 · 0 repositories · arXiv:2308.02522
-
ChatHome: Development and Evaluation of a Domain-Specific Language Model for Home Renovation 28 Jul 2023 · 1 repository · arXiv:2307.15290Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 17 harvested samples) · 1 pointer-only (licence)
-
RSGPT: A Remote Sensing Vision Language Model and Benchmark 28 Jul 2023 · 2 repositories · arXiv:2307.15266
-
LLMediator: GPT-4 Assisted Online Dispute Resolution 27 Jul 2023 · 0 repositories · arXiv:2307.16732
-
SuperCLUE: A Comprehensive Chinese Large Language Model Benchmark 27 Jul 2023 · 0 repositories · arXiv:2307.15020
-
Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data 26 Jul 2023 · 1 repository · arXiv:2307.14385
-
Unveiling Security, Privacy, and Ethical Concerns of ChatGPT 26 Jul 2023 · 0 repositories · arXiv:2307.14192
-
ARB: Advanced Reasoning Benchmark for Large Language Models 25 Jul 2023 · 0 repositories · arXiv:2307.13692
-
How Can Large Language Models Help Humans in Design and Manufacturing? 25 Jul 2023 · 0 repositories · arXiv:2307.14377
-
Is GPT a Computational Model of Emotion? Detailed Analysis 25 Jul 2023 · 0 repositories · arXiv:2307.13779
-
Predicting Code Coverage without Execution 25 Jul 2023 · 1 repository · arXiv:2307.13383
-
Performance of Large Language Models in a Computer Science Degree Program 24 Jul 2023 · 0 repositories · arXiv:2308.02432
-
Enhancing CLIP with GPT-4: Harnessing Visual Descriptions as Prompts 21 Jul 2023 · 1 repository · arXiv:2307.11661Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
GPT-4 Can't Reason 21 Jul 2023 · 0 repositories · arXiv:2308.03762
-
Assessing Large Language Models' ability to predict how humans balance self-interest and the interest of others 21 Jul 2023 · 0 repositories · arXiv:2307.12776
-
A LLM Assisted Exploitation of AI-Guardian 20 Jul 2023 · 0 repositories · arXiv:2307.15008
-
Instruction-following Evaluation through Verbalizer Manipulation 20 Jul 2023 · 0 repositories · arXiv:2307.10558
-
L-Eval: Instituting Standardized Evaluation for Long Context Language Models 20 Jul 2023 · 3 repositories · arXiv:2307.11088Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models 20 Jul 2023 · 2 repositories · arXiv:2307.11137
-
PharmacyGPT: The AI Pharmacist 19 Jul 2023 · 0 repositories · arXiv:2307.10432
-
ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning 18 Jul 2023 · 0 repositories · arXiv:2307.09474
-
Emotional Intelligence of Large Language Models 18 Jul 2023 · 0 repositories · arXiv:2307.09042
-
How is ChatGPT's behavior changing over time? 18 Jul 2023 · 4 repositories · arXiv:2307.09009Syntology official (archive's flag): 1 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 5 pointer-only (licence)
-
Abductive Reasoning with the GPT-4 Language Model: Case studies from criminal investigation, medical practice, scientific research 17 Jul 2023 · 0 repositories · arXiv:2307.10250
-
AlpaGasus: Training A Better Alpaca with Fewer Data 17 Jul 2023 · 3 repositories · arXiv:2307.08701Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ChatGPT is Good but Bing Chat is Better for Vietnamese Students 17 Jul 2023 · 0 repositories · arXiv:2307.08272
-
COLLIE: Systematic Construction of Constrained Text Generation Tasks 17 Jul 2023 · 1 repository · arXiv:2307.08689Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Using an LLM to Help With Code Understanding 17 Jul 2023 · 0 repositories · arXiv:2307.08177
-
On the application of Large Language Models for language teaching and assessment technology 17 Jul 2023 · 0 repositories · arXiv:2307.08393
-
Assessing the Quality of Multiple-Choice Questions Using GPT-4 and Rule-Based Methods 16 Jul 2023 · 1 repository · arXiv:2307.08161
-
The Potential and Pitfalls of using a Large Language Model such as ChatGPT or GPT-4 as a Clinical Assistant 16 Jul 2023 · 0 repositories · arXiv:2307.08152
-
Creating a Dataset for High-Performance Computing Code Translation using LLMs: A Bridge Between OpenMP Fortran and C++ 15 Jul 2023 · 1 repository · arXiv:2307.07686
-
Leveraging Large Language Models to Generate Answer Set Programs 15 Jul 2023 · 1 repository · arXiv:2307.07699Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph 15 Jul 2023 · 3 repositories · arXiv:2307.07697Syntology official (archive's flag): 11 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 15 pointer-only (licence)
-
Large Language Models Understand and Can be Enhanced by Emotional Stimuli 14 Jul 2023 · 0 repositories · arXiv:2307.11760
-
Exploring the Integration of Large Language Models into Automatic Speech Recognition Systems: An Empirical Study 13 Jul 2023 · 0 repositories · arXiv:2307.06530
-
Distilling Large Language Models for Biomedical Knowledge Extraction: A Case Study on Adverse Drug Events 12 Jul 2023 · 0 repositories · arXiv:2307.06439
-
Prompt Generate Train (PGT): Few-shot Domain Adaption of Retrieval Augmented Generation Models for Open Book Question-Answering 12 Jul 2023 · 0 repositories · arXiv:2307.05915
-
Argumentative Segmentation Enhancement for Legal Summarization 11 Jul 2023 · 0 repositories · arXiv:2307.05081
-
Explaining Competitive-Level Programming Solutions using LLMs 11 Jul 2023 · 0 repositories · arXiv:2307.05337