Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 25
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 25 of 29: papers 2,401 to 2,500 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Aligning Large Multimodal Models with Factually Augmented RLHF 25 Sep 2023 · 0 repositories · arXiv:2309.14525
-
Evaluating Cognitive Maps and Planning in Large Language Models with CogEval 25 Sep 2023 · 0 repositories · arXiv:2309.15129
-
Guess & Sketch: Language Model Guided Transpilation 25 Sep 2023 · 0 repositories · arXiv:2309.14396
-
Physics of Language Models: Part 3.2, Knowledge Manipulation 25 Sep 2023 · 0 repositories · arXiv:2309.14402
-
Watch Your Language: Investigating Content Moderation with Large Language Models 25 Sep 2023 · 0 repositories · arXiv:2309.14517
-
Natural Language based Context Modeling and Reasoning for Ubiquitous Computing with Large Language Models: A Tutorial 24 Sep 2023 · 0 repositories · arXiv:2309.15074
-
A Chat About Boring Problems: Studying GPT-based text normalization 23 Sep 2023 · 0 repositories · arXiv:2309.13426
-
Probing the Moral Development of Large Language Models through Defining Issues Test 23 Sep 2023 · 0 repositories · arXiv:2309.13356
-
GlotScript: A Resource and Tool for Low Resource Writing System Identification 23 Sep 2023 · 1 repository · arXiv:2309.13320Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs 22 Sep 2023 · 2 repositories · arXiv:2309.13007Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
AceGPT, Localizing Large Language Models in Arabic 21 Sep 2023 · 1 repository · arXiv:2309.12053
-
Can LLMs Augment Low-Resource Reading Comprehension Datasets? Opportunities and Challenges 21 Sep 2023 · 0 repositories · arXiv:2309.12426
-
Code Soliloquies for Accurate Calculations in Large Language Models 21 Sep 2023 · 1 repository · arXiv:2309.12161Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
LLMR: Real-time Prompting of Interactive Worlds using Large Language Models 21 Sep 2023 · 0 repositories · arXiv:2309.12276
-
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset 21 Sep 2023 · 5 repositories · arXiv:2309.11998
-
MiChao-HuaFen 1.0: A Specialized Pre-trained Corpus Dataset for Domain-specific Large Models 21 Sep 2023 · 0 repositories · arXiv:2309.13079
-
Multimodal Deep Learning for Scientific Imaging Interpretation 21 Sep 2023 · 0 repositories · arXiv:2309.12460
-
Random-Access Infinite Context Length for Transformers 21 Sep 2023 · 1 repository
-
SPRING: Studying Papers and Reasoning to play Games 21 Sep 2023 · 0 repositories
-
The Cambridge Law Corpus: A Dataset for Legal AI Research 21 Sep 2023 · 0 repositories · arXiv:2309.12269
-
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" 21 Sep 2023 · 2 repositories · arXiv:2309.12288Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Toward Re-Identifying Any Animal 21 Sep 2023 · 0 repositories
-
Generative AI in Mafia-like Game Simulation 20 Sep 2023 · 0 repositories · arXiv:2309.11672
-
Is GPT4 a Good Trader? 20 Sep 2023 · 0 repositories · arXiv:2309.10982
-
An Evaluation of GPT-4 on the ETHICS Dataset 19 Sep 2023 · 0 repositories · arXiv:2309.10492
-
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models 19 Sep 2023 · 1 repository · arXiv:2309.10444Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Generative AI vs. AGI: The Cognitive Strengths and Weaknesses of Modern LLMs 19 Sep 2023 · 0 repositories · arXiv:2309.10371
-
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition 19 Sep 2023 · 0 repositories · arXiv:2309.10294
-
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback 19 Sep 2023 · 1 repository · arXiv:2309.10691
-
PolicyGPT: Automated Analysis of Privacy Policies with Large Language Models 19 Sep 2023 · 0 repositories · arXiv:2309.10238
-
Facilitating NSFW Text Detection in Open-Domain Dialogue Systems via Knowledge Distillation 18 Sep 2023 · 1 repository · arXiv:2309.09749
-
Embrace Divergence for Richer Insights: A Multi-document Summarization Benchmark and a Case Study on Summarizing Diverse Information from News Articles 17 Sep 2023 · 1 repository · arXiv:2309.09369
-
Performance of the Pre-Trained Large Language Model GPT-4 on Automated Short Answer Grading 17 Sep 2023 · 0 repositories · arXiv:2309.09338
-
Examining the Influence of Varied Levels of Domain Knowledge Base Inclusion in GPT-based Intelligent Tutors 16 Sep 2023 · 1 repository · arXiv:2309.12367
-
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? 16 Sep 2023 · 1 repository · arXiv:2309.08963
-
GPT-Lab: Next Generation Of Optimal Chemistry Discovery By GPT Driven Robotic Lab 15 Sep 2023 · 0 repositories · arXiv:2309.16721
-
ICLEF: In-Context Learning with Expert Feedback for Explainable Style Transfer 15 Sep 2023 · 1 repository · arXiv:2309.08583
-
InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning 15 Sep 2023 · 1 repository · arXiv:2309.13064Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
OpenAI Cribbed Our Tax Example, But Can GPT-4 Really Do Tax? 15 Sep 2023 · 0 repositories · arXiv:2309.09992
-
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation? 14 Sep 2023 · 0 repositories · arXiv:2309.07462
-
Generative AI 13 Sep 2023 · 0 repositories · arXiv:2309.07930
-
In-Contextual Gender Bias Suppression for Large Language Models 13 Sep 2023 · 1 repository · arXiv:2309.07251
-
Large Language Models Can Infer Psychological Dispositions of Social Media Users 13 Sep 2023 · 0 repositories · arXiv:2309.08631
-
RAIN: Your Language Models Can Align Themselves without Finetuning 13 Sep 2023 · 1 repository · arXiv:2309.07124Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SafetyBench: Evaluating the Safety of Large Language Models 13 Sep 2023 · 1 repository · arXiv:2309.07045Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models 12 Sep 2023 · 3 repositories · arXiv:2309.06085
-
Leveraging Large Language Models and Weak Supervision for Social Media data annotation: an evaluation using COVID-19 self-reported vaccination tweets 12 Sep 2023 · 0 repositories · arXiv:2309.06503
-
Strategic Behavior of Large Language Models: Game Structure vs. Contextual Framing 12 Sep 2023 · 0 repositories · arXiv:2309.05898
-
The Moral Machine Experiment on Large Language Models 12 Sep 2023 · 1 repository · arXiv:2309.05958
-
An Empirical Study of NetOps Capability of Pre-Trained Large Language Models 11 Sep 2023 · 0 repositories · arXiv:2309.05557
-
Black-Box Analysis: GPTs Across Time in Legal Textual Entailment Task 11 Sep 2023 · 0 repositories · arXiv:2309.05501
-
Large Language Model for Science: A Study on P vs. NP 11 Sep 2023 · 1 repository · arXiv:2309.05689
-
Efficient Finetuning Large Language Models For Vietnamese Chatbot 9 Sep 2023 · 0 repositories · arXiv:2309.04646
-
FIMO: A Challenge Formal Dataset for Automated Theorem Proving 8 Sep 2023 · 1 repository · arXiv:2309.04295
-
From Sparse to Dense: GPT-4 Summarization with Chain of Density Prompting 8 Sep 2023 · 0 repositories · arXiv:2309.04269
-
NESTLE: a No-Code Tool for Statistical Analysis of Legal Corpus 8 Sep 2023 · 1 repository · arXiv:2309.04146
-
Enhancing Pipeline-Based Conversational Agents with Large Language Models 7 Sep 2023 · 0 repositories · arXiv:2309.03748
-
Supervised Learning and Large Language Model Benchmarks on Mental Health Datasets: Cognitive Distortions and Suicidal Risks in Chinese Social Media 7 Sep 2023 · 2 repositories · arXiv:2309.03564
-
Evaluation of large language models for discovery of gene set function 7 Sep 2023 · 1 repository · arXiv:2309.04019
-
Zero-Shot Audio Captioning via Audibility Guidance 7 Sep 2023 · 0 repositories · arXiv:2309.03884
-
GPT Can Solve Mathematical Problems Without a Calculator 6 Sep 2023 · 1 repository · arXiv:2309.03241Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Knowledge Solver: Teaching LLMs to Search for Domain Knowledge from Knowledge Graphs 6 Sep 2023 · 0 repositories · arXiv:2309.03118
-
Large Language Models for Automated Open-domain Scientific Hypotheses Discovery 6 Sep 2023 · 1 repository · arXiv:2309.02726Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models 5 Sep 2023 · 1 repository · arXiv:2309.01940
-
Data-Juicer: A One-Stop Data Processing System for Large Language Models 5 Sep 2023 · 2 repositories · arXiv:2309.02033
-
Bias Testing and Mitigation in LLM-based Code Generation 3 Sep 2023 · 0 repositories · arXiv:2309.14345
-
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties 2 Sep 2023 · 1 repository · arXiv:2309.00779
-
Large Language Models for Semantic Monitoring of Corporate Disclosures: A Case Study on Korea's Top 50 KOSPI Companies 1 Sep 2023 · 0 repositories · arXiv:2309.00208
-
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes 1 Sep 2023 · 1 repository · arXiv:2309.00237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models 31 Aug 2023 · 1 repository · arXiv:2308.16458Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Enhancing Subtask Performance of Multi-modal Large Language Model 31 Aug 2023 · 0 repositories · arXiv:2308.16474
-
GPT has become financially literate: Insights from financial literacy tests of GPT and a preliminary test of how people use it as a source of advice 31 Aug 2023 · 0 repositories · arXiv:2309.00649
-
Linking microblogging sentiments to stock price movement: An application of GPT-4 31 Aug 2023 · 0 repositories · arXiv:2308.16771
-
TouchStone: Evaluating Vision-Language Models by Language Models 31 Aug 2023 · 1 repository · arXiv:2308.16890
-
Large Language Models as Data Preprocessors 30 Aug 2023 · 0 repositories · arXiv:2308.16361
-
Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models 29 Aug 2023 · 0 repositories · arXiv:2308.15022
-
Breaking the Bank with ChatGPT: Few-Shot Text Classification for Finance 28 Aug 2023 · 0 repositories · arXiv:2308.14634
-
Large Language Models Streamline Automated Machine Learning for Clinical Studies 27 Aug 2023 · 1 repository · arXiv:2308.14120
-
Examining User-Friendly and Open-Sourced Large GPT Models: A Survey on Language, Multimodal, and Scientific GPT Models 27 Aug 2023 · 1 repository · arXiv:2308.14149
-
MedAlign: A Clinician-Generated Dataset for Instruction Following with Electronic Medical Records 27 Aug 2023 · 0 repositories · arXiv:2308.14089
-
A Wide Evaluation of ChatGPT on Affective Computing Tasks 26 Aug 2023 · 1 repository · arXiv:2308.13911
-
Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content 26 Aug 2023 · 0 repositories · arXiv:2308.13768
-
Exploring Large Language Models for Knowledge Graph Completion 26 Aug 2023 · 1 repository · arXiv:2308.13916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions 25 Aug 2023 · 1 repository · arXiv:2309.12342
-
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs 25 Aug 2023 · 1 repository · arXiv:2308.13387Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MLLM-DataEngine: An Iterative Refinement Approach for MLLM 25 Aug 2023 · 1 repository · arXiv:2308.13566Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Prompting a Large Language Model to Generate Diverse Motivational Messages: A Comparison with Human-Written Messages 25 Aug 2023 · 0 repositories · arXiv:2308.13479
-
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research 25 Aug 2023 · 2 repositories · arXiv:2308.13149Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
GPTEval: A Survey on Assessments of ChatGPT and GPT-4 24 Aug 2023 · 0 repositories · arXiv:2308.12488
-
Harnessing the Power of David against Goliath: Exploring Instruction Data Generation without Using Closed-Source Models 24 Aug 2023 · 0 repositories · arXiv:2308.12711
-
Language as Reality: A Co-Creative Storytelling Game Experience in 1001 Nights using Generative AI 24 Aug 2023 · 0 repositories · arXiv:2308.12915
-
Mind vs. Mouth: On Measuring Re-judge Inconsistency of Social Bias in Large Language Models 24 Aug 2023 · 0 repositories · arXiv:2308.12578
-
VIGC: Visual Instruction Generation and Correction 24 Aug 2023 · 2 repositories · arXiv:2308.12714Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Are ChatGPT and GPT-4 Good Poker Players? -- A Pre-Flop Analysis 23 Aug 2023 · 0 repositories · arXiv:2308.12466
-
Diagnosing Infeasible Optimization Problems Using Large Language Models 23 Aug 2023 · 1 repository · arXiv:2308.12923
-
InstructionGPT-4: A 200-Instruction Paradigm for Fine-Tuning MiniGPT-4 23 Aug 2023 · 3 repositories · arXiv:2308.12067Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments 23 Aug 2023 · 1 repository · arXiv:2308.12086Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Prompt-Based Length Controlled Generation with Reinforcement Learning 23 Aug 2023 · 0 repositories · arXiv:2308.12030
-
Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis 22 Aug 2023 · 1 repository · arXiv:2308.11224Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
GPT-in-the-Loop: Adaptive Decision-Making for Multiagent Systems 21 Aug 2023 · 0 repositories · arXiv:2308.10435