Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 24
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 24 of 29: papers 2,301 to 2,400 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Diversity of Thought Improves Reasoning Abilities of LLMs 11 Oct 2023 · 0 repositories · arXiv:2310.07088
-
Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs 11 Oct 2023 · 0 repositories · arXiv:2310.07251
-
Do Large Language Models have Shared Weaknesses in Medical Question Answering? 11 Oct 2023 · 0 repositories · arXiv:2310.07225
-
Large Language Models Are Zero-Shot Time Series Forecasters 11 Oct 2023 · 2 repositories · arXiv:2310.07820Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
From Supervised to Generative: A Novel Paradigm for Tabular Deep Learning with Large Language Models 11 Oct 2023 · 0 repositories · arXiv:2310.07338
-
Diffusion Models for Wireless Communications 11 Oct 2023 · 0 repositories · arXiv:2310.07312
-
Automated clinical coding using off-the-shelf large language models 10 Oct 2023 · 0 repositories · arXiv:2310.06552
-
Generating and Evaluating Tests for K-12 Students with Language Model Simulations: A Case Study on Sentence Reading Efficiency 10 Oct 2023 · 0 repositories · arXiv:2310.06837
-
GPT-4 as an Agronomist Assistant? Answering Agriculture Exams Using Large Language Models 10 Oct 2023 · 0 repositories · arXiv:2310.06225
-
Large Language Models for Propaganda Detection 10 Oct 2023 · 2 repositories · arXiv:2310.06422
-
LLMs as Potential Brainstorming Partners for Math and Science Problems 10 Oct 2023 · 0 repositories · arXiv:2310.10677
-
Multilingual Jailbreak Challenges in Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06474
-
NEWTON: Are Large Language Models Capable of Physical Reasoning? 10 Oct 2023 · 0 repositories · arXiv:2310.07018
-
SWE-bench: Can Language Models Resolve Real-World GitHub Issues? 10 Oct 2023 · 8 repositories · arXiv:2310.06770Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06627
-
FireAct: Toward Language Agent Fine-tuning 9 Oct 2023 · 0 repositories · arXiv:2310.05915
-
Integrating Graphs with Large Language Models: Methods and Prospects 9 Oct 2023 · 0 repositories · arXiv:2310.05499
-
Integrating Stock Features and Global Information via Large Language Models for Enhanced Stock Return Prediction 9 Oct 2023 · 0 repositories · arXiv:2310.05627
-
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena 9 Oct 2023 · 1 repository · arXiv:2310.05746Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Reinforcement Learning in the Era of LLMs: What is Essential? What is needed? An RL Perspective on RLHF, Prompting, and Beyond 9 Oct 2023 · 2 repositories · arXiv:2310.06147Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SC-Safety: A Multi-round Open-ended Question Adversarial Safety Benchmark for Large Language Models in Chinese 9 Oct 2023 · 0 repositories · arXiv:2310.05818
-
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models 9 Oct 2023 · 0 repositories · arXiv:2310.06117
-
The Importance of Prompt Tuning for Automated Neuron Explanations 9 Oct 2023 · 0 repositories · arXiv:2310.06200
-
ChatRadio-Valuer: A Chat Large Language Model for Generalizable Radiology Report Generation Based on Multi-institution and Multi-system Data 8 Oct 2023 · 0 repositories · arXiv:2310.05242
-
Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements 8 Oct 2023 · 1 repository · arXiv:2310.05140Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
LLM4VV: Developing LLM-Driven Testsuite for Compiler Validation 8 Oct 2023 · 1 repository · arXiv:2310.04963
-
Zero-Shot Detection of Machine-Generated Codes 8 Oct 2023 · 1 repository · arXiv:2310.05103
-
DiffNAS: Bootstrapping Diffusion Models by Prompting for Better Architectures 7 Oct 2023 · 0 repositories · arXiv:2310.04750
-
Large Language Models for Spatial Trajectory Patterns Mining 7 Oct 2023 · 0 repositories · arXiv:2310.04942
-
An In-Context Learning Agent for Formal Theorem-Proving 6 Oct 2023 · 1 repository · arXiv:2310.04353
-
Coding by Design: GPT-4 empowers Agile Model Driven Development 6 Oct 2023 · 0 repositories · arXiv:2310.04304
-
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models 6 Oct 2023 · 2 repositories · arXiv:2310.04406Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Adversarial Machine Learning for Social Good: Reframing the Adversary as an Ally 5 Oct 2023 · 0 repositories · arXiv:2310.03614
-
Automating Human Tutor-Style Programming Feedback: Leveraging GPT-4 Tutor Model for Hint Generation and GPT-3.5 Student Model for Hint Validation 5 Oct 2023 · 2 repositories · arXiv:2310.03780Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Benchmarking a foundation LLM on its ability to re-label structure names in accordance with the AAPM TG-263 report 5 Oct 2023 · 0 repositories · arXiv:2310.03874
-
MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation 5 Oct 2023 · 2 repositories · arXiv:2310.03302Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 6 pointer-only (licence)
-
Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning 5 Oct 2023 · 1 repository · arXiv:2310.03249Syntology official (archive's flag): 28 ran · 28 ran (of which 0 constructed an object rather than computing a result; 28 with no instrument failure: 0 honoured, 0 violated, 28 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 28 harvested samples) · 28 pointer-only (licence)
-
Evaluating Hallucinations in Chinese Large Language Models 5 Oct 2023 · 3 repositories · arXiv:2310.03368Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
SteP: Stacked LLM Policies for Web Actions 5 Oct 2023 · 0 repositories · arXiv:2310.03720
-
Learning Personalized Alignment for Evaluating Open-ended Text Generation 5 Oct 2023 · 0 repositories · arXiv:2310.03304
-
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning 5 Oct 2023 · 1 repository · arXiv:2310.03731Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Procedural Text Mining with Large Language Models 5 Oct 2023 · 1 repository · arXiv:2310.03376
-
Reformulating Domain Adaptation of Large Language Models as Adapt-Retrieve-Revise: A Case Study on Chinese Legal Domain 5 Oct 2023 · 1 repository · arXiv:2310.03328
-
Can Language Models Employ the Socratic Method? Experiments with Code Debugging 4 Oct 2023 · 1 repository · arXiv:2310.03210
-
CITING: Large Language Models Create Curriculum for Instruction Tuning 4 Oct 2023 · 0 repositories · arXiv:2310.02527
-
DQ-LoRe: Dual Queries with Low Rank Approximation Re-ranking for In-Context Learning 4 Oct 2023 · 1 repository · arXiv:2310.02954Syntology official (archive's flag): 2 ran · 5 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
GPT-4 as an interface between researchers and computational software: improving usability and reproducibility 4 Oct 2023 · 0 repositories · arXiv:2310.11458
-
How FaR Are Large Language Models From Agents with Theory-of-Mind? 4 Oct 2023 · 1 repository · arXiv:2310.03051
-
Large Language Model Cascades with Mixture of Thoughts Representations for Cost-efficient Reasoning 4 Oct 2023 · 1 repository · arXiv:2310.03094
-
Multimodal Question Answering for Unified Information Extraction 4 Oct 2023 · 1 repository · arXiv:2310.03017
-
Robust and Interpretable Medical Image Classifiers via Concept Bottleneck Models 4 Oct 2023 · 0 repositories · arXiv:2310.03182
-
T³Bench: Benchmarking Current Progress in Text-to-3D Generation 4 Oct 2023 · 1 repository · arXiv:2310.02977Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions 4 Oct 2023 · 0 repositories · arXiv:2310.03016
-
Benchmarking and Improving Generator-Validator Consistency of Language Models 3 Oct 2023 · 0 repositories · arXiv:2310.01846
-
Can GPT-4 Replicate Empirical Software Engineering Research? 3 Oct 2023 · 0 repositories · arXiv:2310.01727
-
Can large language models provide useful feedback on research papers? A large-scale empirical analysis 3 Oct 2023 · 1 repository · arXiv:2310.01783
-
Automatic Pair Construction for Contrastive Post-training 3 Oct 2023 · 1 repository · arXiv:2310.02263
-
EcoAssistant: Using LLM Assistant More Affordably and Accurately 3 Oct 2023 · 1 repository · arXiv:2310.03046Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Editing Personality for Large Language Models 3 Oct 2023 · 1 repository · arXiv:2310.02168
-
HallE-Control: Controlling Object Hallucination in Large Multimodal Models 3 Oct 2023 · 2 repositories · arXiv:2310.01779Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Instances Need More Care: Rewriting Prompts for Instances with LLMs in the Loop Yields Better Zero-Shot Performance 3 Oct 2023 · 1 repository · arXiv:2310.02107Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Investigating Large Language Models' Perception of Emotion Using Appraisal Theory 3 Oct 2023 · 0 repositories · arXiv:2310.04450
-
Low-Resource Languages Jailbreak GPT-4 3 Oct 2023 · 0 repositories · arXiv:2310.02446
-
Reinforcement Learning from Automatic Feedback for High-Quality Unit Test Generation 3 Oct 2023 · 0 repositories · arXiv:2310.02368
-
Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation 3 Oct 2023 · 1 repository · arXiv:2310.02304Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
What's Next in Affective Modeling? Large Language Models 3 Oct 2023 · 0 repositories · arXiv:2310.18322
-
Large Language Model-Powered Smart Contract Vulnerability Detection: New Perspectives 2 Oct 2023 · 1 repository · arXiv:2310.01152
-
LoFT: Local Proxy Fine-tuning For Improving Transferability Of Adversarial Attacks Against Large Language Model 2 Oct 2023 · 0 repositories · arXiv:2310.04445
-
Probing the Multi-turn Planning Capabilities of LLMs via 20 Question Games 2 Oct 2023 · 1 repository · arXiv:2310.01468Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
TRAM: Benchmarking Temporal Reasoning for Large Language Models 2 Oct 2023 · 1 repository · arXiv:2310.00835
-
UltraFeedback: Boosting Language Models with Scaled AI Feedback 2 Oct 2023 · 4 repositories · arXiv:2310.01377Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Who is ChatGPT? Benchmarking LLMs' Psychological Portrayal Using PsychoBench 2 Oct 2023 · 1 repository · arXiv:2310.01386Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
BooookScore: A systematic exploration of book-length summarization in the era of LLMs 1 Oct 2023 · 2 repositories · arXiv:2310.00785Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
TIGERScore: Towards Building Explainable Metric for All Text Generation Tasks 1 Oct 2023 · 1 repository · arXiv:2310.00752Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples)
-
AutomaTikZ: Text-Guided Synthesis of Scientific Vector Graphics with TikZ 30 Sep 2023 · 1 repository · arXiv:2310.00367Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
Graph Neural Architecture Search with GPT-4 30 Sep 2023 · 0 repositories · arXiv:2310.01436
-
ValueDCG: Measuring Comprehensive Human Value Understanding Ability of Language Models 30 Sep 2023 · 1 repository · arXiv:2310.00378
-
UPAR: A Kantian-Inspired Prompting Framework for Enhancing Large Language Model Capabilities 30 Sep 2023 · 0 repositories · arXiv:2310.01441
-
A Large Language Model Approach to Educational Survey Feedback Analysis 29 Sep 2023 · 0 repositories · arXiv:2309.17447
-
Benchmarking the Abilities of Large Language Models for RDF Knowledge Graph Creation and Comprehension: How Well Do LLMs Speak Turtle? 29 Sep 2023 · 3 repositories · arXiv:2309.17122
-
CRAFT: Customizing LLMs by Creating and Retrieving from Specialized Toolsets 29 Sep 2023 · 2 repositories · arXiv:2309.17428Syntology official: harvested, nothing ran · 0 ran · 5 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks 29 Sep 2023 · 1 repository · arXiv:2309.17167
-
Enhancing Large Language Models in Coding Through Multi-Perspective Self-Consistency 29 Sep 2023 · 1 repository · arXiv:2309.17272Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples)
-
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation 29 Sep 2023 · 2 repositories · arXiv:2309.17234Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SCALE: Synergized Collaboration of Asymmetric Language Translation Engines 29 Sep 2023 · 1 repository · arXiv:2309.17061
-
SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation 29 Sep 2023 · 1 repository · arXiv:2310.00074
-
Split and Merge: Aligning Position Biases in LLM-based Evaluators 29 Sep 2023 · 0 repositories · arXiv:2310.01432
-
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4 29 Sep 2023 · 1 repository · arXiv:2309.17277Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving 29 Sep 2023 · 1 repository · arXiv:2309.17452Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 9 pointer-only (licence)
-
AE-GPT: Using Large Language Models to Extract Adverse Events from Surveillance Reports-A Use Case with Influenza Vaccine Adverse Events 28 Sep 2023 · 0 repositories · arXiv:2309.16150
-
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 28 Sep 2023 · 1 repository · arXiv:2309.16583Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
LawBench: Benchmarking Legal Knowledge of Large Language Models 28 Sep 2023 · 1 repository · arXiv:2309.16289Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Neuro Symbolic Reasoning for Planning: Counterexample Guided Inductive Synthesis using Large Language Models and Satisfiability Solving 28 Sep 2023 · 0 repositories · arXiv:2309.16436
-
ChatCounselor: A Large Language Models for Mental Health Support 27 Sep 2023 · 1 repository · arXiv:2309.15461
-
Benchmarking Large Language Models on CMExam - A comprehensive Chinese Medical Exam Dataset 26 Sep 2023 · 1 repository
-
CAPP-130: A Corpus of Chinese Application Privacy Policy Summarization and Interpretation 26 Sep 2023 · 1 repository
-
Question-Answering Approach to Evaluating Legal Summaries 26 Sep 2023 · 1 repository · arXiv:2309.15016
-
RankVicuna: Zero-Shot Listwise Document Reranking with Open-Source Large Language Models 26 Sep 2023 · 3 repositories · arXiv:2309.15088
-
Supersonic: Learning to Generate Source Code Optimizations in C/C++ 26 Sep 2023 · 1 repository · arXiv:2309.14846
-
VisIT-Bench: A Dynamic Benchmark for Evaluating Instruction-Following Vision-and-Language Models 26 Sep 2023 · 0 repositories