Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 15
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 15 of 29: papers 1,401 to 1,500 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
AI-Enhanced Cognitive Behavioral Therapy: Deep Learning and Large Language Models for Extracting Cognitive Pathways from Social Media Texts 17 Apr 2024 · 1 repository · arXiv:2404.11449
-
Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent 17 Apr 2024 · 0 repositories · arXiv:2404.11459
-
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding 17 Apr 2024 · 0 repositories · arXiv:2404.11589
-
Towards Data-Centric Automatic R&D 17 Apr 2024 · 1 repository · arXiv:2404.11276
-
Can Language Models Solve Olympiad Programming? 16 Apr 2024 · 1 repository · arXiv:2404.10952Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CoTAR: Chain-of-Thought Attribution Reasoning with Multi-level Granularity 16 Apr 2024 · 0 repositories · arXiv:2404.10513
-
ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence 16 Apr 2024 · 1 repository · arXiv:2404.10198
-
Incubating Text Classifiers Following User Instruction with Nothing but LLM 16 Apr 2024 · 1 repository · arXiv:2404.10877
-
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents 16 Apr 2024 · 2 repositories · arXiv:2404.10774Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples)
-
Grounded Language Agent for Product Search via Intelligent Web Interactions 16 Apr 2024 · 1 repository · arXiv:2404.10887Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Self-Supervised Visual Preference Alignment 16 Apr 2024 · 1 repository · arXiv:2404.10501Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback 16 Apr 2024 · 0 repositories · arXiv:2404.10271
-
Learn Your Reference Model for Real Good Alignment 15 Apr 2024 · 0 repositories · arXiv:2404.09656
-
LLM Evaluators Recognize and Favor Their Own Generations 15 Apr 2024 · 0 repositories · arXiv:2404.13076
-
Are Medium-Sized Transformers Models still Relevant for Medical Records Processing? 15 Apr 2024 · 0 repositories · arXiv:2404.10171
-
Unveiling Imitation Learning: Exploring the Impact of Data Falsity to Large Language Model 15 Apr 2024 · 0 repositories · arXiv:2404.09717
-
Zero-shot Building Age Classification from Facade Image Using GPT-4 15 Apr 2024 · 1 repository · arXiv:2404.09921
-
Constrained C-Test Generation via Mixed-Integer Programming 12 Apr 2024 · 1 repository · arXiv:2404.08821
-
Dataset Reset Policy Optimization for RLHF 12 Apr 2024 · 1 repository · arXiv:2404.08495Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
"Don't forget to put the milk back!" Dataset for Enabling Embodied Agents to Detect Anomalous Situations 12 Apr 2024 · 0 repositories · arXiv:2404.08827
-
Small Models Are (Still) Effective Cross-Domain Argument Extractors 12 Apr 2024 · 1 repository · arXiv:2404.08579
-
Automatic Generation and Evaluation of Reading Comprehension Test Items with Large Language Models 11 Apr 2024 · 2 repositories · arXiv:2404.07720
-
Comments as Natural Logic Pivots: Improve Code Generation via Comment Perspective 11 Apr 2024 · 1 repository · arXiv:2404.07549
-
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation 11 Apr 2024 · 1 repository · arXiv:2404.07917Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples 11 Apr 2024 · 1 repository · arXiv:2404.07544Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Human Latency Conversational Turns for Spoken Avatar Systems 11 Apr 2024 · 0 repositories · arXiv:2404.16053
-
LLM Agents can Autonomously Exploit One-day Vulnerabilities 11 Apr 2024 · 0 repositories · arXiv:2404.08144
-
MM-PhyQA: Multimodal Physics Question-Answering With Multi-Image CoT Prompting 11 Apr 2024 · 0 repositories · arXiv:2404.08704
-
Dynamic Generation of Personalities with Large Language Models 10 Apr 2024 · 1 repository · arXiv:2404.07084
-
Characterizing Multimodal Long-form Summarization: A Case Study on Financial Reports 9 Apr 2024 · 0 repositories · arXiv:2404.06162
-
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders 9 Apr 2024 · 1 repository · arXiv:2404.05961
-
LLMs' Reading Comprehension Is Affected by Parametric Knowledge and Struggles with Hypothetical Statements 9 Apr 2024 · 0 repositories · arXiv:2404.06283
-
Sandwich attack: Multi-language Mixture Adaptive Attack on LLMs 9 Apr 2024 · 0 repositories · arXiv:2404.07242
-
Evaluation of an LLM in Identifying Logical Fallacies: A Call for Rigor When Adopting LLMs in HCI Research 8 Apr 2024 · 0 repositories · arXiv:2404.05213
-
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models 8 Apr 2024 · 1 repository · arXiv:2404.05221Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Relation Extraction Using Large Language Models: A Case Study on Acupuncture Point Locations 8 Apr 2024 · 0 repositories · arXiv:2404.05415
-
Use of a Structured Knowledge Base Enhances Metadata Curation by Large Language Models 8 Apr 2024 · 1 repository · arXiv:2404.05893
-
Xiwu: A Basis Flexible and Learnable LLM for High Energy Physics 8 Apr 2024 · 1 repository · arXiv:2404.08001
-
MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification 7 Apr 2024 · 1 repository · arXiv:2404.05091
-
Initial Exploration of Zero-Shot Privacy Utility Tradeoffs in Tabular Data Using GPT-4 7 Apr 2024 · 0 repositories · arXiv:2404.05047
-
MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems 6 Apr 2024 · 1 repository · arXiv:2404.04735Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Cleared for Takeoff? Compositional & Conditional Reasoning may be the Achilles Heel to (Flight-Booking) Language Agents 5 Apr 2024 · 0 repositories · arXiv:2404.04237
-
Effects of Different Prompts on the Quality of GPT-4 Responses to Dementia Care Questions 5 Apr 2024 · 0 repositories · arXiv:2404.08674
-
Scope Ambiguities in Large Language Models 5 Apr 2024 · 1 repository · arXiv:2404.04332
-
AutoWebGLM: A Large Language Model-based Web Navigating Agent 4 Apr 2024 · 1 repository · arXiv:2404.03648Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Capabilities of Large Language Models in Control Engineering: A Benchmark Study on GPT-4, Claude 3 Opus, and Gemini 1.0 Ultra 4 Apr 2024 · 0 repositories · arXiv:2404.03647
-
Conversational Disease Diagnosis via External Planner-Controlled Large Language Models 4 Apr 2024 · 1 repository · arXiv:2404.04292
-
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences 4 Apr 2024 · 0 repositories · arXiv:2404.03715
-
Evaluating LLMs at Detecting Errors in LLM Responses 4 Apr 2024 · 1 repository · arXiv:2404.03602Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Probing Large Language Models for Scalar Adjective Lexical Semantics and Scalar Diversity Pragmatics 4 Apr 2024 · 1 repository · arXiv:2404.03301
-
Reason from Fallacy: Enhancing Large Language Models' Logical Reasoning through Logical Fallacy Understanding 4 Apr 2024 · 0 repositories · arXiv:2404.04293
-
Attributions toward Artificial Agents in a modified Moral Turing Test 3 Apr 2024 · 0 repositories · arXiv:2406.11854
-
BCAmirs at SemEval-2024 Task 4: Beyond Words: A Multimodal and Multilingual Exploration of Persuasion in Memes 3 Apr 2024 · 1 repository · arXiv:2404.03022
-
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT 3 Apr 2024 · 1 repository · arXiv:2404.02403
-
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models 3 Apr 2024 · 1 repository · arXiv:2404.02823Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Task Agnostic Architecture for Algorithm Induction via Implicit Composition 3 Apr 2024 · 0 repositories · arXiv:2404.02450
-
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers? 3 Apr 2024 · 1 repository · arXiv:2404.02474Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Automated User Story Generation with Test Case Specification Using Large Language Model 2 Apr 2024 · 0 repositories · arXiv:2404.01558
-
Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack 2 Apr 2024 · 0 repositories · arXiv:2404.01833
-
IndoCulture: Exploring Geographically-Influenced Cultural Commonsense Reasoning Across Eleven Indonesian Provinces 2 Apr 2024 · 0 repositories · arXiv:2404.01854
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks 2 Apr 2024 · 1 repository · arXiv:2404.02151Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
METAL: Towards Multilingual Meta-Evaluation 2 Apr 2024 · 0 repositories · arXiv:2404.01667
-
Octopus: On-device language model for function calling of software APIs 2 Apr 2024 · 0 repositories · arXiv:2404.01549
-
Octopus v2: On-device language model for super agent 2 Apr 2024 · 0 repositories · arXiv:2404.01744
-
PATCH! Psychometrics-Assis{T}ed Ben{CH}marking of Large Language Models against Human Populations: A Case Study of Proficiency in 8th Grade Mathematics 2 Apr 2024 · 1 repository · arXiv:2404.01799
-
Toward Informal Language Processing: Knowledge of Slang in Large Language Models 2 Apr 2024 · 1 repository · arXiv:2404.02323
-
Automated Assessment of Encouragement and Warmth in Classrooms Leveraging Multimodal Emotional Features and ChatGPT 1 Apr 2024 · 0 repositories · arXiv:2404.15310
-
Forklift: An Extensible Neural Lifter 1 Apr 2024 · 0 repositories · arXiv:2404.16041
-
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations 1 Apr 2024 · 0 repositories · arXiv:2404.01266
-
LLM-RadJudge: Achieving Radiologist-Level Evaluation for X-Ray Report Generation 1 Apr 2024 · 0 repositories · arXiv:2404.00998
-
Unveiling Divergent Inductive Biases of LLMs on Temporal Data 1 Apr 2024 · 1 repository · arXiv:2404.01453
-
Algorithmic Collusion by Large Language Models 31 Mar 2024 · 0 repositories · arXiv:2404.00806
-
CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs 31 Mar 2024 · 1 repository · arXiv:2404.01343
-
CoUDA: Coherence Evaluation via Unified Data Augmentation 31 Mar 2024 · 1 repository · arXiv:2404.00681
-
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories 31 Mar 2024 · 1 repository · arXiv:2404.00599Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 3 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 4 pointer-only (licence)
-
Extracting Social Determinants of Health from Pediatric Patient Notes Using Large Language Models: Novel Corpus and Methods 31 Mar 2024 · 1 repository · arXiv:2404.00826
-
How Much are Large Language Models Contaminated? A Comprehensive Survey and the LLMSanitize Library 31 Mar 2024 · 1 repository · arXiv:2404.00699Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange 30 Mar 2024 · 1 repository · arXiv:2404.00344Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4 30 Mar 2024 · 1 repository · arXiv:2404.00484
-
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning 30 Mar 2024 · 0 repositories · arXiv:2404.00213
-
Small Language Models Learn Enhanced Reasoning Skills from Medical Textbooks 30 Mar 2024 · 0 repositories · arXiv:2404.00376
-
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning 29 Mar 2024 · 1 repository · arXiv:2403.19962Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Latxa: An Open Language Model and Evaluation Suite for Basque 29 Mar 2024 · 1 repository · arXiv:2403.20266
-
MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models 29 Mar 2024 · 1 repository · arXiv:2403.19913Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
On-the-fly Definition Augmentation of LLMs for Biomedical NER 29 Mar 2024 · 1 repository · arXiv:2404.00152Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
ReALM: Reference Resolution As Language Modeling 29 Mar 2024 · 0 repositories · arXiv:2403.20329
-
Checkpoint Merging via Bayesian Optimization in LLM Pretraining 28 Mar 2024 · 0 repositories · arXiv:2403.19390
-
Generating Multi-Aspect Queries for Conversational Search 28 Mar 2024 · 0 repositories · arXiv:2403.19302
-
MATEval: A Multi-Agent Discussion Framework for Advancing Open-Ended Text Evaluation 28 Mar 2024 · 1 repository · arXiv:2403.19305
-
BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text 27 Mar 2024 · 1 repository · arXiv:2403.18421Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models 27 Mar 2024 · 0 repositories · arXiv:2403.18365
-
Evaluating Large Language Models for Health-Related Text Classification Tasks with Public Social Media Data 27 Mar 2024 · 0 repositories · arXiv:2403.19031
-
Long-form factuality in large language models 27 Mar 2024 · 3 repositories · arXiv:2403.18802Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models 27 Mar 2024 · 2 repositories · arXiv:2403.18814Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Vulnerability Detection with Code Language Models: How Far Are We? 27 Mar 2024 · 1 repository · arXiv:2403.18624Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons 26 Mar 2024 · 1 repository · arXiv:2403.17760
-
ELLEN: Extremely Lightly Supervised Learning For Efficient Named Entity Recognition 26 Mar 2024 · 1 repository · arXiv:2403.17385
-
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models 26 Mar 2024 · 0 repositories · arXiv:2403.18093
-
InternLM2 Technical Report 26 Mar 2024 · 3 repositories · arXiv:2403.17297Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Large Language Models Are State-of-the-Art Evaluator for Grammatical Error Correction 26 Mar 2024 · 0 repositories · arXiv:2403.17540