Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 18
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 18 of 29: papers 1,701 to 1,800 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The Impact of Demonstrations on Multilingual In-Context Learning: A Multidimensional Analysis 20 Feb 2024 · 1 repository · arXiv:2402.12976
-
TofuEval: Evaluating Hallucinations of LLMs on Topic-Focused Dialogue Summarization 20 Feb 2024 · 1 repository · arXiv:2402.13249
-
A Critical Evaluation of AI Feedback for Aligning Large Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12366Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs 19 Feb 2024 · 1 repository · arXiv:2402.11753Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Creating a Fine Grained Entity Type Taxonomy Using LLMs 19 Feb 2024 · 0 repositories · arXiv:2402.12557
-
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models 19 Feb 2024 · 0 repositories · arXiv:2402.13291
-
Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models 19 Feb 2024 · 0 repositories · arXiv:2402.12147
-
Evaluation of ChatGPT's Smart Contract Auditing Capabilities Based on Chain of Thought 19 Feb 2024 · 0 repositories · arXiv:2402.12023
-
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations 19 Feb 2024 · 2 repositories · arXiv:2402.12348
-
IMBUE: Improving Interpersonal Effectiveness through Simulation and Just-in-time Feedback with Human-Language Model Interaction 19 Feb 2024 · 0 repositories · arXiv:2402.12556
-
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports 19 Feb 2024 · 0 repositories · arXiv:2402.12298
-
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking 19 Feb 2024 · 0 repositories · arXiv:2402.12146
-
Cofca: A Step-Wise Counterfactual Multi-hop QA benchmark 19 Feb 2024 · 0 repositories · arXiv:2402.11924
-
Robust CLIP: Unsupervised Adversarial Fine-Tuning of Vision Embeddings for Robust Large Vision-Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12336Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples)
-
Shallow Synthesis of Knowledge in GPT-Generated Texts: A Case Study in Automatic Related Work Composition 19 Feb 2024 · 0 repositories · arXiv:2402.12255
-
SPML: A DSL for Defending Language Models Against Prompt Attacks 19 Feb 2024 · 0 repositories · arXiv:2402.11755
-
Standardize: Aligning Language Models with Expert-Defined Standards for Content Generation 19 Feb 2024 · 1 repository · arXiv:2402.12593
-
Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One 19 Feb 2024 · 0 repositories · arXiv:2402.12150
-
Can Deception Detection Go Deeper? Dataset, Evaluation, and Benchmark for Deception Reasoning 18 Feb 2024 · 0 repositories · arXiv:2402.11432
-
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection 18 Feb 2024 · 0 repositories · arXiv:2402.11621
-
DictLLM: Harnessing Key-Value Data Structures with Large Language Models for Enhanced Medical Diagnostics 18 Feb 2024 · 0 repositories · arXiv:2402.11481
-
EventRL: Enhancing Event Extraction with Outcome Supervision for Large Language Models 18 Feb 2024 · 1 repository · arXiv:2402.11430
-
FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence 18 Feb 2024 · 1 repository · arXiv:2402.11456Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
KMMLU: Measuring Massive Multitask Language Understanding in Korean 18 Feb 2024 · 0 repositories · arXiv:2402.11548
-
Learning From Failure: Integrating Negative Examples when Fine-tuning Large Language Models as Agents 18 Feb 2024 · 1 repository · arXiv:2402.11651Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
LongAgent: Scaling Language Models to 128k Context through Multi-Agent Collaboration 18 Feb 2024 · 1 repository · arXiv:2402.11550
-
Multi-dimensional Evaluation of Empathetic Dialog Responses 18 Feb 2024 · 0 repositories · arXiv:2402.11409
-
Multi-Task Inference: Can Large Language Models Follow Multiple Instructions at Once? 18 Feb 2024 · 1 repository · arXiv:2402.11597Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Ploutos: Towards interpretable stock movement prediction with financial large language model 18 Feb 2024 · 0 repositories · arXiv:2403.00782
-
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning 18 Feb 2024 · 0 repositories · arXiv:2402.11690
-
Boosting of Thoughts: Trial-and-Error Problem Solving with Large Language Models 17 Feb 2024 · 2 repositories · arXiv:2402.11140
-
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges 17 Feb 2024 · 0 repositories · arXiv:2402.11203
-
GenDec: A robust generative Question-decomposition method for Multi-hop reasoning 17 Feb 2024 · 0 repositories · arXiv:2402.11166
-
Reasoning before Comparison: LLM-Enhanced Semantic Similarity Metrics for Domain Specialized Text Analysis 17 Feb 2024 · 0 repositories · arXiv:2402.11398
-
ZeroG: Investigating Cross-dataset Zero-shot Transferability in Graphs 17 Feb 2024 · 1 repository · arXiv:2402.11235Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Assessing the Reasoning Abilities of ChatGPT in the Context of Claim Verification 16 Feb 2024 · 0 repositories · arXiv:2402.10735
-
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements 16 Feb 2024 · 1 repository · arXiv:2402.10614Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Can Separators Improve Chain-of-Thought Prompting? 16 Feb 2024 · 0 repositories · arXiv:2402.10645
-
Emoji Driven Crypto Assets Market Reactions 16 Feb 2024 · 0 repositories · arXiv:2402.10481
-
FinTral: A Family of GPT-4 Level Multimodal Financial Large Language Models 16 Feb 2024 · 0 repositories · arXiv:2402.10986
-
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data 16 Feb 2024 · 1 repository · arXiv:2402.10675
-
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs? 16 Feb 2024 · 0 repositories · arXiv:2402.10770
-
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss 16 Feb 2024 · 2 repositories · arXiv:2402.10790
-
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers 16 Feb 2024 · 1 repository · arXiv:2402.10601Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling 16 Feb 2024 · 1 repository · arXiv:2402.10466Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives 16 Feb 2024 · 0 repositories · arXiv:2402.11051
-
LinkNER: Linking Local Named Entity Recognition Models to Large Language Models using Uncertainty 16 Feb 2024 · 1 repository · arXiv:2402.10573
-
ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages 16 Feb 2024 · 1 repository · arXiv:2402.10753
-
A StrongREJECT for Empty Jailbreaks 15 Feb 2024 · 2 repositories · arXiv:2402.10260Syntology official: harvested, nothing ran · 0 ran · 4 unverified (of 4 harvested samples)
-
An Analysis of Language Frequency and Error Correction for Esperanto 15 Feb 2024 · 0 repositories · arXiv:2402.09696
-
Data Engineering for Scaling Language Models to 128K Context 15 Feb 2024 · 1 repository · arXiv:2402.10171Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4 15 Feb 2024 · 0 repositories · arXiv:2402.10083
-
GPT-4's assessment of its performance in a USMLE-based case study 15 Feb 2024 · 0 repositories · arXiv:2402.09654
-
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset 15 Feb 2024 · 1 repository · arXiv:2402.10176
-
Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models 15 Feb 2024 · 0 repositories · arXiv:2402.10353
-
Unlocking Structure Measuring: Introducing PDD, an Automatic Metric for Positional Discourse Coherence 15 Feb 2024 · 1 repository · arXiv:2402.10175
-
X-lifecycle Learning for Cloud Incident Management using LLMs 15 Feb 2024 · 0 repositories · arXiv:2404.03662
-
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation 14 Feb 2024 · 1 repository · arXiv:2402.09615
-
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability 14 Feb 2024 · 1 repository · arXiv:2402.09404
-
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects 14 Feb 2024 · 0 repositories · arXiv:2402.09052
-
Leveraging Large Language Models for Enhanced NLP Task Performance through Knowledge Distillation and Optimized Training Strategies 14 Feb 2024 · 0 repositories · arXiv:2402.09282
-
LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset 14 Feb 2024 · 1 repository · arXiv:2402.09391Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
AutoTutor meets Large Language Models: A Language Model Tutor with Rich Pedagogy and Guardrails 14 Feb 2024 · 1 repository · arXiv:2402.09216Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models 13 Feb 2024 · 1 repository · arXiv:2402.08219Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Combining Insights From Multiple Large Language Models Improves Diagnostic Accuracy 13 Feb 2024 · 0 repositories · arXiv:2402.08806
-
eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data 13 Feb 2024 · 1 repository · arXiv:2402.08831Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
InstructGraph: Boosting Large Language Models via Graph-centric Instruction Tuning and Preference Alignment 13 Feb 2024 · 1 repository · arXiv:2402.08785
-
Large Language Models for the Automated Analysis of Optimization Algorithms 13 Feb 2024 · 1 repository · arXiv:2402.08472
-
LLaGA: Large Language and Graph Assistant 13 Feb 2024 · 2 repositories · arXiv:2402.08170Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling 13 Feb 2024 · 1 repository · arXiv:2402.08702Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
The Last JITAI? Exploring Large Language Models for Issuing Just-in-Time Adaptive Interventions: Fostering Physical Activity in a Conceptual Cardiac Rehabilitation Setting 13 Feb 2024 · 0 repositories · arXiv:2402.08658
-
Addressing cognitive bias in medical language models 12 Feb 2024 · 1 repository · arXiv:2402.08113Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension 12 Feb 2024 · 1 repository · arXiv:2402.07729
-
Dólares or Dollars? Unraveling the Bilingual Prowess of Financial LLMs Between Spanish and English 12 Feb 2024 · 1 repository · arXiv:2402.07405
-
Enhancing Multi-Criteria Decision Analysis with AI: Integrating Analytic Hierarchy Process and GPT-4 for Automated Decision Support 12 Feb 2024 · 0 repositories · arXiv:2402.07404
-
Enhancing Programming Error Messages in Real Time with Generative AI 12 Feb 2024 · 0 repositories · arXiv:2402.08072
-
Large Language Models "Ad Referendum": How Good Are They at Machine Translation in the Legal Domain? 12 Feb 2024 · 0 repositories · arXiv:2402.07681
-
Large Language Models are Few-shot Generators: Proposing Hybrid Prompt Algorithm To Generate Webshell Escape Samples 12 Feb 2024 · 1 repository · arXiv:2402.07408
-
Leveraging AI to Advance Science and Computing Education across Africa: Challenges, Progress and Opportunities 12 Feb 2024 · 0 repositories · arXiv:2402.07397
-
On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks 12 Feb 2024 · 0 repositories · arXiv:2402.08115
-
Secret Collusion among Generative AI Agents: Multi-Agent Deception via Steganography 12 Feb 2024 · 0 repositories · arXiv:2402.07510
-
Suppressing Pink Elephants with Direct Principle Feedback 12 Feb 2024 · 0 repositories · arXiv:2402.07896
-
VCR: Video representation for Contextual Retrieval 12 Feb 2024 · 1 repository · arXiv:2402.07466
-
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness? 11 Feb 2024 · 0 repositories · arXiv:2402.07282
-
Natural Language Reinforcement Learning 11 Feb 2024 · 0 repositories · arXiv:2402.07157
-
ChemLLM: A Chemical Large Language Model 10 Feb 2024 · 1 repository · arXiv:2402.06852
-
Gemini Goes to Med School: Exploring the Capabilities of Multimodal Large Language Models on Medical Challenge Problems & Hallucinations 10 Feb 2024 · 1 repository · arXiv:2402.07023
-
OpenFedLLM: Training Large Language Models on Decentralized Private Data via Federated Learning 10 Feb 2024 · 3 repositories · arXiv:2402.06954
-
UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction 10 Feb 2024 · 1 repository · arXiv:2402.06861Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Bryndza at ClimateActivism 2024: Stance, Target and Hate Event Detection via Retrieval-Augmented GPT-4 and LLaMA 9 Feb 2024 · 2 repositories · arXiv:2402.06549
-
CultureLLM: Incorporating Cultural Differences into Large Language Models 9 Feb 2024 · 2 repositories · arXiv:2402.10946Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LLaVA-Docent: Instruction Tuning with Multimodal Large Language Model to Support Art Appreciation Education 9 Feb 2024 · 0 repositories · arXiv:2402.06264
-
RareBench: Can LLMs Serve as Rare Diseases Specialists? 9 Feb 2024 · 1 repository · arXiv:2402.06341
-
ResumeFlow: An LLM-facilitated Pipeline for Personalized Resume Generation and Refinement 9 Feb 2024 · 1 repository · arXiv:2402.06221Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
A Prompt Response to the Demand for Automatic Gender-Neutral Translation 8 Feb 2024 · 1 repository · arXiv:2402.06041
-
GPT-4 Generated Narratives of Life Events using a Structured Narrative Prompt: A Validation Study 8 Feb 2024 · 0 repositories · arXiv:2402.05435
-
How Well Can LLMs Negotiate? NegotiationArena Platform and Analysis 8 Feb 2024 · 1 repository · arXiv:2402.05863Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
In-Context Principle Learning from Mistakes 8 Feb 2024 · 1 repository · arXiv:2402.05403
-
Large Language Models for Psycholinguistic Plausibility Pretesting 8 Feb 2024 · 0 repositories · arXiv:2402.05455
-
Limits of Transformer Language Models on Learning to Compose Algorithms 8 Feb 2024 · 1 repository · arXiv:2402.05785Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 8 harvested samples)