Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 22
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 22 of 29: papers 2,101 to 2,200 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
From Text to Image: Exploring GPT-4Vision's Potential in Advanced Radiological Analysis across Subspecialties 24 Nov 2023 · 0 repositories · arXiv:2311.14777
-
Cultural Bias and Cultural Alignment of Large Language Models 23 Nov 2023 · 0 repositories · arXiv:2311.14096
-
Evaluating GPT-4's Vision Capabilities on Brazilian University Admission Exams 23 Nov 2023 · 1 repository · arXiv:2311.14169
-
Surpassing GPT-4 Medical Coding with a Two-Stage Approach 22 Nov 2023 · 0 repositories · arXiv:2311.13735
-
Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs 22 Nov 2023 · 1 repository · arXiv:2311.13194Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
@ve: A Chatbot for Latin 22 Nov 2023 · 0 repositories · arXiv:2311.14741
-
GeoLocator: a location-integrated large multimodal model for inferring geo-privacy 21 Nov 2023 · 0 repositories · arXiv:2311.13018
-
From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.13063
-
GAIA: a benchmark for General AI Assistants 21 Nov 2023 · 2 repositories · arXiv:2311.12983
-
GPT4Motion: Scripting Physical Motions in Text-to-Video Generation via Blender-Oriented GPT Planning 21 Nov 2023 · 0 repositories · arXiv:2311.12631
-
Oasis: Data Curation and Assessment System for Pretraining of Large Language Models 21 Nov 2023 · 1 repository · arXiv:2311.12537
-
Evil Geniuses: Delving into the Safety of LLM-based Agents 20 Nov 2023 · 1 repository · arXiv:2311.11855Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Generating Valid and Natural Adversarial Examples with Large Language Models 20 Nov 2023 · 0 repositories · arXiv:2311.11861
-
GPQA: A Graduate-Level Google-Proof Q&A Benchmark 20 Nov 2023 · 3 repositories · arXiv:2311.12022Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples)
-
Towards Human-Level Text Coding with LLMs: The Case of Fatherhood Roles in Public Policy Documents 20 Nov 2023 · 1 repository · arXiv:2311.11844
-
Meta Prompting for AI Systems 20 Nov 2023 · 1 repository · arXiv:2311.11482Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Which AI Technique Is Better to Classify Requirements? An Experiment with SVM, LSTM, and ChatGPT 20 Nov 2023 · 1 repository · arXiv:2311.11547
-
Behavior Optimized Image Generation 18 Nov 2023 · 0 repositories · arXiv:2311.10995
-
Visual AI and Linguistic Intelligence Through Steerability and Composability 18 Nov 2023 · 0 repositories · arXiv:2312.12383
-
EduQuick: A Dataset Toward Evaluating Summarization of Informal Educational Content for Social Media 17 Nov 2023 · 0 repositories
-
TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in LLMs through Translation-Assisted Chain-of-Thought Processes 17 Nov 2023 · 1 repository · arXiv:2311.10797
-
BLT: Can Large Language Models Handle Basic Legal Text? 16 Nov 2023 · 1 repository · arXiv:2311.09693
-
Prompt-based Pseudo-labeling Strategy for Sample-Efficient Semi-Supervised Extractive Summarization 16 Nov 2023 · 0 repositories · arXiv:2311.09559
-
GEE! Grammar Error Explanation with Large Language Models 16 Nov 2023 · 1 repository · arXiv:2311.09517
-
HuatuoGPT-II, One-stage Training for Medical Adaption of LLMs 16 Nov 2023 · 1 repository · arXiv:2311.09774Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Human Still Wins over LLM: An Empirical Study of Active Learning on Domain-Specific Annotation Tasks 16 Nov 2023 · 0 repositories · arXiv:2311.09825
-
Investigating Data Contamination in Modern Benchmarks for Large Language Models 16 Nov 2023 · 0 repositories · arXiv:2311.09783
-
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains 16 Nov 2023 · 1 repository · arXiv:2311.09797Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Large Language Models for Propaganda Span Annotation 16 Nov 2023 · 1 repository · arXiv:2311.09812
-
ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code 16 Nov 2023 · 1 repository · arXiv:2311.09835Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
On Evaluating the Integration of Reasoning and Action in LLM Agents with Database Question Answering 16 Nov 2023 · 0 repositories · arXiv:2311.09721
-
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology 16 Nov 2023 · 0 repositories · arXiv:2311.09861
-
Self-Contradictory Reasoning Evaluation and Detection 16 Nov 2023 · 1 repository · arXiv:2311.09603Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Structured Chemistry Reasoning with Large Language Models 16 Nov 2023 · 1 repository · arXiv:2311.09656Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Towards Autonomous Hypothesis Verification via Language Models with Minimal Guidance 16 Nov 2023 · 0 repositories · arXiv:2311.09706
-
UnifiedVisionGPT: Streamlining Vision-Oriented AI through Generalized Multimodal Framework 16 Nov 2023 · 1 repository · arXiv:2311.10125
-
Can Large Language Models Follow Concept Annotation Guidelines? A Case Study on Scientific and Financial Domains 15 Nov 2023 · 1 repository · arXiv:2311.08704
-
Enhancing Machine Translation through Advanced In-Context Learning: A Methodological Strategy for GPT-4 Improvement 15 Nov 2023 · 0 repositories · arXiv:2311.10765
-
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers 15 Nov 2023 · 2 repositories · arXiv:2311.09000Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
GENEVA: GENErating and Visualizing branching narratives using LLMs 15 Nov 2023 · 0 repositories · arXiv:2311.09213
-
I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots 15 Nov 2023 · 1 repository · arXiv:2311.08957
-
Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts 15 Nov 2023 · 0 repositories · arXiv:2311.09127
-
Llamas Know What GPTs Don't Show: Surrogate Models for Confidence Estimation 15 Nov 2023 · 0 repositories · arXiv:2311.08877
-
MELA: Multilingual Evaluation of Linguistic Acceptability 15 Nov 2023 · 1 repository · arXiv:2311.09033
-
Safer-Instruct: Aligning Language Models with Automated Preference Data 15 Nov 2023 · 1 repository · arXiv:2311.08685
-
ToolTalk: Evaluating Tool-Usage in a Conversational Setting 15 Nov 2023 · 1 repository · arXiv:2311.10775Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
X-Eval: Generalizable Multi-aspect Text Evaluation via Augmented Instruction Tuning with Auxiliary Evaluation Aspects 15 Nov 2023 · 0 repositories · arXiv:2311.08788
-
A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily 14 Nov 2023 · 1 repository · arXiv:2311.08268Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
Automated title and abstract screening for scoping reviews using the GPT-4 Large Language Model 14 Nov 2023 · 1 repository · arXiv:2311.07918
-
Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks 14 Nov 2023 · 0 repositories · arXiv:2311.09247
-
Evaluating LLMs on Document-Based QA: Exact Answer Selection and Numerical Extraction using Cogtale dataset 14 Nov 2023 · 0 repositories · arXiv:2311.07878
-
How good are Large Language Models on African Languages? 14 Nov 2023 · 0 repositories · arXiv:2311.07978
-
MAgIC: Investigation of Large Language Model Powered Multi-Agent in Cognition, Adaptability, Rationality and Collaboration 14 Nov 2023 · 1 repository · arXiv:2311.08562Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models 14 Nov 2023 · 0 repositories · arXiv:2311.08370
-
A Benchmark to Understand the Role of Knowledge Graphs on Large Language Model's Accuracy for Question Answering on Enterprise SQL Databases 13 Nov 2023 · 0 repositories · arXiv:2311.07509
-
Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study 13 Nov 2023 · 1 repository · arXiv:2311.07387Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
LM-Polygraph: Uncertainty Estimation for Language Models 13 Nov 2023 · 0 repositories · arXiv:2311.07383
-
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks 13 Nov 2023 · 0 repositories · arXiv:2311.07463
-
Speech-based Slot Filling using Large Language Models 13 Nov 2023 · 0 repositories · arXiv:2311.07418
-
The Impact of Large Language Models on Scientific Discovery: a Preliminary Study using GPT-4 13 Nov 2023 · 0 repositories · arXiv:2311.07361
-
VerityMath: Advancing Mathematical Reasoning by Self-Verification Through Unit Consistency 13 Nov 2023 · 1 repository · arXiv:2311.07172Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Detecting and Correcting Hate Speech in Multimodal Memes with Large Visual Language Model 12 Nov 2023 · 0 repositories · arXiv:2311.06737
-
Evaluation of GPT-4 for chest X-ray impression generation: A reader study on performance and perception 12 Nov 2023 · 0 repositories · arXiv:2311.06815
-
Flames: Benchmarking Value Alignment of LLMs in Chinese 12 Nov 2023 · 1 repository · arXiv:2311.06899
-
Large Language Models' Understanding of Math: Source Criticism and Extrapolation 12 Nov 2023 · 0 repositories · arXiv:2311.07618
-
Intentional Biases in LLM Responses 11 Nov 2023 · 0 repositories · arXiv:2311.07611
-
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models 10 Nov 2023 · 2 repositories · arXiv:2311.06233Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model 10 Nov 2023 · 0 repositories · arXiv:2311.07594
-
Language Models can be Logical Solvers 10 Nov 2023 · 0 repositories · arXiv:2311.06158
-
Making LLMs Worth Every Penny: Resource-Limited Text Classification in Banking 10 Nov 2023 · 0 repositories · arXiv:2311.06102
-
Conic10K: A Challenging Math Problem Understanding and Reasoning Dataset 9 Nov 2023 · 1 repository · arXiv:2311.05113
-
Challenging the Validity of Personality Tests for Large Language Models 9 Nov 2023 · 0 repositories · arXiv:2311.05297
-
Large Language Models can Strategically Deceive their Users when Put Under Pressure 9 Nov 2023 · 1 repository · arXiv:2311.07590Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples 8 Nov 2023 · 1 repository · arXiv:2311.04850Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Black-Box Prompt Optimization: Aligning Large Language Models without Model Training 7 Nov 2023 · 1 repository · arXiv:2311.04155Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Evaluating Large Language Models in Ophthalmology 7 Nov 2023 · 0 repositories · arXiv:2311.04933
-
Evaluating multiple large language models in pediatric ophthalmology 7 Nov 2023 · 0 repositories · arXiv:2311.04368
-
Identifying and Mitigating Vulnerabilities in LLM-Integrated Applications 7 Nov 2023 · 0 repositories · arXiv:2311.16153
-
Leveraging Large Language Models for Automated Proof Synthesis in Rust 7 Nov 2023 · 0 repositories · arXiv:2311.03739
-
Which is better? Exploring Prompting Strategy For LLM-based Metrics 7 Nov 2023 · 1 repository · arXiv:2311.03754
-
Can LLMs Follow Simple Rules? 6 Nov 2023 · 1 repository · arXiv:2311.04235Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
DeepInception: Hypnotize Large Language Model to Be Jailbreaker 6 Nov 2023 · 1 repository · arXiv:2311.03191Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Nexus at ArAIEval Shared Task: Fine-Tuning Arabic Language Models for Propaganda and Disinformation Detection 6 Nov 2023 · 0 repositories · arXiv:2311.03184
-
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation 6 Nov 2023 · 0 repositories · arXiv:2311.03348
-
Evaluating the Potential of Leading Large Language Models in Reasoning Biology Questions 5 Nov 2023 · 0 repositories · arXiv:2311.07582
-
GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection 5 Nov 2023 · 1 repository · arXiv:2311.02612
-
FloodBrain: Flood Disaster Reporting by Web-based Retrieval Augmented Generation with an LLM 5 Nov 2023 · 0 repositories · arXiv:2311.02597
-
MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning 4 Nov 2023 · 1 repository · arXiv:2311.02303
-
DialogBench: Evaluating LLMs as Human-like Dialogue Systems 3 Nov 2023 · 1 repository · arXiv:2311.01677
-
PPTC Benchmark: Evaluating Large Language Models for PowerPoint Task Completion 3 Nov 2023 · 1 repository · arXiv:2311.01767Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples)
-
The risks of risk-based AI regulation: taking liability seriously 3 Nov 2023 · 0 repositories · arXiv:2311.14684
-
Generative Input: Towards Next-Generation Input Methods Paradigm 2 Nov 2023 · 0 repositories · arXiv:2311.01166
-
Are Large Language Models Reliable Judges? A Study on the Factuality Evaluation Capabilities of LLMs 1 Nov 2023 · 0 repositories · arXiv:2311.00681
-
Can Large Language Models Capture Public Opinion about Global Warming? An Empirical Assessment of Algorithmic Fidelity and Bias 1 Nov 2023 · 0 repositories · arXiv:2311.00217
-
From Text to Structure: Using Large Language Models to Support the Development of Legal Expert Systems 1 Nov 2023 · 1 repository · arXiv:2311.04911
-
ChipNeMo: Domain-Adapted LLMs for Chip Design 31 Oct 2023 · 0 repositories · arXiv:2311.00176
-
Does GPT-4 pass the Turing test? 31 Oct 2023 · 0 repositories · arXiv:2310.20216
-
Efficient Classification of Student Help Requests in Programming Courses Using Large Language Models 31 Oct 2023 · 0 repositories · arXiv:2310.20105
-
Learning From Mistakes Makes LLM Better Reasoner 31 Oct 2023 · 1 repository · arXiv:2310.20689Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
BioInstruct: Instruction Tuning of Large Language Models for Biomedical Natural Language Processing 30 Oct 2023 · 0 repositories · arXiv:2310.19975