Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 21
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 21 of 29: papers 2,001 to 2,100 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Large Language Model (LLM) Bias Index -- LLMBI 22 Dec 2023 · 0 repositories · arXiv:2312.14769
-
MMGPL: Multimodal Medical Data Analysis with Graph Prompt Learning 22 Dec 2023 · 0 repositories · arXiv:2312.14574
-
Towards a Unified Multimodal Reasoning Framework 22 Dec 2023 · 1 repository · arXiv:2312.15021
-
Voila-A: Aligning Vision-Language Models with User's Gaze Attention 22 Dec 2023 · 0 repositories · arXiv:2401.09454
-
Exploiting Novel GPT-4 APIs 21 Dec 2023 · 1 repository · arXiv:2312.14302
-
LingoQA: Visual Question Answering for Autonomous Driving 21 Dec 2023 · 2 repositories · arXiv:2312.14115Syntology community repositories only · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Shai: A large language model for asset management 21 Dec 2023 · 0 repositories · arXiv:2312.14203
-
Preparing to Integrate Generative Pretrained Transformer Series 4 models into Genetic Variant Assessment Workflows: Assessing Performance, Drift, and Nondeterminism Characteristics Relative to Classifying Functional Evidence in Literature 21 Dec 2023 · 0 repositories · arXiv:2312.13521
-
Benchmarking and Analyzing In-context Learning, Fine-tuning and Supervised Learning for Biomedical Knowledge Curation: a focused study on chemical entities of biological interest 20 Dec 2023 · 0 repositories · arXiv:2312.12989
-
Scaling Down to Scale Up: A Cost-Benefit Analysis of Replacing OpenAI's LLM with Open Source SLMs in Production 20 Dec 2023 · 1 repository · arXiv:2312.14972
-
Large Language Models in Medical Term Classification and Unexpected Misalignment Between Response and Reasoning 19 Dec 2023 · 0 repositories · arXiv:2312.14184
-
Evaluating and Enhancing Large Language Models for Conversational Reasoning on Knowledge Graphs 18 Dec 2023 · 1 repository · arXiv:2312.11282
-
MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL 18 Dec 2023 · 1 repository · arXiv:2312.11242Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation 18 Dec 2023 · 1 repository · arXiv:2312.11361Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
An Evaluation of GPT-4V and Gemini in Online VQA 17 Dec 2023 · 0 repositories · arXiv:2312.10637
-
CEIR: Concept-based Explainable Image Representation Learning 17 Dec 2023 · 0 repositories · arXiv:2312.10747
-
A Comparative Analysis of Large Language Models for Code Documentation Generation 16 Dec 2023 · 0 repositories · arXiv:2312.10349
-
DeepArt: A Benchmark to Advance Fidelity Research in AI-Generated Content 16 Dec 2023 · 0 repositories · arXiv:2312.10407
-
RecPrompt: A Self-tuning Prompting Framework for News Recommendation Using Large Language Models 16 Dec 2023 · 1 repository · arXiv:2312.10463
-
Binary Code Summarization: Benchmarking ChatGPT/GPT-4 and Other Large Language Models 15 Dec 2023 · 1 repository · arXiv:2312.09601
-
Distilling Large Language Models for Matching Patients to Clinical Trials 15 Dec 2023 · 0 repositories · arXiv:2312.09958
-
Integrating AI and Learning Analytics for Data-Driven Pedagogical Decisions and Personalized Interventions in Education 15 Dec 2023 · 0 repositories · arXiv:2312.09548
-
OpenMedCalc: Augmentation of ChatGPT with Clinician-Informed Tools Improves Performance on Medical Calculation Tasks 15 Dec 2023 · 1 repository
-
Heterogeneous Graph Neural Architecture Search with GPT-4 14 Dec 2023 · 1 repository · arXiv:2312.08680
-
Holodeck: Language Guided Generation of 3D Embodied AI Environments 14 Dec 2023 · 1 repository · arXiv:2312.09067Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Modeling Complex Mathematical Reasoning via Large Language Model based MathAgent 14 Dec 2023 · 1 repository · arXiv:2312.08926
-
Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision 14 Dec 2023 · 0 repositories · arXiv:2312.09390
-
Assessing GPT4-V on Structured Reasoning Tasks 13 Dec 2023 · 0 repositories · arXiv:2312.11524
-
Beyond English: Evaluating LLMs for Arabic Grammatical Error Correction 13 Dec 2023 · 0 repositories · arXiv:2312.08400
-
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers 13 Dec 2023 · 2 repositories · arXiv:2312.08168Syntology official (archive's flag): 6 ran · 11 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 5 pointer-only (licence)
-
CoIE: Chain-of-Instruct Editing for Multi-Attribute Face Manipulation 13 Dec 2023 · 0 repositories · arXiv:2312.07879
-
High-throughput Biomedical Relation Extraction for Semi-Structured Web Articles Empowered by Large Language Models 13 Dec 2023 · 0 repositories · arXiv:2312.08274
-
Native Language Identification with Large Language Models 13 Dec 2023 · 0 repositories · arXiv:2312.07819
-
Prompt Engineering-assisted Malware Dynamic Analysis Using GPT-4 13 Dec 2023 · 1 repository · arXiv:2312.08317
-
AI Control: Improving Safety Despite Intentional Subversion 12 Dec 2023 · 1 repository · arXiv:2312.06942Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Context Matters: Data-Efficient Augmentation of Large Language Models for Scientific Applications 12 Dec 2023 · 2 repositories · arXiv:2312.07069
-
Exploring Large Language Models to Facilitate Variable Autonomy for Human-Robot Teaming 12 Dec 2023 · 0 repositories · arXiv:2312.07214
-
Large Foundation Models for Power Systems 12 Dec 2023 · 1 repository · arXiv:2312.07044
-
LLMEval: A Preliminary Study on How to Evaluate Large Language Models 12 Dec 2023 · 0 repositories · arXiv:2312.07398
-
Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack 12 Dec 2023 · 1 repository · arXiv:2312.06924Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Audio-Visual LLM for Video Understanding 11 Dec 2023 · 0 repositories · arXiv:2312.06720
-
Genixer: Empowering Multimodal Large Language Models as a Powerful Data Generator 11 Dec 2023 · 1 repository · arXiv:2312.06731Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
GPTBIAS: A Comprehensive Framework for Evaluating Bias in Large Language Models 11 Dec 2023 · 0 repositories · arXiv:2312.06315
-
GTA: Gated Toxicity Avoidance for LM Performance Preservation 11 Dec 2023 · 1 repository · arXiv:2312.06122
-
Interactive Planning Using Large Language Models for Partially Observable Robotics Tasks 11 Dec 2023 · 0 repositories · arXiv:2312.06876
-
KnowGPT: Knowledge Graph based Prompting for Large Language Models 11 Dec 2023 · 0 repositories · arXiv:2312.06185
-
Context Tuning for Retrieval Augmented Generation 9 Dec 2023 · 0 repositories · arXiv:2312.05708
-
GPT-4 and Safety Case Generation: An Exploratory Analysis 9 Dec 2023 · 0 repositories · arXiv:2312.05696
-
Sim-GPT: Text Similarity via GPT Annotated Data 9 Dec 2023 · 1 repository · arXiv:2312.05603
-
Learning to Break: Knowledge-Enhanced Reasoning in Multi-Agent Debate System 8 Dec 2023 · 2 repositories · arXiv:2312.04854
-
Exploring the Limits of ChatGPT in Software Security Applications 8 Dec 2023 · 0 repositories · arXiv:2312.05275
-
KwaiAgents: Generalized Information-seeking Agent System with Large Language Models 8 Dec 2023 · 1 repository · arXiv:2312.04889Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
PixLore: A Dataset-driven Approach to Rich Image Captioning 8 Dec 2023 · 1 repository · arXiv:2312.05349
-
Cost-Effective In-Context Learning for Entity Resolution: A Design Space Exploration 7 Dec 2023 · 1 repository · arXiv:2312.03987
-
Fortify the Shortest Stave in Attention: Enhancing Context Awareness of Large Language Models for Effective Tool Use 7 Dec 2023 · 1 repository · arXiv:2312.04455Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition 7 Dec 2023 · 1 repository · arXiv:2312.04293Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives 7 Dec 2023 · 1 repository · arXiv:2312.04314
-
LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs 7 Dec 2023 · 1 repository · arXiv:2312.04372
-
On Sarcasm Detection with OpenAI GPT-based Models 7 Dec 2023 · 0 repositories · arXiv:2312.04642
-
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos 7 Dec 2023 · 2 repositories · arXiv:2312.04746Syntology 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples)
-
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models 6 Dec 2023 · 1 repository · arXiv:2312.03543
-
XAIQA: Explainer-Based Data Augmentation for Extractive Question Answering 6 Dec 2023 · 0 repositories · arXiv:2312.03567
-
A Comparative Study of AI-Generated (GPT-4) and Human-crafted MCQs in Programming Education 5 Dec 2023 · 0 repositories · arXiv:2312.03173
-
GPT vs Human for Scientific Reviews: A Dual Source Review on Applications of ChatGPT in Science 5 Dec 2023 · 0 repositories · arXiv:2312.03769
-
Let the LLMs Talk: Simulating Human-to-Human Conversational QA via Zero-Shot LLM-to-LLM Interactions 5 Dec 2023 · 1 repository · arXiv:2312.02913
-
Rank-without-GPT: Building GPT-Independent Listwise Rerankers on Open-Source Large Language Models 5 Dec 2023 · 0 repositories · arXiv:2312.02969
-
RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze! 5 Dec 2023 · 2 repositories · arXiv:2312.02724
-
Competition-Level Problems are Effective LLM Evaluators 4 Dec 2023 · 0 repositories · arXiv:2312.02143
-
Explore, Select, Derive, and Recall: Augmenting LLM with Human-like Memory for Mobile Task Automation 4 Dec 2023 · 0 repositories · arXiv:2312.03003
-
Fine-Tuning Language Models for Context-Specific SQL Query Generation 4 Dec 2023 · 0 repositories · arXiv:2312.02251
-
InstructTA: Instruction-Tuned Targeted Attack for Large Vision-Language Models 4 Dec 2023 · 1 repository · arXiv:2312.01886Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Retrieval-augmented Multi-modal Chain-of-Thoughts Reasoning for Large Language Models 4 Dec 2023 · 0 repositories · arXiv:2312.01714
-
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically 4 Dec 2023 · 2 repositories · arXiv:2312.02119Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
D-Bot: Database Diagnosis System using Large Language Models 3 Dec 2023 · 1 repository · arXiv:2312.01454Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Axiomatic Preference Modeling for Longform Question Answering 2 Dec 2023 · 0 repositories · arXiv:2312.02206
-
From Voices to Validity: Leveraging Large Language Models (LLMs) for Textual Analysis of Policy Stakeholder Interviews 2 Dec 2023 · 0 repositories · arXiv:2312.01202
-
A Bayesian approach for prompt optimization in pre-trained language models 1 Dec 2023 · 0 repositories · arXiv:2312.00471
-
Applying Large Language Models and Chain-of-Thought for Automatic Scoring 30 Nov 2023 · 0 repositories · arXiv:2312.03748
-
CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation 30 Nov 2023 · 2 repositories · arXiv:2311.18702Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Unnatural Error Correction: GPT-4 Can Almost Perfectly Handle Unnatural Scrambled Text 30 Nov 2023 · 1 repository · arXiv:2311.18805Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
AviationGPT: A Large Language Model for the Aviation Domain 29 Nov 2023 · 0 repositories · arXiv:2311.17686
-
Biomedical knowledge graph-optimized prompt generation for large language models 29 Nov 2023 · 1 repository · arXiv:2311.17330Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples)
-
M²Chat: Empowering VLM for Multimodal LLM Interleaved Text-Image Generation 29 Nov 2023 · 1 repository · arXiv:2311.17963
-
Grounding Foundation Models through Federated Transfer Learning: A General Framework 29 Nov 2023 · 0 repositories · arXiv:2311.17431
-
MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning 29 Nov 2023 · 0 repositories · arXiv:2311.17435
-
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models 29 Nov 2023 · 1 repository · arXiv:2311.17667
-
Can Generalist Foundation Models Outcompete Special-Purpose Tuning? Case Study in Medicine 28 Nov 2023 · 2 repositories · arXiv:2311.16452Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design 28 Nov 2023 · 0 repositories · arXiv:2311.16974
-
General-Purpose vs. Domain-Adapted Large Language Models for Extraction of Structured Data from Chest Radiology Reports 28 Nov 2023 · 0 repositories · arXiv:2311.17213
-
Positioning Political Texts with Large Language Models by Asking and Averaging 28 Nov 2023 · 0 repositories · arXiv:2311.16639
-
The Falcon Series of Open Language Models 28 Nov 2023 · 0 repositories · arXiv:2311.16867
-
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models 27 Nov 2023 · 1 repository · arXiv:2311.15596
-
ChartLlama: A Multimodal LLM for Chart Understanding and Generation 27 Nov 2023 · 0 repositories · arXiv:2311.16483
-
Decoding Logic Errors: A Comparative Study on Bug Detection by Students and Large Language Models 27 Nov 2023 · 0 repositories · arXiv:2311.16017
-
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition? 27 Nov 2023 · 2 repositories · arXiv:2311.15732Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Instruct2Attack: Language-Guided Semantic Adversarial Attacks 27 Nov 2023 · 0 repositories · arXiv:2311.15551
-
MEDITRON-70B: Scaling Medical Pretraining for Large Language Models 27 Nov 2023 · 1 repository · arXiv:2311.16079Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 14 harvested samples)
-
Towards Vision Enhancing LLMs: Empowering Multimodal Knowledge Storage and Sharing in LLMs 27 Nov 2023 · 0 repositories · arXiv:2311.15759
-
Comparative Analysis of ChatGPT, GPT-4, and Microsoft Bing Chatbots for GRE Test 26 Nov 2023 · 0 repositories · arXiv:2312.03719
-
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering 25 Nov 2023 · 1 repository · arXiv:2311.14906