Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 23
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 23 of 29: papers 2,201 to 2,300 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Building Real-World Meeting Summarization Systems using Large Language Models: A Practical Perspective 30 Oct 2023 · 0 repositories · arXiv:2310.19233
-
Constituency Parsing using LLMs 30 Oct 2023 · 0 repositories · arXiv:2310.19462
-
Dynamics of Instruction Tuning: Each Ability of Large Language Models Has Its Own Growth Pace 30 Oct 2023 · 1 repository · arXiv:2310.19651
-
Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck 30 Oct 2023 · 1 repository · arXiv:2310.19660
-
Multimodal ChatGPT for Medical Applications: an Experimental Study of GPT-4V 29 Oct 2023 · 1 repository · arXiv:2310.19061
-
Using Large Language Models to Support Thematic Analysis in Empirical Legal Studies 28 Oct 2023 · 0 repositories · arXiv:2310.18729
-
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory 27 Oct 2023 · 1 repository · arXiv:2310.17884
-
GPT-4 Vision on Medical Image Classification -- A Case Study on COVID-19 Dataset 27 Oct 2023 · 0 repositories · arXiv:2310.18498
-
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method 27 Oct 2023 · 0 repositories · arXiv:2310.17918
-
Large language models for aspect-based sentiment analysis 27 Oct 2023 · 1 repository · arXiv:2310.18025Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SOUL: Towards Sentiment and Opinion Understanding of Language 27 Oct 2023 · 1 repository · arXiv:2310.17924Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
A Framework for Automated Measurement of Responsible AI Harms in Generative AI Applications 26 Oct 2023 · 0 repositories · arXiv:2310.17750
-
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages 26 Oct 2023 · 0 repositories · arXiv:2310.17526
-
Can LLMs Grade Short-Answer Reading Comprehension Questions : An Empirical Study with a Novel Dataset 26 Oct 2023 · 0 repositories · arXiv:2310.18373
-
CompeteAI: Understanding the Competition Dynamics in Large Language Model-based Agents 26 Oct 2023 · 1 repository · arXiv:2310.17512Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 5 harvested samples)
-
Cultural Adaptation of Recipes 26 Oct 2023 · 0 repositories · arXiv:2310.17353
-
Is Explanation the Cure? Misinformation Mitigation in the Short Term and Long Term 26 Oct 2023 · 0 repositories · arXiv:2310.17711
-
Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models 26 Oct 2023 · 0 repositories · arXiv:2310.17567
-
"You Are An Expert Linguistic Annotator": Limits of LLMs as Analyzers of Abstract Meaning Representation 26 Oct 2023 · 0 repositories · arXiv:2310.17793
-
Evaluating, Understanding, and Improving Constrained Text Generation for Large Language Models 25 Oct 2023 · 0 repositories · arXiv:2310.16343
-
An Early Evaluation of GPT-4V(ision) 25 Oct 2023 · 1 repository · arXiv:2310.16534
-
Can GPT models Follow Human Summarization Guidelines? Evaluating ChatGPT and GPT-4 for Dialogue Summarization 25 Oct 2023 · 0 repositories · arXiv:2310.16810
-
Is ChatGPT a Good Multi-Party Conversation Solver? 25 Oct 2023 · 1 repository · arXiv:2310.16301
-
LLM Performance Predictors are good initializers for Architecture Search 25 Oct 2023 · 1 repository · arXiv:2310.16712Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
netFound: Foundation Model for Network Security 25 Oct 2023 · 1 repository · arXiv:2310.17025
-
OccuQuest: Mitigating Occupational Bias for Inclusive Large Language Models 25 Oct 2023 · 1 repository · arXiv:2310.16517
-
SuperHF: Supervised Iterative Learning from Human Feedback 25 Oct 2023 · 1 repository · arXiv:2310.16763Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Using GPT-4 to Augment Unbalanced Data for Automatic Scoring 25 Oct 2023 · 0 repositories · arXiv:2310.18365
-
Confounder Balancing in Adversarial Domain Adaptation for Pre-Trained Large Models Fine-Tuning 24 Oct 2023 · 0 repositories · arXiv:2310.16062
-
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning 24 Oct 2023 · 3 repositories · arXiv:2310.16049Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
NoteChat: A Dataset of Synthetic Doctor-Patient Conversations Conditioned on Clinical Notes 24 Oct 2023 · 1 repository · arXiv:2310.15959
-
UI Layout Generation with LLMs Guided by UI Grammar 24 Oct 2023 · 0 repositories · arXiv:2310.15455
-
AlpaCare:Instruction-tuned Large Language Models for Medical Application 23 Oct 2023 · 1 repository · arXiv:2310.14558Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Analyzing Multilingual Competency of LLMs in Multi-Turn Instruction Following: A Case Study of Arabic 23 Oct 2023 · 0 repositories · arXiv:2310.14819
-
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation 23 Oct 2023 · 0 repositories · arXiv:2310.15123
-
Causal Inference Using LLM-Guided Discovery 23 Oct 2023 · 0 repositories · arXiv:2310.15117
-
Evaluating Spatial Understanding of Large Language Models 23 Oct 2023 · 1 repository · arXiv:2310.14540Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Evaluating the Knowledge Base Completion Potential of GPT 23 Oct 2023 · 0 repositories · arXiv:2310.14771
-
Exploring the Boundaries of GPT-4 in Radiology 23 Oct 2023 · 0 repositories · arXiv:2310.14573
-
GPT-4 as an Effective Zero-Shot Evaluator for Scientific Figure Captions 23 Oct 2023 · 0 repositories · arXiv:2310.15405
-
InstructExcel: A Benchmark for Natural Language Instruction in Excel 23 Oct 2023 · 0 repositories · arXiv:2310.14495
-
Large Language Models can Share Images, Too! 23 Oct 2023 · 2 repositories · arXiv:2310.14804
-
LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers 23 Oct 2023 · 1 repository · arXiv:2310.15164Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Location-Aware Visual Question Generation with Lightweight Models 23 Oct 2023 · 1 repository · arXiv:2310.15129
-
TeleQnA: A Benchmark Dataset to Assess Large Language Models Telecommunications Knowledge 23 Oct 2023 · 1 repository · arXiv:2310.15051
-
Unleashing the potential of prompt engineering for large language models 23 Oct 2023 · 0 repositories · arXiv:2310.14735
-
CXR-LLAVA: a multimodal large language model for interpreting chest X-ray images 22 Oct 2023 · 1 repository · arXiv:2310.18341Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Large Language Models are biased to overestimate profoundness 22 Oct 2023 · 1 repository · arXiv:2310.14422
-
GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4 21 Oct 2023 · 1 repository · arXiv:2310.13988Syntology 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Small Language Models Fine-tuned to Coordinate Larger Language Models improve Complex Reasoning 21 Oct 2023 · 1 repository · arXiv:2310.18338Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
AllTogether: Investigating the Efficacy of Spliced Prompt for Web Navigation using Large Language Models 20 Oct 2023 · 0 repositories · arXiv:2310.18331
-
Ask Language Model to Clean Your Noisy Translation Data 20 Oct 2023 · 0 repositories · arXiv:2310.13469
-
BotChat: Evaluating LLMs' Capabilities of Having Multi-Turn Dialogues 20 Oct 2023 · 1 repository · arXiv:2310.13650Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Cache me if you Can: an Online Cost-aware Teacher-Student framework to Reduce the Calls to Large Language Models 20 Oct 2023 · 1 repository · arXiv:2310.13395Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
She had Cobalt Blue Eyes: Prompt Testing to Create Aligned and Sustainable Language Models 20 Oct 2023 · 0 repositories · arXiv:2310.18333
-
Evaluation Metrics in the Era of GPT-4: Reliably Evaluating Large Language Models on Sequence to Sequence Tasks 20 Oct 2023 · 1 repository · arXiv:2310.13800Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
MoqaGPT : Zero-Shot Multi-modal Open-domain Question Answering with Large Language Model 20 Oct 2023 · 1 repository · arXiv:2310.13265
-
The Perils & Promises of Fact-checking with Large Language Models 20 Oct 2023 · 0 repositories · arXiv:2310.13549
-
Tuna: Instruction Tuning using Feedback from Large Language Models 20 Oct 2023 · 1 repository · arXiv:2310.13385
-
AgentTuning: Enabling Generalized Agent Abilities for LLMs 19 Oct 2023 · 1 repository · arXiv:2310.12823Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks 19 Oct 2023 · 0 repositories · arXiv:2310.12516
-
AutoMix: Automatically Mixing Language Models 19 Oct 2023 · 1 repository · arXiv:2310.12963Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Eureka: Human-Level Reward Design via Coding Large Language Models 19 Oct 2023 · 1 repository · arXiv:2310.12931Syntology official (archive's flag): 8 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Experimental Narratives: A Comparison of Human Crowdsourced Storytelling and AI Storytelling 19 Oct 2023 · 0 repositories · arXiv:2310.12902
-
Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language Models 19 Oct 2023 · 0 repositories · arXiv:2310.12481
-
ExtractGPT: Exploring the Potential of Large Language Models for Product Attribute Value Extraction 19 Oct 2023 · 1 repository · arXiv:2310.12537
-
The Foundation Model Transparency Index 19 Oct 2023 · 1 repository · arXiv:2310.12941
-
Evaluating the Symbol Binding Ability of Large Language Models for Multiple-Choice Questions in Vietnamese General Education 18 Oct 2023 · 0 repositories · arXiv:2310.12059
-
Quantifying Self-diagnostic Atomic Knowledge in Chinese Medical Foundation Model: A Computational Analysis 18 Oct 2023 · 1 repository · arXiv:2310.11722
-
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents 18 Oct 2023 · 2 repositories · arXiv:2310.11667
-
CoMPosT: Characterizing and Evaluating Caricature in LLM Simulations 17 Oct 2023 · 1 repository · arXiv:2310.11501Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Entity Matching using Large Language Models 17 Oct 2023 · 1 repository · arXiv:2310.11244
-
LLMs as Hackers: Autonomous Linux Privilege Escalation Attacks 17 Oct 2023 · 1 repository · arXiv:2310.11409
-
Experimenting AI Technologies for Disinformation Combat: the IDMO Project 17 Oct 2023 · 0 repositories · arXiv:2310.11097
-
Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament 17 Oct 2023 · 0 repositories · arXiv:2310.13014
-
Learning from Red Teaming: Gender Bias Provocation and Mitigation in Large Language Models 17 Oct 2023 · 0 repositories · arXiv:2310.11079
-
Probing the Creativity of Large Language Models: Can models produce divergent semantic association? 17 Oct 2023 · 1 repository · arXiv:2310.11158Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Alexpaca: Learning Factual Clarification Question Generation Without Examples 17 Oct 2023 · 0 repositories · arXiv:2310.11571
-
Battle of the Large Language Models: Dolly vs LLaMA vs Vicuna vs Guanaco vs Bard vs ChatGPT -- A Text-to-SQL Parsing Comparison 16 Oct 2023 · 0 repositories · arXiv:2310.10190
-
BiomedJourney: Counterfactual Biomedical Image Generation by Instruction-Learning from Multimodal Patient Journeys 16 Oct 2023 · 0 repositories · arXiv:2310.10765
-
BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology 16 Oct 2023 · 1 repository · arXiv:2310.10632Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Factored Verification: Detecting and Reducing Hallucination in Summaries of Academic Papers 16 Oct 2023 · 1 repository · arXiv:2310.10627
-
Prompt Packer: Deceiving LLMs through Compositional Instruction with Hidden Attacks 16 Oct 2023 · 0 repositories · arXiv:2310.10077
-
TRIGO: Benchmarking Formal Mathematical Proof Reduction for Generative Language Models 16 Oct 2023 · 1 repository · arXiv:2310.10180Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Verbosity Bias in Preference Labeling by Large Language Models 16 Oct 2023 · 0 repositories · arXiv:2310.10076
-
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer 15 Oct 2023 · 0 repositories · arXiv:2310.09762
-
Large Language Models for In-Context Student Modeling: Synthesizing Student's Behavior in Visual Programming 15 Oct 2023 · 1 repository · arXiv:2310.10690
-
Instruction Tuning with Human Curriculum 14 Oct 2023 · 1 repository · arXiv:2310.09518
-
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning 13 Oct 2023 · 1 repository · arXiv:2310.09430Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Automated Claim Matching with Large Language Models: Empowering Fact-Checkers in the Fight Against Misinformation 13 Oct 2023 · 0 repositories · arXiv:2310.09223
-
Dont Add, dont Miss: Effective Content Preserving Generation from Pre-Selected Text Spans 13 Oct 2023 · 1 repository · arXiv:2310.09017Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
GLoRE: Evaluating Logical Reasoning of Large Language Models 13 Oct 2023 · 1 repository · arXiv:2310.09107
-
Can GPT models be Financial Analysts? An Evaluation of ChatGPT and GPT-4 on mock CFA Exams 12 Oct 2023 · 0 repositories · arXiv:2310.08678
-
Can Large Language Models Really Improve by Self-critiquing Their Own Plans? 12 Oct 2023 · 0 repositories · arXiv:2310.08118
-
Interpreting Learned Feedback Patterns in Large Language Models 12 Oct 2023 · 1 repository · arXiv:2310.08164Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Large language models can replicate cross-cultural differences in personality 12 Oct 2023 · 0 repositories · arXiv:2310.10679
-
Octopus: Embodied Vision-Language Programmer from Environmental Feedback 12 Oct 2023 · 1 repository · arXiv:2310.08588Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models 12 Oct 2023 · 3 repositories · arXiv:2310.08491Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 9 harvested samples) · 5 pointer-only (licence)
-
QASiNa: Religious Domain Question Answering using Sirah Nabawiyah 12 Oct 2023 · 1 repository · arXiv:2310.08102
-
Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning 12 Oct 2023 · 0 repositories · arXiv:2310.08166