Methods › Natural Language Processing › Transformers › GPT-3 › Papers, page 14
GPT-3
Papers archive 2025-07-28
archive papers tagged: 1,906 · with a code link: 866 · where Syntology ran a sample: 319 (259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (319 of 1,906 tagged: 259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument)
Page 14 of 20: papers 1,301 to 1,400 of 1,906, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Towards Explainable Conversational Recommender Systems 27 May 2023 · 1 repository · arXiv:2305.18363
-
What can Large Language Models do in chemistry? A comprehensive benchmark on eight tasks 27 May 2023 · 1 repository · arXiv:2305.18365
-
Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance 26 May 2023 · 1 repository · arXiv:2305.17306Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
ChatGPT: A Study on its Utility for Ubiquitous Software Engineering Tasks 26 May 2023 · 0 repositories · arXiv:2305.16837
-
Counterfactual reasoning: Testing language models' understanding of hypothetical scenarios 26 May 2023 · 1 repository · arXiv:2305.16572
-
Distinguishing Human Generated Text From ChatGPT Generated Text Using Machine Learning 26 May 2023 · 0 repositories · arXiv:2306.01761
-
Do GPTs Produce Less Literal Translations? 26 May 2023 · 1 repository · arXiv:2305.16806
-
Evaluation of Question Generation Needs More References 26 May 2023 · 0 repositories · arXiv:2305.16626
-
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing 26 May 2023 · 0 repositories · arXiv:2305.16635
-
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model 26 May 2023 · 0 repositories · arXiv:2305.17116
-
Large Language Models as Tool Makers 26 May 2023 · 1 repository · arXiv:2305.17126
-
Playing repeated games with Large Language Models 26 May 2023 · 0 repositories · arXiv:2305.16867
-
Linguistic Properties of Truthful Response 25 May 2023 · 1 repository · arXiv:2305.15875
-
Not wacky vs. definitely wacky: A study of scalar adverbs in pretrained language models 25 May 2023 · 0 repositories · arXiv:2305.16426
-
A Causal View of Entity Bias in (Large) Language Models 24 May 2023 · 1 repository · arXiv:2305.14695Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
LAraBench: Benchmarking Arabic AI with Large Language Models 24 May 2023 · 0 repositories · arXiv:2305.14982
-
Chain-of-Questions Training with Latent Answers for Robust Multistep Question Answering 24 May 2023 · 0 repositories · arXiv:2305.14901
-
ChatAgri: Exploring Potentials of ChatGPT on Cross-linguistic Agricultural Text Classification 24 May 2023 · 1 repository · arXiv:2305.15024
-
Don't Take This Out of Context! On the Need for Contextual Models and Evaluations for Stylistic Rewriting 24 May 2023 · 0 repositories · arXiv:2305.14755
-
ExpertPrompting: Instructing Large Language Models to be Distinguished Experts 24 May 2023 · 2 repositories · arXiv:2305.14688Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Harnessing the Power of Large Language Models for Natural Language to First-Order Logic Translation 24 May 2023 · 1 repository · arXiv:2305.15541Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Psychological Metrics for Dialog System Evaluation 24 May 2023 · 0 repositories · arXiv:2305.14757
-
I Spy a Metaphor: Large Language Models and Diffusion Models Co-Create Visual Metaphors 24 May 2023 · 1 repository · arXiv:2305.14724
-
Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning 24 May 2023 · 1 repository · arXiv:2305.15065Syntology official (archive's flag): 9 ran · 9 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
Enabling and Analyzing How to Efficiently Extract Information from Hybrid Long Documents with LLMs 24 May 2023 · 0 repositories · arXiv:2305.16344
-
Mastering the ABCDs of Complex Questions: Answer-Based Claim Decomposition for Fine-grained Self-Evaluation 24 May 2023 · 0 repositories · arXiv:2305.14750
-
Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering 24 May 2023 · 1 repository · arXiv:2305.15387Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples)
-
Self-Checker: Plug-and-Play Modules for Fact-Checking with Large Language Models 24 May 2023 · 1 repository · arXiv:2305.14623Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Testing Causal Models of Word Meaning in GPT-3 and -4 24 May 2023 · 1 repository · arXiv:2305.14630
-
ToMChallenges: A Principle-Guided Dataset and Diverse Evaluation Tasks for Exploring Theory of Mind 24 May 2023 · 1 repository · arXiv:2305.15068
-
Fine-tuned LLMs Know More, Hallucinate Less with Few-Shot Sequence-to-Sequence Semantic Parsing over Wikidata 23 May 2023 · 1 repository · arXiv:2305.14202
-
Dancing Between Success and Failure: Edit-level Simplification Evaluation using SALSA 23 May 2023 · 0 repositories · arXiv:2305.14458
-
Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data Curation 23 May 2023 · 1 repository · arXiv:2305.14327Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples)
-
Few-Shot Data Synthesis for Open Domain Multi-Hop Question Answering 23 May 2023 · 0 repositories · arXiv:2305.13691
-
Enhancing Black-Box Few-Shot Text Classification with Prompt-Based Data Augmentation 23 May 2023 · 0 repositories · arXiv:2305.13785
-
HumBEL: A Human-in-the-Loop Approach for Evaluating Demographic Factors of Language Models in Human-Machine Conversations 23 May 2023 · 1 repository · arXiv:2305.14195
-
IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions 23 May 2023 · 0 repositories · arXiv:2305.14010
-
Images in Language Space: Exploring the Suitability of Large Language Models for Vision & Language Tasks 23 May 2023 · 1 repository · arXiv:2305.13782
-
INSTRUCTSCORE: Explainable Text Generation Evaluation with Finegrained Feedback 23 May 2023 · 2 repositories · arXiv:2305.14282
-
Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video Chain-of-Thought 23 May 2023 · 1 repository · arXiv:2305.13903
-
MathDial: A Dialogue Tutoring Dataset with Rich Pedagogical Properties Grounded in Math Reasoning Problems 23 May 2023 · 1 repository · arXiv:2305.14536Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
NAIL: Lexical Retrieval Indices with Efficient Non-Autoregressive Decoders 23 May 2023 · 0 repositories · arXiv:2305.14499
-
Sources of Hallucination by Large Language Models on Inference Tasks 23 May 2023 · 1 repository · arXiv:2305.14552
-
Two Failures of Self-Consistency in the Multi-Step Reasoning of LLMs 23 May 2023 · 0 repositories · arXiv:2305.14279
-
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia 23 May 2023 · 1 repository · arXiv:2305.14292Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Let GPT be a Math Tutor: Teaching Math Word Problem Solvers with Customized Exercise Generation 22 May 2023 · 0 repositories · arXiv:2305.14386
-
A Study of Generative Large Language Model for Medical Research and Healthcare 22 May 2023 · 1 repository · arXiv:2305.13523
-
Cognitive network science reveals bias in GPT-3, ChatGPT, and GPT-4 mirroring math anxiety in high-school students 22 May 2023 · 0 repositories · arXiv:2305.18320
-
Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study 22 May 2023 · 1 repository · arXiv:2305.13062
-
InheritSumm: A General, Versatile and Compact Summarizer by Distilling from GPT 22 May 2023 · 0 repositories · arXiv:2305.13083
-
MAILEX: Email Event and Argument Extraction 22 May 2023 · 1 repository · arXiv:2305.13469
-
Measuring Inductive Biases of In-Context Learning with Underspecified Demonstrations 22 May 2023 · 1 repository · arXiv:2305.13299
-
Investigating Agency of LLMs in Human-AI Collaboration Tasks 22 May 2023 · 0 repositories · arXiv:2305.12815
-
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers 21 May 2023 · 0 repositories · arXiv:2305.12563
-
BiasAsker: Measuring the Bias in Conversational AI System 21 May 2023 · 1 repository · arXiv:2305.12434
-
Abstract Meaning Representation-Based Logic-Driven Data Augmentation for Logical Reasoning 21 May 2023 · 1 repository · arXiv:2305.12599
-
GPT-3.5, GPT-4, or BARD? Evaluating LLMs Reasoning Ability in Zero-Shot Setting and Performance Boosting Through Prompts 21 May 2023 · 0 repositories · arXiv:2305.12477
-
Can NLP Models Correctly Reason Over Contexts that Break the Common Assumptions? 20 May 2023 · 0 repositories · arXiv:2305.12096
-
LogiCoT: Logical Chain-of-Thought Instruction-Tuning 20 May 2023 · 1 repository · arXiv:2305.12147
-
Practical PCG Through Large Language Models 20 May 2023 · 0 repositories · arXiv:2305.18243
-
AutoTrial: Prompting Language Models for Clinical Trial Design 19 May 2023 · 0 repositories · arXiv:2305.11366
-
Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding 19 May 2023 · 2 repositories · arXiv:2305.12031Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Exploring the Upper Limits of Text-Based Collaborative Filtering Using Large Language Models: Discoveries and Insights 19 May 2023 · 0 repositories · arXiv:2305.11700
-
SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative Models 19 May 2023 · 1 repository · arXiv:2305.11840
-
Self-Agreement: A Framework for Fine-tuning Language Models to Find Agreement among Diverse Opinions 19 May 2023 · 0 repositories · arXiv:2305.11460
-
AIwriting: Relations Between Image Generation and Digital Writing 18 May 2023 · 0 repositories · arXiv:2305.10834
-
Deep Learning Methods for Extracting Metaphorical Names of Flowers and Plants 18 May 2023 · 0 repositories · arXiv:2305.10833
-
Generalized Multiple Intent Conditioned Slot Filling 18 May 2023 · 0 repositories · arXiv:2305.11023
-
Generalized Planning in PDDL Domains with Pretrained Large Language Models 18 May 2023 · 1 repository · arXiv:2305.11014Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Large Language Models can be Guided to Evade AI-Generated Text Detection 18 May 2023 · 1 repository · arXiv:2305.10847
-
CoEdIT: Text Editing by Task-Specific Instruction Tuning 17 May 2023 · 1 repository · arXiv:2305.09857
-
From chocolate bunny to chocolate crocodile: Do Language Models Understand Noun Compounds? 17 May 2023 · 0 repositories · arXiv:2305.10568
-
Knowledge Graph Completion Models are Few-shot Learners: An Empirical Study of Relation Labeling in E-commerce with LLMs 17 May 2023 · 0 repositories · arXiv:2305.09858
-
M3KE: A Massive Multi-Level Multi-Subject Knowledge Evaluation Benchmark for Chinese Large Language Models 17 May 2023 · 1 repository · arXiv:2305.10263
-
When Gradient Descent Meets Derivative-Free Optimization: A Match Made in Black-Box Scenario 17 May 2023 · 0 repositories · arXiv:2305.10013
-
A Preliminary Analysis on the Code Generation Capabilities of GPT-3.5 and Bard AI Models for Java Functions 16 May 2023 · 0 repositories · arXiv:2305.09402
-
Knowledge Rumination for Pre-trained Language Models 15 May 2023 · 1 repository · arXiv:2305.08732
-
RL4F: Generating Natural Language Feedback with Reinforcement Learning for Repairing Model Outputs 15 May 2023 · 1 repository · arXiv:2305.08844Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 6 pointer-only (licence)
-
Schema-adaptable Knowledge Graph Construction 15 May 2023 · 1 repository · arXiv:2305.08703
-
Similarity-weighted Construction of Contextualized Commonsense Knowledge Graphs for Knowledge-intense Argumentation Tasks 15 May 2023 · 1 repository · arXiv:2305.08495
-
Small Models are Valuable Plug-ins for Large Language Models 15 May 2023 · 1 repository · arXiv:2305.08848Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Text Classification via Large Language Models 15 May 2023 · 1 repository · arXiv:2305.08377Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Assessing the potential of LLM-assisted annotation for corpus-based pragmatics and discourse analysis: The case of apology 15 May 2023 · 0 repositories · arXiv:2305.08339
-
The Machine Psychology of Cooperation: Can GPT models operationalise prompts for altruism, cooperation, competitiveness and selfishness in economic games? 13 May 2023 · 2 repositories · arXiv:2305.07970Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? 12 May 2023 · 8 repositories · arXiv:2305.07759Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples)
-
When Giant Language Brains Just Aren't Enough! Domain Pizzazz with Knowledge Sparkle Dust 12 May 2023 · 0 repositories · arXiv:2305.07230
-
Spear Phishing With Large Language Models 11 May 2023 · 0 repositories · arXiv:2305.06972
-
Overinformative Question Answering by Humans and Machines 11 May 2023 · 0 repositories · arXiv:2305.07151
-
Recommendation as Instruction Following: A Large Language Model Empowered Recommendation Approach 11 May 2023 · 0 repositories · arXiv:2305.07001
-
Bits of Grass: Does GPT already know how to write like Whitman? 10 May 2023 · 0 repositories · arXiv:2305.11064
-
Davinci the Dualist: the mind-body divide in large language models and in human learners 10 May 2023 · 0 repositories · arXiv:2305.07667
-
Generating medically-accurate summaries of patient-provider dialogue: A multi-stage approach using large language models 10 May 2023 · 0 repositories · arXiv:2305.05982
-
Benchmarking large language models for biomedical natural language processing applications and recommendations 10 May 2023 · 1 repository · arXiv:2305.16326
-
Summarizing, Simplifying, and Synthesizing Medical Evidence Using GPT-3 (with Varying Success) 10 May 2023 · 1 repository · arXiv:2305.06299
-
CodeIE: Large Code Generation Models are Better Few-Shot Information Extractors 9 May 2023 · 1 repository · arXiv:2305.05711Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Towards an Automatic Optimisation Model Generator Assisted with Generative Pre-trained Transformer 9 May 2023 · 0 repositories · arXiv:2305.05811
-
Do Large Language Models Show Decision Heuristics Similar to Humans? A Case Study Using GPT-3.5 8 May 2023 · 0 repositories · arXiv:2305.04400
-
Explanation-based Finetuning Makes Models More Robust to Spurious Cues 8 May 2023 · 1 repository · arXiv:2305.04990
-
GersteinLab at MEDIQA-Chat 2023: Clinical Note Summarization from Doctor-Patient Conversations through Fine-tuning and In-context Learning 8 May 2023 · 0 repositories · arXiv:2305.05001
-
NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge 8 May 2023 · 1 repository · arXiv:2305.04978