Methods › Natural Language Processing › Autoregressive Transformers › GPT-2 › Papers, page 4
GPT-2
Papers archive 2025-07-28
archive papers tagged: 768 · with a code link: 339 · where Syntology ran a sample: 125 (99 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (125 of 768 tagged: 99 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument)
Page 4 of 8: papers 301 to 400 of 768, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Copy Suppression: Comprehensively Understanding an Attention Head 6 Oct 2023 · 1 repository · arXiv:2310.04625Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
NOLA: Compressing LoRA using Linear Combination of Random Basis 4 Oct 2023 · 1 repository · arXiv:2310.02556Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels 2 Oct 2023 · 0 repositories · arXiv:2310.01655
-
AE-GPT: Using Large Language Models to Extract Adverse Events from Surveillance Reports-A Use Case with Influenza Vaccine Adverse Events 28 Sep 2023 · 0 repositories · arXiv:2309.16150
-
Constraints First: A New MDD-based Model to Generate Sentences Under Constraints 21 Sep 2023 · 0 repositories · arXiv:2309.12415
-
The Languini Kitchen: Enabling Language Modelling Research at Different Scales of Compute 20 Sep 2023 · 1 repository · arXiv:2309.11197
-
Rigorously Assessing Natural Language Explanations of Neurons 19 Sep 2023 · 0 repositories · arXiv:2309.10312
-
RECAP: Retrieval-Augmented Audio Captioning 18 Sep 2023 · 1 repository · arXiv:2309.09836Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
A Modern Turkish Poet: Fine-Tuned GPT-2 15 Sep 2023 · 1 repository
-
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models 15 Sep 2023 · 1 repository · arXiv:2309.08573
-
Traveling Words: A Geometric Interpretation of Transformers 13 Sep 2023 · 1 repository · arXiv:2309.07315
-
Circuit Breaking: Removing Model Behaviors with Targeted Ablation 12 Sep 2023 · 1 repository · arXiv:2309.05973Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Characterizing Latent Perspectives of Media Houses Towards Public Figures 12 Sep 2023 · 0 repositories · arXiv:2309.06112
-
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models 11 Sep 2023 · 1 repository · arXiv:2309.05605Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Zero-Shot Audio Captioning via Audibility Guidance 7 Sep 2023 · 0 repositories · arXiv:2309.03884
-
Why do universal adversarial attacks work on large language models?: Geometry might be the answer 1 Sep 2023 · 0 repositories · arXiv:2309.00254
-
Target-independent XLA optimization using Reinforcement Learning 28 Aug 2023 · 0 repositories · arXiv:2308.14364
-
Steering Language Models With Activation Engineering 20 Aug 2023 · 2 repositories · arXiv:2308.10248Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
How Good Are LLMs at Out-of-Distribution Detection? 20 Aug 2023 · 1 repository · arXiv:2308.10261Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
A tailored Handwritten-Text-Recognition System for Medieval Latin 18 Aug 2023 · 0 repositories · arXiv:2308.09368
-
A Preliminary Study on a Conceptual Game Feature Generation and Recommendation System 16 Aug 2023 · 0 repositories · arXiv:2308.13538
-
Generating Individual Trajectories Using GPT-2 Trained from Scratch on Encoded Spatiotemporal Data 14 Aug 2023 · 0 repositories · arXiv:2308.07940
-
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining 10 Aug 2023 · 2 repositories · arXiv:2308.05734Syntology official (archive's flag): 8 ran · 16 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 2 honoured, 3 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 27 harvested samples) · 19 pointer-only (licence)
-
I-WAS: a Data Augmentation Method with GPT-2 for Simile Detection 8 Aug 2023 · 0 repositories · arXiv:2308.04109
-
Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty 3 Aug 2023 · 1 repository · arXiv:2308.02019
-
Unveiling Gender Bias in Terms of Profession Across LLMs: Analyzing and Addressing Sociological Implications 18 Jul 2023 · 0 repositories · arXiv:2307.09162
-
A mixed policy to improve performance of language models on math problems 17 Jul 2023 · 1 repository · arXiv:2307.08767
-
MorphPiece : A Linguistic Tokenizer for Large Language Models 14 Jul 2023 · 0 repositories · arXiv:2307.07262
-
SimpleMTOD: A Simple Language Model for Multimodal Task-Oriented Dialogue with Symbolic Scene Representation 10 Jul 2023 · 0 repositories · arXiv:2307.04907
-
Assessing the efficacy of large language models in generating accurate teacher responses 9 Jul 2023 · 0 repositories · arXiv:2307.04274
-
CAME: Confidence-guided Adaptive Memory Efficient Optimization 5 Jul 2023 · 2 repositories · arXiv:2307.02047Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 1 pointer-only (licence)
-
Evaluating the Effectiveness of Large Language Models in Representing Textual Descriptions of Geometry and Spatial Relations 5 Jul 2023 · 0 repositories · arXiv:2307.03678
-
TensorGPT: Efficient Compression of Large Language Models based on Tensor-Train Decomposition 2 Jul 2023 · 0 repositories · arXiv:2307.00526
-
Stay on topic with Classifier-Free Guidance 30 Jun 2023 · 0 repositories · arXiv:2306.17806
-
A negation detection assessment of GPTs: analysis with the xNot360 dataset 29 Jun 2023 · 0 repositories · arXiv:2306.16638
-
Is Pre-training Truly Better Than Meta-Learning? 24 Jun 2023 · 0 repositories · arXiv:2306.13841
-
Investigating Pre-trained Language Models on Cross-Domain Datasets, a Step Closer to General AI 21 Jun 2023 · 0 repositories · arXiv:2306.12205
-
InRank: Incremental Low-Rank Learning 20 Jun 2023 · 1 repository · arXiv:2306.11250
-
Learning to Generate Better Than Your LLM 20 Jun 2023 · 1 repository · arXiv:2306.11816
-
Generative Sequential Recommendation with GPTRec 19 Jun 2023 · 0 repositories · arXiv:2306.11114
-
Explore, Establish, Exploit: Red Teaming Language Models from Scratch 15 Jun 2023 · 3 repositories · arXiv:2306.09442Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The pop song generator: designing an online course to teach collaborative, creative AI 15 Jun 2023 · 0 repositories · arXiv:2306.10069
-
On the N-gram Approximation of Pre-trained Language Models 12 Jun 2023 · 0 repositories · arXiv:2306.06892
-
The BEA 2023 Shared Task on Generating AI Teacher Responses in Educational Dialogues 12 Jun 2023 · 0 repositories · arXiv:2306.06941
-
Language Models Can Learn Exceptions to Syntactic Rules 9 Jun 2023 · 1 repository · arXiv:2306.05969
-
Understanding Telecom Language Through Large Language Models 9 Jun 2023 · 0 repositories · arXiv:2306.07933
-
An Empirical Analysis of Parameter-Efficient Methods for Debiasing Pre-Trained Language Models 6 Jun 2023 · 1 repository · arXiv:2306.04067
-
Language acquisition: do children and language models follow similar learning stages? 6 Jun 2023 · 0 repositories · arXiv:2306.03586
-
Efficient GPT Model Pre-training using Tensor Train Matrix Representation 5 Jun 2023 · 0 repositories · arXiv:2306.02697
-
Can Contextual Biasing Remain Effective with Whisper and GPT-2? 2 Jun 2023 · 1 repository · arXiv:2306.01942
-
TopEx: Topic-based Explanations for Model Comparison 1 Jun 2023 · 0 repositories · arXiv:2306.00976
-
Test-Time Training on Nearest Neighbors for Large Language Models 29 May 2023 · 1 repository · arXiv:2305.18466Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Transformer Language Models Handle Word Frequency in Prediction Head 29 May 2023 · 0 repositories · arXiv:2305.18294
-
The Curse of Recursion: Training on Generated Data Makes Models Forget 27 May 2023 · 1 repository · arXiv:2305.17493
-
Backpack Language Models 26 May 2023 · 1 repository · arXiv:2305.16765Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 4 harvested samples)
-
Learning and Leveraging Verifiers to Improve Planning Capabilities of Pre-trained Language Models 26 May 2023 · 0 repositories · arXiv:2305.17077
-
Not wacky vs. definitely wacky: A study of scalar adverbs in pretrained language models 25 May 2023 · 0 repositories · arXiv:2305.16426
-
Editing Common Sense in Transformers 24 May 2023 · 1 repository · arXiv:2305.14956Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Inference-Time Policy Adapters (IPA): Tailoring Extreme-Scale LMs without Fine-tuning 24 May 2023 · 1 repository · arXiv:2305.15065Syntology official (archive's flag): 9 ran · 9 ran (of which 3 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples)
-
LLMDet: A Third Party Large Language Models Generated Text Detection Tool 24 May 2023 · 1 repository · arXiv:2305.15004Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Advancing Precise Outline-Conditioned Text Generation with Task Duality and Explicit Outline Control 23 May 2023 · 0 repositories · arXiv:2305.14459
-
On Robustness of Finetuned Transformer-based NLP Models 23 May 2023 · 1 repository · arXiv:2305.14453
-
Probing Brain Context-Sensitivity with Masked-Attention Generation 23 May 2023 · 0 repositories · arXiv:2305.13863
-
Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training 23 May 2023 · 7 repositories · arXiv:2305.14342Syntology 13 ran (of which 4 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 19 harvested samples)
-
Revisiting the Architectures like Pointer Networks to Efficiently Improve the Next Word Distribution, Summarization Factuality, and Beyond 20 May 2023 · 1 repository · arXiv:2305.12289
-
Scaling laws for language encoding models in fMRI 19 May 2023 · 1 repository · arXiv:2305.11863
-
Learning to Reason over Scene Graphs: A Case Study of Finetuning GPT-2 into a Robot Language Model for Grounded Task Planning 12 May 2023 · 0 repositories · arXiv:2305.07716
-
TinyStories: How Small Can Language Models Be and Still Speak Coherent English? 12 May 2023 · 8 repositories · arXiv:2305.07759Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples)
-
Effects of sub-word segmentation on performance of transformer language models 9 May 2023 · 0 repositories · arXiv:2305.05480
-
NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge 8 May 2023 · 1 repository · arXiv:2305.04978
-
Adapting Transformer Language Models for Predictive Typing in Brain-Computer Interfaces 5 May 2023 · 0 repositories · arXiv:2305.03819
-
How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language model 30 Apr 2023 · 3 repositories · arXiv:2305.00586
-
Towards Automated Circuit Discovery for Mechanistic Interpretability 28 Apr 2023 · 4 repositories · arXiv:2304.14997Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
How to Do Things with Deep Learning Code 19 Apr 2023 · 0 repositories · arXiv:2304.09406
-
An Empirical Study of Multitask Learning to Improve Open Domain Dialogue Systems 17 Apr 2023 · 1 repository · arXiv:2304.08115
-
Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models 15 Apr 2023 · 0 repositories · arXiv:2304.07619
-
Stochastic Code Generation 14 Apr 2023 · 0 repositories · arXiv:2304.08243
-
PGTask: Introducing the Task of Profile Generation from Dialogues 13 Apr 2023 · 1 repository · arXiv:2304.06634
-
Localizing Model Behavior with Path Patching 12 Apr 2023 · 1 repository · arXiv:2304.05969Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples)
-
On the Possibilities of AI-Generated Text Detection 10 Apr 2023 · 0 repositories · arXiv:2304.04736
-
GPT4Rec: A Generative Framework for Personalized Recommendation and User Interests Interpretation 8 Apr 2023 · 0 repositories · arXiv:2304.03879
-
LLMMaps -- A Visual Metaphor for Stratified Evaluation of Large Language Models 2 Apr 2023 · 1 repository · arXiv:2304.00457
-
How do decoding algorithms distribute information in dialogue responses? 29 Mar 2023 · 0 repositories · arXiv:2303.17006
-
Zero-Shot Generalizable End-to-End Task-Oriented Dialog System using Context Summarization and Domain Schema 28 Mar 2023 · 1 repository · arXiv:2303.16252
-
Personalizing Task-oriented Dialog Systems via Zero-shot Generalizable Reward Function 24 Mar 2023 · 0 repositories · arXiv:2303.13797
-
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations 16 Mar 2023 · 2 repositories · arXiv:2303.09435Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
GCRE-GPT: A Generative Model for Comparative Relation Extraction 15 Mar 2023 · 0 repositories · arXiv:2303.08601
-
Learning Combinatorial Prompts for Universal Controllable Image Captioning 11 Mar 2023 · 0 repositories · arXiv:2303.06338
-
Open-Ended Medical Visual Question Answering Through Prefix Tuning of Language Models 10 Mar 2023 · 1 repository · arXiv:2303.05977Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Stealing the Decoding Algorithms of Language Models 8 Mar 2023 · 1 repository · arXiv:2303.04729
-
Towards Zero-Shot Functional Compositionality of Language Models 6 Mar 2023 · 1 repository · arXiv:2303.03103
-
Industry Risk Assessment via Hierarchical Financial Data Using Stock Market Sentiment Indicators 5 Mar 2023 · 0 repositories · arXiv:2303.02707
-
Information-Restricted Neural Language Models Reveal Different Brain Regions' Sensitivity to Semantics, Syntax and Context 28 Feb 2023 · 1 repository · arXiv:2302.14389Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Inseq: An Interpretability Toolkit for Sequence Generation Models 27 Feb 2023 · 2 repositories · arXiv:2302.13942
-
Fast Attention Requires Bounded Entries 26 Feb 2023 · 0 repositories · arXiv:2302.13214
-
Conveying the Predicted Future to Users: A Case Study of Story Plot Prediction 17 Feb 2023 · 1 repository · arXiv:2302.09122
-
Tree-Based Representation and Generation of Natural and Mathematical Language 15 Feb 2023 · 1 repository · arXiv:2302.07974
-
FairPy: A Toolkit for Evaluation of Prediction Biases and their Mitigation in Large Language Models 10 Feb 2023 · 1 repository · arXiv:2302.05508
-
Nationality Bias in Text Generation 5 Feb 2023 · 0 repositories · arXiv:2302.02463
-
REaLTabFormer: Generating Realistic Relational and Tabular Data using Transformers 4 Feb 2023 · 3 repositories · arXiv:2302.02041Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)