Methods › Natural Language Processing › Autoregressive Transformers › GPT-2 › Papers, page 3
GPT-2
Papers archive 2025-07-28
archive papers tagged: 768 · with a code link: 339 · where Syntology ran a sample: 125 (99 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (125 of 768 tagged: 99 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument)
Page 3 of 8: papers 201 to 300 of 768, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Understanding Intrinsic Socioeconomic Biases in Large Language Models 28 May 2024 · 0 repositories · arXiv:2405.18662
-
InversionView: A General-Purpose Method for Reading Information from Neural Activations 27 May 2024 · 1 repository · arXiv:2405.17653Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; the one sample that ran constructed an object rather than computing a result (of 6 harvested samples) · 6 pointer-only (licence)
-
The Scaling Law in Stellar Light Curves 27 May 2024 · 0 repositories · arXiv:2405.17156
-
Incremental Comprehension of Garden-Path Sentences by Large Language Models: Semantic Interpretation, Syntactic Re-Analysis, and Attention 25 May 2024 · 0 repositories · arXiv:2405.16042
-
The Buffer Mechanism for Multi-Step Information Reasoning in Language Models 24 May 2024 · 0 repositories · arXiv:2405.15302
-
From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step 23 May 2024 · 1 repository · arXiv:2405.14838Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Not All Language Model Features Are Linear 23 May 2024 · 1 repository · arXiv:2405.14860Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Automatically Identifying Local and Global Circuits with Linear Computation Graphs 22 May 2024 · 0 repositories · arXiv:2405.13868
-
Quantifying Semantic Emergence in Language Models 21 May 2024 · 1 repository · arXiv:2405.12617Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 6 harvested samples)
-
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models 16 May 2024 · 2 repositories · arXiv:2405.10299
-
Matching domain experts by training from scratch on domain knowledge 15 May 2024 · 0 repositories · arXiv:2405.09395
-
Beyond Scaling Laws: Understanding Transformer Performance with Associative Memory 14 May 2024 · 0 repositories · arXiv:2405.08707
-
Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control 14 May 2024 · 0 repositories · arXiv:2405.08366
-
Open-vocabulary Auditory Neural Decoding Using fMRI-prompted LLM 13 May 2024 · 0 repositories · arXiv:2405.07840
-
Quite Good, but Not Enough: Nationality Bias in Large Language Models -- A Case Study of ChatGPT 11 May 2024 · 1 repository · arXiv:2405.06996
-
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning 11 May 2024 · 0 repositories · arXiv:2405.07046
-
How does GPT-2 Predict Acronyms? Extracting and Understanding a Circuit via Mechanistic Interpretability 7 May 2024 · 1 repository · arXiv:2405.04156Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Anchored Answers: Unravelling Positional Bias in GPT-2's Multiple-Choice Questions 6 May 2024 · 1 repository · arXiv:2405.03205
-
Early Transformers: A study on Efficient Training of Transformer Models through Early-Bird Lottery Tickets 2 May 2024 · 0 repositories · arXiv:2405.02353
-
Investigating Wit, Creativity, and Detectability of Large Language Models in Domain-Specific Writing Style Adaptation of Reddit's Showerthoughts 2 May 2024 · 1 repository · arXiv:2405.01660
-
Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper 26 Apr 2024 · 0 repositories · arXiv:2404.17143
-
Augmenting emotion features in irony detection with Large language modeling 18 Apr 2024 · 0 repositories · arXiv:2404.12291
-
Inheritune: Training Smaller Yet More Attentive Language Models 12 Apr 2024 · 1 repository · arXiv:2404.08634Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Guiding Large Language Models to Generate Computer-Parsable Content 8 Apr 2024 · 0 repositories · arXiv:2404.05499
-
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws 8 Apr 2024 · 0 repositories · arXiv:2404.05405
-
Scope Ambiguities in Large Language Models 5 Apr 2024 · 1 repository · arXiv:2404.04332
-
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction 3 Apr 2024 · 3 repositories · arXiv:2404.02905Syntology community repositories only · 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
Collapse of Self-trained Language Models 2 Apr 2024 · 1 repository · arXiv:2404.02305Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models 29 Mar 2024 · 0 repositories · arXiv:2403.20158
-
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models 28 Mar 2024 · 1 repository · arXiv:2403.19521Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs 26 Mar 2024 · 0 repositories · arXiv:2403.17299
-
CodeShell Technical Report 23 Mar 2024 · 0 repositories · arXiv:2403.15747
-
Adapprox: Adaptive Approximation in Adam Optimization via Randomized Low-Rank Matrices 22 Mar 2024 · 0 repositories · arXiv:2403.14958
-
On Zero-Shot Counterspeech Generation by LLMs 22 Mar 2024 · 1 repository · arXiv:2403.14938
-
Selecting Query-bag as Pseudo Relevance Feedback for Information-seeking Conversations 22 Mar 2024 · 0 repositories · arXiv:2404.04272
-
AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts 20 Mar 2024 · 0 repositories · arXiv:2403.13678
-
Incentivizing News Consumption on Social Media Platforms Using Large Language Models and Realistic Bot Accounts 20 Mar 2024 · 1 repository · arXiv:2403.13362
-
Embedded Named Entity Recognition using Probing Classifiers 18 Mar 2024 · 2 repositories · arXiv:2403.11747Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Leap: molecular synthesisability scoring with intermediates 14 Mar 2024 · 0 repositories · arXiv:2403.13005
-
Optimistic Verifiable Training by Controlling Hardware Nondeterminism 14 Mar 2024 · 1 repository · arXiv:2403.09603Syntology official (archive's flag): 9 ran · 9 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale 13 Mar 2024 · 2 repositories · arXiv:2403.08293Syntology official (archive's flag): 2 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Variational Learning is Effective for Large Deep Networks 27 Feb 2024 · 1 repository · arXiv:2402.17641Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
An Integrated Data Processing Framework for Pretraining Foundation Models 26 Feb 2024 · 2 repositories · arXiv:2402.16358
-
Advancing Parameter Efficiency in Fine-tuning via Representation Editing 23 Feb 2024 · 2 repositories · arXiv:2402.15179Syntology official (archive's flag): 1 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 1 pointer-only (licence)
-
Knowledge Graph Enhanced Large Language Model Editing 21 Feb 2024 · 0 repositories · arXiv:2402.13593
-
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization 21 Feb 2024 · 1 repository · arXiv:2402.13919
-
Can GNN be Good Adapter for LLMs? 20 Feb 2024 · 2 repositories · arXiv:2402.12984Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
PromptKD: Distilling Student-Friendly Knowledge for Generative Language Models via Prompt Tuning 20 Feb 2024 · 1 repository · arXiv:2402.12842Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Reflect-RL: Two-Player Online RL Fine-Tuning for LMs 20 Feb 2024 · 1 repository · arXiv:2402.12621Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
GNNavi: Navigating the Information Flow in Large Language Models by Graph Neural Network 18 Feb 2024 · 1 repository · arXiv:2402.11709
-
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss 16 Feb 2024 · 2 repositories · arXiv:2402.10790
-
The Hedgehog & the Porcupine: Expressive Linear Attentions with Softmax Mimicry 6 Feb 2024 · 1 repository · arXiv:2402.04347
-
Generation, Distillation and Evaluation of Motivational Interviewing-Style Reflections with a Foundational Language Model 1 Feb 2024 · 0 repositories · arXiv:2402.01051
-
Human-mediated Large Language Models for Robotic Intervention in Children with Autism Spectrum Disorders 1 Feb 2024 · 0 repositories · arXiv:2402.00260
-
ConSmax: Hardware-Friendly Alternative Softmax with Learnable Parameters 31 Jan 2024 · 1 repository · arXiv:2402.10930
-
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding 31 Jan 2024 · 0 repositories · arXiv:2401.17692
-
Uncertainty-Aware Explainable Recommendation with Large Language Models 31 Jan 2024 · 0 repositories · arXiv:2402.03366
-
A Unified Approach to Emotion Detection and Task-Oriented Dialogue Modeling 24 Jan 2024 · 1 repository · arXiv:2401.13789
-
Leveraging Biases in Large Language Models: "bias-kNN'' for Effective Few-Shot Learning 18 Jan 2024 · 0 repositories · arXiv:2401.09783
-
When Neural Code Completion Models Size up the Situation: Attaining Cheaper and Faster Completion through Dynamic Model Inference 18 Jan 2024 · 1 repository · arXiv:2401.09964
-
Learning from Implicit User Feedback, Emotions and Demographic Information in Task-Oriented and Document-Grounded Dialogues 17 Jan 2024 · 1 repository · arXiv:2401.09248
-
MADA: Meta-Adaptive Optimizers through hyper-gradient Descent 17 Jan 2024 · 0 repositories · arXiv:2401.08893
-
Mission: Impossible Language Models 12 Jan 2024 · 1 repository · arXiv:2401.06416Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
Investigating Data Contamination for Pre-training Language Models 11 Jan 2024 · 0 repositories · arXiv:2401.06059
-
Monte Carlo Tree Search for Recipe Generation using GPT-2 10 Jan 2024 · 0 repositories · arXiv:2401.05199
-
Reinforcement Learning for Optimizing RAG for Domain Chatbots 10 Jan 2024 · 0 repositories · arXiv:2401.06800
-
PIXAR: Auto-Regressive Language Modeling in Pixel Space 6 Jan 2024 · 0 repositories · arXiv:2401.03321
-
Are LLMs Robust for Spoken Dialogues? 4 Jan 2024 · 0 repositories · arXiv:2401.02297
-
Fairness-Aware Structured Pruning in Transformers 24 Dec 2023 · 1 repository · arXiv:2312.15398Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
How Smooth Is Attention? 22 Dec 2023 · 0 repositories · arXiv:2312.14820
-
Can Transformers Learn Sequential Function Classes In Context? 19 Dec 2023 · 0 repositories · arXiv:2312.12655
-
Exploring Multi-Level Threats in Telegram Data with AI-Human Annotation: A Preliminary Study 15 Dec 2023 · 0 repositories
-
Inter-Layer Scheduling Space Exploration for Multi-model Inference on Heterogeneous Chiplets 14 Dec 2023 · 0 repositories · arXiv:2312.09401
-
Successor Heads: Recurring, Interpretable Attention Heads In The Wild 14 Dec 2023 · 0 repositories · arXiv:2312.09230
-
User-Aware Prefix-Tuning is a Good Learner for Personalized Image Captioning 8 Dec 2023 · 0 repositories · arXiv:2312.04793
-
A ripple in time: a discontinuity in American history 2 Dec 2023 · 1 repository · arXiv:2312.01185
-
On Retrieval Augmentation and the Limitations of Language Model Training 16 Nov 2023 · 0 repositories · arXiv:2311.09615
-
Language Model-In-The-Loop: Data Optimal Approach to Learn-To-Recommend Actions in Text Games 13 Nov 2023 · 0 repositories · arXiv:2311.07687
-
On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models 13 Nov 2023 · 0 repositories · arXiv:2311.07692
-
Vision Encoder-Decoder Models for AI Coaching 9 Nov 2023 · 2 repositories · arXiv:2311.16161
-
Massive Editing for Large Language Models via Meta Learning 8 Nov 2023 · 1 repository · arXiv:2311.04661Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models 7 Nov 2023 · 1 repository · arXiv:2311.04131Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Unraveling Downstream Gender Bias from Large Language Models: A Study on AI Educational Writing Assistance 6 Nov 2023 · 1 repository · arXiv:2311.03311
-
Efficient kernel surrogates for neural network-based regression 28 Oct 2023 · 0 repositories · arXiv:2310.18612
-
BabyStories: Can Reinforcement Learning Teach Baby Language Models to Write Better Stories? 25 Oct 2023 · 1 repository · arXiv:2310.16681Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution 25 Oct 2023 · 4 repositories · arXiv:2310.16834Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 0 violated, 13 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 18 harvested samples) · 15 pointer-only (licence)
-
A Language Model with Limited Memory Capacity Captures Interference in Human Sentence Processing 24 Oct 2023 · 0 repositories · arXiv:2310.16142
-
Learning From Free-Text Human Feedback -- Collect New Datasets Or Extend Existing Ones? 24 Oct 2023 · 1 repository · arXiv:2310.15758Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Generative Pre-trained Transformer for Vietnamese Community-based COVID-19 Question Answering 23 Oct 2023 · 0 repositories · arXiv:2310.14602
-
Prefix-Tuning Based Unsupervised Text Style Transfer 23 Oct 2023 · 0 repositories · arXiv:2310.14599
-
Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models 23 Oct 2023 · 2 repositories · arXiv:2310.14491Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Foundation Model's Embedded Representations May Detect Distribution Shift 20 Oct 2023 · 0 repositories · arXiv:2310.13836
-
Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model 19 Oct 2023 · 1 repository · arXiv:2310.12611Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Efficient Model-Agnostic Multi-Group Equivariant Networks 14 Oct 2023 · 0 repositories · arXiv:2310.09675
-
From Words and Exercises to Wellness: Farsi Chatbot for Self-Attachment Technique 13 Oct 2023 · 0 repositories · arXiv:2310.09362
-
Promptor: A Conversational and Autonomous Prompt Generation Agent for Intelligent Text Entry Techniques 12 Oct 2023 · 0 repositories · arXiv:2310.08101
-
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog 11 Oct 2023 · 2 repositories · arXiv:2310.07259Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Humans and language models diverge when predicting repeating text 10 Oct 2023 · 1 repository · arXiv:2310.06408Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Distantly-Supervised Joint Extraction with Noise-Robust Learning 8 Oct 2023 · 1 repository · arXiv:2310.04994
-
Do self-supervised speech and language models extract similar representations as human brain? 7 Oct 2023 · 0 repositories · arXiv:2310.04645