Methods › Natural Language Processing › Autoregressive Transformers › GPT-2 › Papers, page 2
GPT-2
Papers archive 2025-07-28
archive papers tagged: 768 · with a code link: 339 · where Syntology ran a sample: 125 (99 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (125 of 768 tagged: 99 with a run with no instrument failure, 26 where every run was a failure of Syntology's instrument)
Page 2 of 8: papers 101 to 200 of 768, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Improving Next Tokens via Second-Last Predictions with Generate and Refine 23 Nov 2024 · 0 repositories · arXiv:2411.15661
-
MARS: Unleashing the Power of Variance Reduction for Training Large Models 15 Nov 2024 · 2 repositories · arXiv:2411.10438Syntology official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Take Package as Language: Anomaly Detection Using Transformer 15 Nov 2024 · 0 repositories · arXiv:2412.04473
-
BabyLM Challenge: Exploring the Effect of Variation Sets on Language Model Training Efficiency 14 Nov 2024 · 0 repositories · arXiv:2411.09587
-
Autonomous Droplet Microfluidic Design Framework with Large Language Models 11 Nov 2024 · 1 repository · arXiv:2411.06691
-
On Active Privacy Auditing in Supervised Fine-tuning for White-Box Language Models 11 Nov 2024 · 0 repositories · arXiv:2411.07070
-
Prompt-Efficient Fine-Tuning for GPT-like Deep Models to Reduce Hallucination and to Improve Reproducibility in Scientific Text Generation Using Stochastic Optimisation Techniques 10 Nov 2024 · 0 repositories · arXiv:2411.06445
-
Adversarial Robustness of In-Context Learning in Transformers for Linear Regression 7 Nov 2024 · 0 repositories · arXiv:2411.05189
-
Can Custom Models Learn In-Context? An Exploration of Hybrid Architecture Performance on In-Context Learning Tasks 6 Nov 2024 · 1 repository · arXiv:2411.03945
-
Towards Interpreting Language Models: A Case Study in Multi-Hop Reasoning 6 Nov 2024 · 1 repository · arXiv:2411.05037Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Enriching Tabular Data with Contextual LLM Embeddings: A Comprehensive Ablation Study for Ensemble Classifiers 3 Nov 2024 · 0 repositories · arXiv:2411.01645
-
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders 2 Nov 2024 · 1 repository · arXiv:2411.01220Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
BLAST: Block-Level Adaptive Structured Matrices for Efficient Deep Neural Network Inference 28 Oct 2024 · 1 repository · arXiv:2410.21262
-
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics 28 Oct 2024 · 0 repositories · arXiv:2410.21353
-
MultiTok: Variable-Length Tokenization for Efficient LLMs Adapted from LZW Compression 28 Oct 2024 · 1 repository · arXiv:2410.21548
-
Scaling up Masked Diffusion Models on Text 24 Oct 2024 · 1 repository · arXiv:2410.18514Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Differentially Private Learning Needs Better Model Initialization and Self-Distillation 23 Oct 2024 · 1 repository · arXiv:2410.17566
-
DNAHLM -- DNA sequence and Human Language mixed large language Model 22 Oct 2024 · 1 repository · arXiv:2410.16917
-
Exploring Possibilities of AI-Powered Legal Assistance in Bangladesh through Large Language Modeling 22 Oct 2024 · 1 repository · arXiv:2410.17210
-
Improving Neuron-level Interpretability with White-box Language Models 21 Oct 2024 · 0 repositories · arXiv:2410.16443
-
Bias Amplification: Language Models as Increasingly Biased Media 19 Oct 2024 · 0 repositories · arXiv:2410.15234
-
Training Compute-Optimal Vision Transformers for Brain Encoding 17 Oct 2024 · 0 repositories · arXiv:2410.19810
-
Context-Scaling versus Task-Scaling in In-Context Learning 16 Oct 2024 · 0 repositories · arXiv:2410.12783
-
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression 16 Oct 2024 · 0 repositories · arXiv:2410.12707
-
Kallini et al. (2024) do not compare impossible languages with constituency-based ones 16 Oct 2024 · 0 repositories · arXiv:2410.12271
-
Can In-context Learning Really Generalize to Out-of-distribution Tasks? 13 Oct 2024 · 0 repositories · arXiv:2410.09695
-
Adam Exploits ℓ_∞-geometry of Loss Landscape via Coordinate-wise Adaptivity 10 Oct 2024 · 1 repository · arXiv:2410.08198Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
SAGE: Scalable Ground Truth Evaluations for Large Sparse Autoencoders 9 Oct 2024 · 0 repositories · arXiv:2410.07456
-
A second-order-like optimizer with adaptive gradient scaling for deep learning 8 Oct 2024 · 1 repository · arXiv:2410.05871
-
Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning 8 Oct 2024 · 1 repository · arXiv:2410.06101Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
LPZero: Language Model Zero-cost Proxy Search from Zero 7 Oct 2024 · 0 repositories · arXiv:2410.04808
-
How Language Models Prioritize Contextual Grammatical Cues? 4 Oct 2024 · 1 repository · arXiv:2410.03447
-
Sparse Attention Decomposition Applied to Circuit Tracing 1 Oct 2024 · 1 repository · arXiv:2410.00340
-
Evaluating the fairness of task-adaptive pretraining on unlabeled test data before few-shot text classification 30 Sep 2024 · 1 repository · arXiv:2410.00179
-
Modelando procesos cognitivos de la lectura natural con GPT-2 30 Sep 2024 · 0 repositories · arXiv:2409.20174
-
Analog In-Memory Computing Attention Mechanism for Fast and Energy-Efficient Large Language Models 28 Sep 2024 · 1 repository · arXiv:2409.19315
-
Experimental Evaluation of Machine Learning Models for Goal-oriented Customer Service Chatbot with Pipeline Architecture 27 Sep 2024 · 0 repositories · arXiv:2409.18568
-
Comparing Unidirectional, Bidirectional, and Word2vec Models for Discovering Vulnerabilities in Compiled Lifted Code 26 Sep 2024 · 0 repositories · arXiv:2409.17513
-
SDBA: A Stealthy and Long-Lasting Durable Backdoor Attack in Federated Learning 23 Sep 2024 · 1 repository · arXiv:2409.14805
-
Drift to Remember 21 Sep 2024 · 0 repositories · arXiv:2409.13997
-
Loop Neural Networks for Parameter Sharing 21 Sep 2024 · 0 repositories · arXiv:2409.14199
-
HUT: A More Computation Efficient Fine-Tuning Method With Hadamard Updated Transformation 20 Sep 2024 · 0 repositories · arXiv:2409.13501
-
A Unified Framework to Classify Business Activities into International Standard Industrial Classification through Large Language Models for Circular Economy 17 Sep 2024 · 0 repositories · arXiv:2409.18988
-
Stable Language Model Pre-training by Reducing Embedding Variability 12 Sep 2024 · 0 repositories · arXiv:2409.07787
-
Accelerating Large Language Model Pretraining via LFR Pedagogy: Learn, Focus, and Review 10 Sep 2024 · 0 repositories · arXiv:2409.06131
-
Bypassing DARCY Defense: Indistinguishable Universal Adversarial Triggers 5 Sep 2024 · 0 repositories · arXiv:2409.03183
-
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small 5 Sep 2024 · 1 repository · arXiv:2409.04478Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering 30 Aug 2024 · 0 repositories · arXiv:2408.17006
-
LLaVA-Chef: A Multi-modal Generative Model for Food Recipes 29 Aug 2024 · 1 repository · arXiv:2408.16889
-
Enhancing Multi-hop Reasoning through Knowledge Erasure in Large Language Model Editing 22 Aug 2024 · 0 repositories · arXiv:2408.12456
-
Mixed Sparsity Training: Achieving 4× FLOP Reduction for Transformer Pretraining 21 Aug 2024 · 0 repositories · arXiv:2408.11746
-
Tracing Privacy Leakage of Language Models to Training Data via Adjusted Influence Functions 20 Aug 2024 · 0 repositories · arXiv:2408.10468
-
ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA 19 Aug 2024 · 1 repository · arXiv:2408.11869
-
Pragmatic inference of scalar implicature by LLMs 13 Aug 2024 · 0 repositories · arXiv:2408.06673
-
Retrieval-augmented code completion for local projects using large language models 9 Aug 2024 · 0 repositories · arXiv:2408.05026
-
Transformer Explainer: Interactive Learning of Text-Generative Models 8 Aug 2024 · 1 repository · arXiv:2408.04619
-
Image-to-LaTeX Converter for Mathematical Formulas and Text 7 Aug 2024 · 1 repository · arXiv:2408.04015
-
Is Child-Directed Speech Effective Training Data for Language Models? 7 Aug 2024 · 1 repository · arXiv:2408.03617Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
SocFedGPT: Federated GPT-based Adaptive Content Filtering System Leveraging User Interactions in Social Networks 7 Aug 2024 · 0 repositories · arXiv:2408.05243
-
TrafficGPT: An LLM Approach for Open-Set Encrypted Traffic Classification 6 Aug 2024 · 1 repository
-
ML-EAT: A Multilevel Embedding Association Test for Interpretable and Transparent Social Science 4 Aug 2024 · 1 repository · arXiv:2408.01966
-
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs 29 Jul 2024 · 1 repository · arXiv:2407.20177Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability 29 Jul 2024 · 1 repository · arXiv:2407.19842
-
Mechanistic interpretability of large language models with applications to the financial services industry 15 Jul 2024 · 0 repositories · arXiv:2407.11215
-
Generating In-store Customer Journeys from Scratch with GPT Architectures 13 Jul 2024 · 0 repositories · arXiv:2407.11081
-
ASTPrompter: Weakly Supervised Automated Language Model Red-Teaming to Identify Low-Perplexity Toxic Prompts 12 Jul 2024 · 1 repository · arXiv:2407.09447
-
ROSA: Random Subspace Adaptation for Efficient Fine-Tuning 10 Jul 2024 · 1 repository · arXiv:2407.07802
-
Raply: A profanity-mitigated rap generator 9 Jul 2024 · 0 repositories · arXiv:2407.06941
-
Slice-100K: A Multimodal Dataset for Extrusion-based 3D Printing 4 Jul 2024 · 0 repositories · arXiv:2407.04180
-
ObfuscaTune: Obfuscated Offsite Fine-tuning and Inference of Proprietary LLMs on Private Datasets 3 Jul 2024 · 0 repositories · arXiv:2407.02960
-
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules 30 Jun 2024 · 1 repository · arXiv:2407.00599
-
Machine Learning Predictors for Min-Entropy Estimation 28 Jun 2024 · 1 repository · arXiv:2406.19983
-
ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting 28 Jun 2024 · 0 repositories · arXiv:2406.19976
-
Fine-tuned network relies on generic representation to solve unseen cognitive task 27 Jun 2024 · 0 repositories · arXiv:2406.18926
-
Improving Entity Recognition Using Ensembles of Deep Learning and Fine-tuned Large Language Models: A Case Study on Adverse Event Extraction from Multiple Sources 26 Jun 2024 · 0 repositories · arXiv:2406.18049
-
Interpreting Attention Layer Outputs with Sparse Autoencoders 25 Jun 2024 · 1 repository · arXiv:2406.17759Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Understanding Language Model Circuits through Knowledge Editing 25 Jun 2024 · 0 repositories · arXiv:2406.17241
-
Finding Transformer Circuits with Edge Pruning 24 Jun 2024 · 1 repository · arXiv:2406.16778Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
Anime Popularity Prediction Before Huge Investments: a Multimodal Approach Using Deep Learning 21 Jun 2024 · 0 repositories · arXiv:2406.16961
-
What Makes Two Language Models Think Alike? 18 Jun 2024 · 0 repositories · arXiv:2406.12620
-
Promises, Outlooks and Challenges of Diffusion Language Modeling 17 Jun 2024 · 0 repositories · arXiv:2406.11473
-
ShareLoRA: Parameter Efficient and Robust Large Language Model Fine-tuning via Shared Low-Rank Adaptation 16 Jun 2024 · 1 repository · arXiv:2406.10785Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation 15 Jun 2024 · 1 repository · arXiv:2406.10591
-
Towards Efficient Pareto Set Approximation via Mixture of Experts Based Model Fusion 14 Jun 2024 · 1 repository · arXiv:2406.09770
-
A More Practical Approach to Machine Unlearning 13 Jun 2024 · 0 repositories · arXiv:2406.09391
-
Talking Heads: Understanding Inter-layer Communication in Transformer Language Models 13 Jun 2024 · 0 repositories · arXiv:2406.09519
-
Compute Better Spent: Replacing Dense Layers with Structured Matrices 10 Jun 2024 · 1 repository · arXiv:2406.06248Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Critical Phase Transition in Large Language Models 8 Jun 2024 · 0 repositories · arXiv:2406.05335
-
VTrans: Accelerating Transformer Compression with Variational Information Bottleneck based Pruning 7 Jun 2024 · 0 repositories · arXiv:2406.05276
-
Simplified and Generalized Masked Diffusion for Discrete Data 6 Jun 2024 · 1 repository · arXiv:2406.04329Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 15 unverified (of 21 harvested samples)
-
Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data 6 Jun 2024 · 2 repositories · arXiv:2406.03736Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 1 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers 5 Jun 2024 · 0 repositories · arXiv:2406.02847
-
RICo: Reddit ideological communities 5 Jun 2024 · 1 repository
-
Too Big to Fail: Larger Language Models are Disproportionately Resilient to Induction of Dementia-Related Linguistic Anomalies 5 Jun 2024 · 1 repository · arXiv:2406.02830
-
In-Context Learning of Physical Properties: Few-Shot Adaptation to Out-of-Distribution Molecular Graphs 3 Jun 2024 · 0 repositories · arXiv:2406.01808
-
LOLAMEME: Logic, Language, Memory, Mechanistic Framework 31 May 2024 · 0 repositories · arXiv:2406.02592
-
Knowledge Graph Tuning: Real-time Large Language Model Personalization based on Human Feedback 30 May 2024 · 0 repositories · arXiv:2405.19686
-
Efficient Model-agnostic Alignment via Bayesian Persuasion 29 May 2024 · 0 repositories · arXiv:2405.18718
-
LMO-DP: Optimizing the Randomization Mechanism for Differentially Private Fine-Tuning (Large) Language Models 29 May 2024 · 0 repositories · arXiv:2405.18776
-
Are PPO-ed Language Models Hackable? 28 May 2024 · 0 repositories · arXiv:2406.02577