Methods › Natural Language Processing › Transformers › DistilBERT › Papers, page 2
DistilBERT
Papers archive 2025-07-28
archive papers tagged: 166 · with a code link: 63 · where Syntology ran a sample: 9 (6 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (9 of 166 tagged: 6 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 166 of 166, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Fast Post-Training Pruning Framework for Transformers 29 Mar 2022 · 2 repositories · arXiv:2204.09656Syntology official (archive's flag): 1 ran · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Build a Robust QA System with Transformer-based Mixture of Experts 20 Mar 2022 · 1 repository · arXiv:2204.09598
-
Using Word Embeddings to Analyze Protests News 11 Mar 2022 · 0 repositories · arXiv:2203.05875
-
Comparison of biomedical relationship extraction methods and models for knowledge graph creation 5 Jan 2022 · 0 repositories · arXiv:2201.01647
-
Automatic Mixed-Precision Quantization Search of BERT 30 Dec 2021 · 0 repositories · arXiv:2112.14938
-
From Scattered Sources to Comprehensive Technology Landscape: A Recommendation-based Retrieval Approach 9 Dec 2021 · 0 repositories · arXiv:2112.04810
-
A Comparative Study of Transformers on Word Sense Disambiguation 30 Nov 2021 · 0 repositories · arXiv:2111.15417
-
Assessing the Coherence Modeling Capabilities of Pretrained Transformer-based Language Models 16 Nov 2021 · 0 repositories
-
Prune Once for All: Sparse Pre-Trained Language Models 10 Nov 2021 · 2 repositories · arXiv:2111.05754
-
Sexism Identification in Tweets and Gabs using Deep Neural Networks 5 Nov 2021 · 0 repositories · arXiv:2111.03612
-
Modeling Performance in Open-Domain Dialogue with PARADISE 21 Oct 2021 · 0 repositories · arXiv:2110.11164
-
ALL-IN-ONE: Multi-Task Learning BERT models for Evaluating Peer Assessments 8 Oct 2021 · 0 repositories · arXiv:2110.03895
-
Compressing Transformer-Based Sequence to Sequence Models With Pre-trained Autoencoders for Text Summarization 29 Sep 2021 · 0 repositories
-
Cross-Architecture Distillation Using Bidirectional CMOW Embeddings 29 Sep 2021 · 0 repositories
-
Improving Sentiment Classification Using 0-Shot Generated Labels for Custom Transformer Embeddings 29 Sep 2021 · 0 repositories
-
Specialized Transformers: Faster, Smaller and more Accurate NLP Models 29 Sep 2021 · 0 repositories
-
Improving Question Answering Performance Using Knowledge Distillation and Active Learning 26 Sep 2021 · 1 repository · arXiv:2109.12662Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
General Cross-Architecture Distillation of Pretrained Language Models into Matrix Embeddings 17 Sep 2021 · 1 repository · arXiv:2109.08449
-
Transformer-based Language Models for Factoid Question Answering at BioASQ9b 15 Sep 2021 · 1 repository · arXiv:2109.07185
-
Bag-of-Words vs. Graph vs. Sequence in Text Classification: Questioning the Necessity of Text-Graphs and the Surprising Strength of a Wide MLP 8 Sep 2021 · 2 repositories · arXiv:2109.03777
-
DKM: Differentiable K-Means Clustering Layer for Neural Network Compression 28 Aug 2021 · 0 repositories · arXiv:2108.12659
-
sigmoidF1: A Smooth F1 Score Surrogate Loss for Multilabel Classification 24 Aug 2021 · 1 repository · arXiv:2108.10566
-
Offensive Language and Hate Speech Detection with Deep Learning and Transfer Learning 6 Aug 2021 · 0 repositories · arXiv:2108.03305
-
AttesTable at SemEval-2021 Task 9: Extending Statement Verification with Tables for Unknown Class, and Semantic Evidence Finding 1 Aug 2021 · 1 repository
-
CSECU-DSG at SemEval-2021 Task 6: Orchestrating Multimodal Neural Architectures for Identifying Persuasion Techniques in Texts and Images 1 Aug 2021 · 0 repositories
-
UoR at SemEval-2021 Task 7: Utilizing Pre-trained DistilBERT Model and Multi-scale CNN for Humor Detection 1 Aug 2021 · 0 repositories
-
AutoTinyBERT: Automatic Hyper-parameter Optimization for Efficient Pre-trained Language Models 29 Jul 2021 · 1 repository · arXiv:2107.13686
-
LanguageRefer: Spatial-Language Model for 3D Visual Grounding 7 Jul 2021 · 0 repositories · arXiv:2107.03438
-
HONEST: Measuring Hurtful Sentence Completion in Language Models 1 Jun 2021 · 1 repository
-
Towards a Comprehensive Understanding and Accurate Evaluation of Societal Biases in Pre-Trained Transformers 1 Jun 2021 · 0 repositories
-
Accelerating BERT Inference for Sequence Labeling via Early-Exit 28 May 2021 · 1 repository · arXiv:2105.13878
-
Exploring Transformers in Emotion Recognition: a comparison of BERT, DistillBERT, RoBERTa, XLNet and ELECTRA 5 Apr 2021 · 0 repositories · arXiv:2104.02041
-
HLE-UPC at SemEval-2021 Task 5: Multi-Depth DistilBERT for Toxic Spans Detection 1 Apr 2021 · 1 repository · arXiv:2104.00639
-
Combat COVID-19 Infodemic Using Explainable Natural Language Processing Models 1 Mar 2021 · 0 repositories · arXiv:2103.00747
-
Parallelizing Legendre Memory Unit Training 22 Feb 2021 · 2 repositories · arXiv:2102.11417
-
Dancing along Battery: Enabling Transformer with Run-time Reconfigurability on Mobile Devices 12 Feb 2021 · 0 repositories · arXiv:2102.06336
-
Scaling Federated Learning for Fine-tuning of Large Language Models 1 Feb 2021 · 0 repositories · arXiv:2102.00875
-
Post-Training Weighted Quantization of Neural Networks for Language Models 1 Jan 2021 · 0 repositories
-
Yelp Review Rating Prediction: Machine Learning and Deep Learning Models 12 Dec 2020 · 1 repository · arXiv:2012.06690
-
Detecting Insincere Questions from Text: A Transfer Learning Approach 7 Dec 2020 · 1 repository · arXiv:2012.07587
-
Automated Detection of Cyberbullying Against Women and Immigrants and Cross-domain Adaptability 4 Dec 2020 · 0 repositories · arXiv:2012.02565
-
BERT at SemEval-2020 Task 8: Using BERT to Analyse Meme Emotions 1 Dec 2020 · 0 repositories
-
KAFK at SemEval-2020 Task 8: Extracting Features from Pre-trained Neural Networks to Classify Internet Memes 1 Dec 2020 · 1 repository
-
Transformer-Based Models for Automatic Identification of Argument Relations: A Cross-Domain Evaluation 26 Nov 2020 · 0 repositories · arXiv:2011.13187
-
Probing for Multilingual Numerical Understanding in Transformer-Based Language Models 13 Oct 2020 · 1 repository · arXiv:2010.06666
-
Chatbot Interaction with Artificial Intelligence: Human Data Augmentation with T5 and Language Transformer Ensemble for Text Classification 12 Oct 2020 · 0 repositories · arXiv:2010.05990
-
Compressing Transformer-Based Semantic Parsing Models using Compositional Code Embeddings 10 Oct 2020 · 0 repositories · arXiv:2010.05002
-
Deep Learning Meets Projective Clustering 8 Oct 2020 · 0 repositories · arXiv:2010.04290
-
AxFormer: Accuracy-driven Approximation of Transformers for Faster, Smaller and more Accurate NLP Models 7 Oct 2020 · 1 repository · arXiv:2010.03688
-
Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation 6 Oct 2020 · 1 repository · arXiv:2010.02705
-
Compositional and Lexical Semantics in RoBERTa, BERT and DistilBERT: A Case Study on CoQA 17 Sep 2020 · 0 repositories · arXiv:2009.08257
-
Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning 17 Sep 2020 · 0 repositories · arXiv:2009.08065
-
Compressed Deep Networks: Goodbye SVD, Hello Robust Low-Rank Approximation 11 Sep 2020 · 1 repository · arXiv:2009.05647
-
Comparative Study of Language Models on Cross-Domain Data with Model Agnostic Explainability 9 Sep 2020 · 0 repositories · arXiv:2009.04095
-
Sentimental LIAR: Extended Corpus and Deep Learning Models for Fake Claim Classification 1 Sep 2020 · 2 repositories · arXiv:2009.01047
-
Top2Vec: Distributed Representations of Topics 19 Aug 2020 · 2 repositories · arXiv:2008.09470Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Cooking Is All About People: Comment Classification On Cookery Channels Using BERT and Classification Models (Malayalam-English Mix-Code) 15 Jun 2020 · 0 repositories · arXiv:2007.04249
-
Accelerating Natural Language Understanding in Task-Oriented Dialog 5 Jun 2020 · 1 repository · arXiv:2006.03701
-
LRG at SemEval-2020 Task 7: Assessing the Ability of BERT and Derivative Models to Perform Short-Edits based Humor Grading 31 May 2020 · 0 repositories · arXiv:2006.00607
-
Language Representation Models for Fine-Grained Sentiment Classification 27 May 2020 · 1 repository · arXiv:2005.13619
-
Establishing Baselines for Text Classification in Low-Resource Languages 5 May 2020 · 1 repository · arXiv:2005.02068
-
Analyzing ELMo and DistilBERT on Socio-political News Classification 1 May 2020 · 0 repositories
-
On the Effect of Dropping Layers of Pre-trained Transformer Models 8 Apr 2020 · 4 repositories · arXiv:2004.03844Syntology official (archive's flag): 5 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Continual Domain-Tuning for Pretrained Language Models 5 Apr 2020 · 0 repositories · arXiv:2004.02288
-
Deep Entity Matching with Pre-Trained Language Models 1 Apr 2020 · 1 repository · arXiv:2004.00584Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter 2 Oct 2019 · 37 repositories · arXiv:1910.01108Syntology official (archive's flag): 1 ran · 21 ran (of which 5 constructed an object rather than computing a result; 13 with no instrument failure: 3 honoured, 1 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 6 unverified (of 27 harvested samples) · 2 pointer-only (licence)