Methods › Natural Language Processing › Tokenizers › SentencePiece › Papers, page 5
SentencePiece
Papers archive 2025-07-28
archive papers tagged: 908 · with a code link: 438 · where Syntology ran a sample: 111 (96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 908 tagged: 96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 5 of 10: papers 401 to 500 of 908, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
NeuralMind-UNICAMP at 2022 TREC NeuCLIR: Large Boring Rerankers for Cross-lingual Retrieval 28 Mar 2023 · 1 repository · arXiv:2303.16145
-
One Adapter for All Programming Languages? Adapter Tuning for Code Search and Summarization 28 Mar 2023 · 1 repository · arXiv:2303.15822
-
Automatic Generation of Multiple-Choice Questions 25 Mar 2023 · 0 repositories · arXiv:2303.14576
-
DBLP-QuAD: A Question Answering Dataset over the DBLP Scholarly Knowledge Graph 23 Mar 2023 · 1 repository · arXiv:2303.13351
-
GETT-QA: Graph Embedding based T2T Transformer for Knowledge Graph Question Answering 23 Mar 2023 · 1 repository · arXiv:2303.13284
-
Open-source Frame Semantic Parsing 22 Mar 2023 · 1 repository · arXiv:2303.12788
-
Bangla Grammatical Error Detection Using T5 Transformer Model 19 Mar 2023 · 2 repositories · arXiv:2303.10612
-
Exploring Distributional Shifts in Large Language Models for Code Analysis 16 Mar 2023 · 0 repositories · arXiv:2303.09128
-
TypeT5: Seq2seq Type Inference using Static Analysis 16 Mar 2023 · 1 repository · arXiv:2303.09564Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
PRESTO: A Multilingual Dataset for Parsing Realistic Task-Oriented Dialogs 15 Mar 2023 · 1 repository · arXiv:2303.08954
-
Proactive Prioritization of App Issues via Contrastive Learning 12 Mar 2023 · 1 repository · arXiv:2303.06586
-
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT 7 Mar 2023 · 1 repository · arXiv:2303.04226
-
Spelling convention sensitivity in neural language models 6 Mar 2023 · 0 repositories · arXiv:2303.03457
-
TrojText: Test-time Invisible Textual Trojan Insertion 3 Mar 2023 · 1 repository · arXiv:2303.02242Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space 1 Mar 2023 · 0 repositories · arXiv:2303.00456
-
Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation 28 Feb 2023 · 1 repository · arXiv:2302.14220
-
Choice Fusion as Knowledge for Zero-Shot Dialogue State Tracking 25 Feb 2023 · 1 repository · arXiv:2302.13013
-
Prompt-based Learning for Text Readability Assessment 25 Feb 2023 · 1 repository · arXiv:2302.13139
-
HULAT at SemEval-2023 Task 10: Data augmentation for pre-trained transformers applied to the detection of sexism in social media 24 Feb 2023 · 1 repository · arXiv:2302.12840
-
Does Deep Learning Learn to Abstract? A Systematic Probing Framework 23 Feb 2023 · 1 repository · arXiv:2302.11978
-
EVJVQA Challenge: Multilingual Visual Question Answering 23 Feb 2023 · 0 repositories · arXiv:2302.11752
-
BBT-Fin: Comprehensive Construction of Chinese Financial Domain Pre-trained Language Model, Corpus and Benchmark 18 Feb 2023 · 2 repositories · arXiv:2302.09432
-
PAC Prediction Sets for Large Language Models of Code 17 Feb 2023 · 1 repository · arXiv:2302.08703
-
Commonsense Reasoning for Conversational AI: A Survey of the State of the Art 15 Feb 2023 · 0 repositories · arXiv:2302.07926
-
Few-shot learning approaches for classifying low resource domain specific software requirements 14 Feb 2023 · 0 repositories · arXiv:2302.06951
-
Linguistic ambiguity analysis in ChatGPT 13 Feb 2023 · 0 repositories · arXiv:2302.06426
-
STREET: A Multi-Task Structured Reasoning and Explanation Benchmark 13 Feb 2023 · 0 repositories · arXiv:2302.06729
-
Controllable Lexical Simplification for English 6 Feb 2023 · 1 repository · arXiv:2302.02900
-
idT5: Indonesian Version of Multilingual T5 Transformer 2 Feb 2023 · 0 repositories · arXiv:2302.00856
-
HunSum-1: an Abstractive Summarization Dataset for Hungarian 1 Feb 2023 · 1 repository · arXiv:2302.00455
-
FLAME: A small language model for spreadsheet formulas 31 Jan 2023 · 0 repositories · arXiv:2301.13779
-
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning 31 Jan 2023 · 1 repository · arXiv:2301.13688
-
Specializing Smaller Language Models towards Multi-Step Reasoning 30 Jan 2023 · 2 repositories · arXiv:2301.12726Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Progressive Prompts: Continual Learning for Language Models 29 Jan 2023 · 2 repositories · arXiv:2301.12314Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 11 harvested samples)
-
Schema-Guided Semantic Accuracy: Faithfulness in Task-Oriented Dialogue Response Generation 29 Jan 2023 · 1 repository · arXiv:2301.12568
-
Bipol: Multi-axes Evaluation of Bias with Explainability in Benchmark Datasets 28 Jan 2023 · 2 repositories · arXiv:2301.12139
-
A benchmark for toxic comment classification on Civil Comments dataset 26 Jan 2023 · 1 repository · arXiv:2301.11125
-
A Stability Analysis of Fine-Tuning a Pre-Trained Model 24 Jan 2023 · 0 repositories · arXiv:2301.09820
-
Multitask Instruction-based Prompting for Fallacy Recognition 24 Jan 2023 · 0 repositories · arXiv:2301.09992
-
Graphix-T5: Mixing Pre-Trained Transformers with Graph-Aware Layers for Text-to-SQL Parsing 18 Jan 2023 · 1 repository · arXiv:2301.07507
-
There is No Big Brother or Small Brother: Knowledge Infusion in Language Models for Link Prediction and Question Answering 10 Jan 2023 · 2 repositories · arXiv:2301.04013
-
Generative Antibody Design for Complementary Chain Pairing Sequences through Encoder-Decoder Language Model 6 Jan 2023 · 0 repositories · arXiv:2301.02748
-
Extending Source Code Pre-Trained Language Models to Summarise Decompiled Binaries 4 Jan 2023 · 1 repository · arXiv:2301.01701
-
Transformer Based Geocoding 2 Jan 2023 · 0 repositories · arXiv:2301.01170
-
Inconsistencies in Masked Language Models 30 Dec 2022 · 1 repository · arXiv:2301.00068
-
Error syntax aware augmentation of feedback comment generation dataset 29 Dec 2022 · 0 repositories · arXiv:2212.14293
-
Analyzing Semantic Faithfulness of Language Models via Input Intervention on Question Answering 21 Dec 2022 · 1 repository · arXiv:2212.10696
-
Uncontrolled Lexical Exposure Leads to Overestimation of Compositional Generalization in Pretrained Models 21 Dec 2022 · 1 repository · arXiv:2212.10769
-
ByGPT5: End-to-End Style-conditioned Poetry Generation with Token-free Language Models 20 Dec 2022 · 1 repository · arXiv:2212.10474Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Do language models have coherent mental models of everyday things? 20 Dec 2022 · 1 repository · arXiv:2212.10029Syntology official: harvested, nothing ran · 0 ran · 5 unverified (of 5 harvested samples)
-
KronA: Parameter Efficient Tuning with Kronecker Adapter 20 Dec 2022 · 0 repositories · arXiv:2212.10650
-
Dissecting Transformer Length Extrapolation via the Lens of Receptive Field Analysis 20 Dec 2022 · 0 repositories · arXiv:2212.10356
-
T-Projection: High Quality Annotation Projection for Sequence Labeling Tasks 20 Dec 2022 · 2 repositories · arXiv:2212.10548
-
Do CoNLL-2003 Named Entity Taggers Still Work Well in 2023? 19 Dec 2022 · 1 repository · arXiv:2212.09747
-
MIGA: A Unified Multi-task Generation Framework for Conversational Text-to-SQL 19 Dec 2022 · 0 repositories · arXiv:2212.09278
-
Mu²SLAM: Multitask, Multilingual Speech and Language Models 19 Dec 2022 · 0 repositories · arXiv:2212.09553
-
Multilingual Sequence-to-Sequence Models for Hebrew NLP 19 Dec 2022 · 0 repositories · arXiv:2212.09682
-
LegalRelectra: Mixed-domain Language Modeling for Long-range Legal Text Comprehension 16 Dec 2022 · 0 repositories · arXiv:2212.08204
-
Teaching Small Language Models to Reason 16 Dec 2022 · 0 repositories · arXiv:2212.08410
-
FiDO: Fusion-in-Decoder optimized for stronger performance and faster inference 15 Dec 2022 · 0 repositories · arXiv:2212.08153
-
Visually-augmented pretrained language models for NLP tasks without images 15 Dec 2022 · 1 repository · arXiv:2212.07937
-
Evaluating Byte and Wordpiece Level Models for Massively Multilingual Semantic Parsing 14 Dec 2022 · 0 repositories · arXiv:2212.07223
-
T5Score: Discriminative Fine-tuning of Generative Evaluation Metrics 12 Dec 2022 · 2 repositories · arXiv:2212.05726
-
Punctuation Restoration for Singaporean Spoken Languages: English, Malay, and Mandarin 10 Dec 2022 · 1 repository · arXiv:2212.05356
-
Sparse Upcycling: Training Mixture-of-Experts from Dense Checkpoints 9 Dec 2022 · 1 repository · arXiv:2212.05055Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
TRBLLmaker -- Transformer Reads Between Lyrics Lines maker 9 Dec 2022 · 0 repositories · arXiv:2212.04917
-
Hierarchical multimodal transformers for Multi-Page DocVQA 7 Dec 2022 · 1 repository · arXiv:2212.05935
-
Learning-To-Embed: Adopting Transformer based models for E-commerce Products Representation Learning 7 Dec 2022 · 0 repositories · arXiv:2212.03725
-
Controlled Text Generation using T5 based Encoder-Decoder Soft Prompt Tuning and Analysis of the Utility of Generated Text in AI 6 Dec 2022 · 0 repositories · arXiv:2212.02924
-
Languages You Know Influence Those You Learn: Impact of Language Characteristics on Multi-Lingual Text-to-Text Transfer 4 Dec 2022 · 0 repositories · arXiv:2212.01757
-
Global memory transformer for processing long documents 3 Dec 2022 · 0 repositories · arXiv:2212.01650
-
FECAM: Frequency Enhanced Channel Attention Mechanism for Time Series Forecasting 2 Dec 2022 · 1 repository · arXiv:2212.01209
-
Data-Efficient Finetuning Using Cross-Task Nearest Neighbors 1 Dec 2022 · 1 repository · arXiv:2212.00196Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Extending the Subwording Model of Multilingual Pretrained Models for New Languages 29 Nov 2022 · 1 repository · arXiv:2211.15965
-
Detect-Localize-Repair: A Unified Framework for Learning to Debug with CodeT5 27 Nov 2022 · 0 repositories · arXiv:2211.14875
-
Coreference Resolution through a seq2seq Transition-Based System 22 Nov 2022 · 1 repository · arXiv:2211.12142
-
HyperTuning: Toward Adapting Large Language Models without Back-propagation 22 Nov 2022 · 0 repositories · arXiv:2211.12485
-
Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language Inference 21 Nov 2022 · 0 repositories · arXiv:2211.11875
-
UnifiedABSA: A Unified ABSA Framework Based on Multi-task Instruction Tuning 20 Nov 2022 · 0 repositories · arXiv:2211.10986
-
GLAMI-1M: A Multilingual Image-Text Fashion Dataset 17 Nov 2022 · 1 repository · arXiv:2211.14451
-
Unified Question Answering in Slovene 16 Nov 2022 · 1 repository · arXiv:2211.09159
-
Breakpoint Transformers for Modeling and Tracking Intermediate Beliefs 15 Nov 2022 · 1 repository · arXiv:2211.07950Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Empowering Language Models with Knowledge Graph Reasoning for Question Answering 15 Nov 2022 · 0 repositories · arXiv:2211.08380
-
CST5: Data Augmentation for Code-Switched Semantic Parsing 14 Nov 2022 · 1 repository · arXiv:2211.07514
-
Technological taxonomies for hypernym and hyponym retrieval in patent texts 14 Nov 2022 · 1 repository · arXiv:2212.06039
-
Dark patterns in e-commerce: a dataset and its baseline evaluations 12 Nov 2022 · 1 repository · arXiv:2211.06543
-
DocuT5: Seq2seq SQL Generation with Table Documentation 11 Nov 2022 · 0 repositories · arXiv:2211.06193
-
Assistive Completion of Agrammatic Aphasic Sentences: A Transfer Learning Approach using Neurolinguistics-based Synthetic Dataset 10 Nov 2022 · 0 repositories · arXiv:2211.05557
-
Large Language Models with Controllable Working Memory 9 Nov 2022 · 0 repositories · arXiv:2211.05110
-
Conciseness: An Overlooked Language Task 8 Nov 2022 · 0 repositories · arXiv:2211.04126
-
Crosslingual Generalization through Multitask Finetuning 3 Nov 2022 · 1 repository · arXiv:2211.01786Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Using Large Pre-Trained Language Model to Assist FDA in Premarket Medical Device 3 Nov 2022 · 0 repositories · arXiv:2212.01217
-
eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers 2 Nov 2022 · 2 repositories · arXiv:2211.01324Syntology 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
FRSUM: Towards Faithful Abstractive Summarization via Enhancing Factual Robustness 1 Nov 2022 · 0 repositories · arXiv:2211.00294
-
Two-stage LLM Fine-tuning with Less Specialization and More Generalization 1 Nov 2022 · 0 repositories · arXiv:2211.00635
-
T5lephone: Bridging Speech and Text Self-supervised Models for Spoken Language Understanding via Phoneme level T5 1 Nov 2022 · 1 repository · arXiv:2211.00586
-
Probing for targeted syntactic knowledge through grammatical error detection 28 Oct 2022 · 1 repository · arXiv:2210.16228
-
Beyond English-Centric Bitexts for Better Multilingual Language Representation Learning 26 Oct 2022 · 0 repositories · arXiv:2210.14867
-
Leveraging Affirmative Interpretations from Negation Improves Natural Language Understanding 26 Oct 2022 · 1 repository · arXiv:2210.14486
-
Amos: An Adam-style Optimizer with Adaptive Weight Decay towards Model-Oriented Scale 21 Oct 2022 · 1 repository · arXiv:2210.11693Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 18 harvested samples)