Methods › Natural Language Processing › Tokenizers › SentencePiece › Papers, page 3
SentencePiece
Papers archive 2025-07-28
archive papers tagged: 908 · with a code link: 438 · where Syntology ran a sample: 111 (96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (111 of 908 tagged: 96 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 3 of 10: papers 201 to 300 of 908, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A synthetic data approach for domain generalization of NLI models 19 Feb 2024 · 0 repositories · arXiv:2402.12368
-
Key ingredients for effective zero-shot cross-lingual knowledge transfer in generative tasks 19 Feb 2024 · 0 repositories · arXiv:2402.12279
-
Emerging Opportunities of Using Large Language Models for Translation Between Drug Molecules and Indications 14 Feb 2024 · 0 repositories · arXiv:2402.09588
-
FGeo-TP: A Language Model-Enhanced Solver for Geometry Problems 14 Feb 2024 · 0 repositories · arXiv:2402.09047
-
Eliciting Personality Traits in Large Language Models 13 Feb 2024 · 0 repositories · arXiv:2402.08341
-
Improving Black-box Robustness with In-Context Rewriting 13 Feb 2024 · 1 repository · arXiv:2402.08225
-
InkSight: Offline-to-Online Handwriting Conversion by Learning to Read and Write 8 Feb 2024 · 1 repository · arXiv:2402.05804
-
Lens: A Foundation Model for Network Traffic 6 Feb 2024 · 0 repositories · arXiv:2402.03646
-
CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks 2 Feb 2024 · 0 repositories · arXiv:2402.01176
-
What Will My Model Forget? Forecasting Forgotten Examples in Language Model Refinement 2 Feb 2024 · 0 repositories · arXiv:2402.01865
-
Improving Semantic Control in Discrete Latent Spaces with Transformer Quantized Variational Autoencoders 1 Feb 2024 · 1 repository · arXiv:2402.00723
-
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs 30 Jan 2024 · 1 repository · arXiv:2401.16638
-
ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling Tasks 29 Jan 2024 · 1 repository · arXiv:2401.16589
-
LPNL: Scalable Link Prediction with Large Language Models 24 Jan 2024 · 0 repositories · arXiv:2401.13227
-
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference 22 Jan 2024 · 1 repository · arXiv:2401.12200Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
The Right Model for the Job: An Evaluation of Legal Multi-Label Classification Baselines 22 Jan 2024 · 0 repositories · arXiv:2401.11852
-
Finding a Needle in the Adversarial Haystack: A Targeted Paraphrasing Approach For Uncovering Edge Cases with Minimal Distribution Distortion 21 Jan 2024 · 1 repository · arXiv:2401.11373
-
LangBridge: Multilingual Reasoning Without Multilingual Supervision 19 Jan 2024 · 1 repository · arXiv:2401.10695Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 7 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Deciphering Textual Authenticity: A Generalized Strategy through the Lens of Large Language Semantics for Detecting Human vs. Machine-Generated Text 17 Jan 2024 · 1 repository · arXiv:2401.09407
-
Mapping Transformer Leveraged Embeddings for Cross-Lingual Document Representation 12 Jan 2024 · 1 repository · arXiv:2401.06583
-
PizzaCommonSense: Learning to Model Commonsense Reasoning about Intermediate Steps in Cooking Recipes 12 Jan 2024 · 1 repository · arXiv:2401.06930
-
An Assessment on Comprehending Mental Health through Large Language Models 9 Jan 2024 · 0 repositories · arXiv:2401.04592
-
AST-T5: Structure-Aware Pretraining for Code Generation and Understanding 5 Jan 2024 · 1 repository · arXiv:2401.03003Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Are LLMs Robust for Spoken Dialogues? 4 Jan 2024 · 0 repositories · arXiv:2401.02297
-
Large Language Models in Mental Health Care: a Scoping Review 1 Jan 2024 · 0 repositories · arXiv:2401.02984
-
Scaling Down, LiTting Up: Efficient Zero-Shot Listwise Reranking with Seq2seq Encoder-Decoder Models 26 Dec 2023 · 2 repositories · arXiv:2312.16098
-
Numerical Reasoning for Financial Reports 22 Dec 2023 · 1 repository · arXiv:2312.14870
-
Theory of Hallucinations based on Equivariance 22 Dec 2023 · 0 repositories · arXiv:2312.14504
-
Aspect-Based Sentiment Analysis with Explicit Sentiment Augmentations 18 Dec 2023 · 0 repositories · arXiv:2312.10961
-
Multilingual large language models leak human stereotypes across language boundaries 12 Dec 2023 · 1 repository · arXiv:2312.07141
-
Translating Natural Language Queries to SQL Using the T5 Model 12 Dec 2023 · 0 repositories · arXiv:2312.12414
-
Aikyam: A Video Conferencing Utility for Deaf and Dumb 10 Dec 2023 · 0 repositories · arXiv:2312.05962
-
Domain Adaptation of a State of the Art Text-to-SQL Model: Lessons Learned and Challenges Found 9 Dec 2023 · 0 repositories · arXiv:2312.05448
-
PerfRL: A Small Language Model Framework for Efficient Code Optimization 9 Dec 2023 · 0 repositories · arXiv:2312.05657
-
Converting Epics/Stories into Pseudocode using Transformers 8 Dec 2023 · 0 repositories · arXiv:2312.05047
-
Making Translators Privacy-aware on the User's Side 7 Dec 2023 · 0 repositories · arXiv:2312.04068
-
A Text-to-Text Model for Multilingual Offensive Language Identification 6 Dec 2023 · 0 repositories · arXiv:2312.03379
-
Compositional Generalization for Data-to-Text Generation 5 Dec 2023 · 0 repositories · arXiv:2312.02748
-
A Machine Learning Approach Towards SKILL Code Autocompletion 4 Dec 2023 · 0 repositories · arXiv:2312.01921
-
Harnessing the Power of Prompt-based Techniques for Generating School-Level Questions using Large Language Models 2 Dec 2023 · 1 repository · arXiv:2312.01032
-
Towards leveraging LLMs for Conditional QA 2 Dec 2023 · 0 repositories · arXiv:2312.01143
-
Data Generation for Post-OCR correction of Cyrillic handwriting 27 Nov 2023 · 2 repositories · arXiv:2311.15896
-
Image Super-Resolution with Text Prompt Diffusion 24 Nov 2023 · 1 repository · arXiv:2311.14282
-
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization 22 Nov 2023 · 1 repository · arXiv:2311.13171Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
A novel transformer-based approach for soil temperature prediction 20 Nov 2023 · 0 repositories · arXiv:2311.11626
-
Bit Cipher -- A Simple yet Powerful Word Representation System that Integrates Efficiently with Language Models 18 Nov 2023 · 0 repositories · arXiv:2311.11012
-
Vashantor: A Large-scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language 18 Nov 2023 · 1 repository · arXiv:2311.11142
-
DynaPipe: Optimizing Multi-task Training through Dynamic Pipelines 17 Nov 2023 · 2 repositories · arXiv:2311.10418
-
Memory Augmented Language Models through Mixture of Word Experts 15 Nov 2023 · 0 repositories · arXiv:2311.10768
-
Spot: A Natural Language Interface for Geospatial Searches in OSM 14 Nov 2023 · 1 repository · arXiv:2311.08093
-
UT5: Pretraining Non autoregressive T5 with unrolled denoising 14 Nov 2023 · 0 repositories · arXiv:2311.08552
-
Controllable Topic-Focused Abstractive Summarization 12 Nov 2023 · 0 repositories · arXiv:2311.06724
-
Argumentation Element Annotation Modeling using XLNet 10 Nov 2023 · 0 repositories · arXiv:2311.06239
-
Exploring Fine-tuning ChatGPT for News Recommendation 10 Nov 2023 · 0 repositories · arXiv:2311.05850
-
Deep Learning Brasil at ABSAPT 2022: Portuguese Transformer Ensemble Approaches 8 Nov 2023 · 1 repository · arXiv:2311.05051
-
Modelling Sentiment Analysis: LLMs and data augmentation techniques 7 Nov 2023 · 1 repository · arXiv:2311.04139
-
Adapting Pre-trained Generative Models for Extractive Question Answering 6 Nov 2023 · 0 repositories · arXiv:2311.02961
-
FaMeSumm: Investigating and Improving Faithfulness of Medical Summarization 3 Nov 2023 · 1 repository · arXiv:2311.02271Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Indo LEGO-ABSA: A Multitask Generative Aspect Based Sentiment Analysis for Indonesian Language 3 Nov 2023 · 1 repository · arXiv:2311.01757
-
Better Together: Enhancing Generative Knowledge Graph Completion with Language Models and Neighborhood Information 2 Nov 2023 · 1 repository · arXiv:2311.01326
-
Learning Defect Prediction from Unrealistic Data 2 Nov 2023 · 0 repositories · arXiv:2311.00931
-
MAAIG: Motion Analysis And Instruction Generation 2 Nov 2023 · 0 repositories · arXiv:2311.00980
-
Attention Alignment and Flexible Positional Embeddings Improve Transformer Length Extrapolation 1 Nov 2023 · 0 repositories · arXiv:2311.00684
-
Data Augmentation for Code Translation with Comparable Corpora and Multiple References 1 Nov 2023 · 1 repository · arXiv:2311.00317Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Generating Medical Prescriptions with Conditional Transformer 30 Oct 2023 · 1 repository · arXiv:2310.19727Syntology official (archive's flag): 10 ran · 10 ran (of which 4 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
An Ensemble Method Based on the Combination of Transformers with Convolutional Neural Networks to Detect Artificially Generated Text 26 Oct 2023 · 0 repositories · arXiv:2310.17312
-
LightLM: A Lightweight Deep and Narrow Language Model for Generative Recommendation 26 Oct 2023 · 1 repository · arXiv:2310.17488Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
RedCoast: A Lightweight Tool to Automate Distributed Training of LLMs on Any GPU/TPUs 25 Oct 2023 · 1 repository · arXiv:2310.16355Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples)
-
Efficient Data Learning for Open Information Extraction with Pre-trained Language Models 23 Oct 2023 · 0 repositories · arXiv:2310.15021
-
Once Upon a Time in Graph: Relative-Time Pretraining for Complex Temporal Reasoning 23 Oct 2023 · 1 repository · arXiv:2310.14709Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Improving Cross-Lingual Transfer through Subtree-Aware Word Reordering 20 Oct 2023 · 1 repository · arXiv:2310.13583
-
Exploring Graph Neural Networks for Indian Legal Judgment Prediction 19 Oct 2023 · 0 repositories · arXiv:2310.12800
-
Empirical study of pretrained multilingual language models for zero-shot cross-lingual knowledge transfer in generation 15 Oct 2023 · 0 repositories · arXiv:2310.09917
-
Retrieval-Generation Alignment for End-to-End Task-Oriented Dialogue System 13 Oct 2023 · 1 repository · arXiv:2310.08877
-
Answer Candidate Type Selection: Text-to-Text Language Model for Closed Book Question Answering Meets Knowledge Graphs 10 Oct 2023 · 0 repositories · arXiv:2310.07008
-
Sparse Fine-tuning for Inference Acceleration of Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06927Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations 9 Oct 2023 · 0 repositories · arXiv:2310.05421
-
Benchmarking Large Language Models with Augmented Instructions for Fine-grained Information Extraction 8 Oct 2023 · 0 repositories · arXiv:2310.05092
-
Effective Slogan Generation with Noise Perturbation 6 Oct 2023 · 1 repository · arXiv:2310.04472
-
Can Large Language Models be Good Path Planners? A Benchmark and Investigation on Spatial-temporal Reasoning 5 Oct 2023 · 1 repository · arXiv:2310.03249Syntology official (archive's flag): 28 ran · 28 ran (of which 0 constructed an object rather than computing a result; 28 with no instrument failure: 0 honoured, 0 violated, 28 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 28 harvested samples) · 28 pointer-only (licence)
-
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines 5 Oct 2023 · 3 repositories · arXiv:2310.03714Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Natural Language Models for Data Visualization Utilizing nvBench Dataset 2 Oct 2023 · 0 repositories · arXiv:2310.00832
-
Testing the Limits of Unified Sequence to Sequence LLM Pretraining on Diverse Table Data Tasks 1 Oct 2023 · 0 repositories · arXiv:2310.00789
-
DeBERTinha: A Multistep Approach to Adapt DebertaV3 XSmall for Brazilian Portuguese Natural Language Processing Task 28 Sep 2023 · 0 repositories · arXiv:2309.16844
-
CAPP-130: A Corpus of Chinese Application Privacy Policy Summarization and Interpretation 26 Sep 2023 · 1 repository
-
Program Repair with Minimal Edits Using CodeT5 26 Sep 2023 · 0 repositories · arXiv:2309.14760
-
Lexical Squad@Multimodal Hate Speech Event Detection 2023: Multimodal Hate Speech Detection using Fused Ensemble Approach 23 Sep 2023 · 1 repository · arXiv:2309.13354
-
On the Relationship between Skill Neurons and Robustness in Prompt Tuning 21 Sep 2023 · 1 repository · arXiv:2309.12263
-
Localize, Retrieve and Fuse: A Generalized Framework for Free-Form Question Answering over Tables 20 Sep 2023 · 0 repositories · arXiv:2309.11049
-
Sequence-to-Sequence Spanish Pre-trained Language Models 20 Sep 2023 · 1 repository · arXiv:2309.11259
-
Empowering In-Browser Deep Learning Inference on Edge Devices with Just-in-Time Kernel Optimizations 16 Sep 2023 · 0 repositories · arXiv:2309.08978
-
Structural Self-Supervised Objectives for Transformers 15 Sep 2023 · 1 repository · arXiv:2309.08272
-
DBLPLink: An Entity Linker for the DBLP Scholarly Knowledge Graph 14 Sep 2023 · 1 repository · arXiv:2309.07545
-
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish 13 Sep 2023 · 1 repository · arXiv:2309.06698
-
RAP-Gen: Retrieval-Augmented Patch Generation with CodeT5 for Automatic Program Repair 12 Sep 2023 · 0 repositories · arXiv:2309.06057
-
Detecting Natural Language Biases with Prompt-based Learning 11 Sep 2023 · 0 repositories · arXiv:2309.05227
-
Can NLP Models 'Identify', 'Distinguish', and 'Justify' Questions that Don't have a Definitive Answer? 8 Sep 2023 · 0 repositories · arXiv:2309.04635
-
nanoT5: A PyTorch Framework for Pre-training and Fine-tuning T5-style Models with Limited Resources 5 Sep 2023 · 1 repository · arXiv:2309.02373Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples)
-
SP³: Enhancing Structured Pruning via PCA Projection 31 Aug 2023 · 1 repository · arXiv:2308.16475Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 1 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Multi-party Goal Tracking with LLMs: Comparing Pre-training, Fine-tuning, and Prompt Engineering 29 Aug 2023 · 1 repository · arXiv:2308.15231