Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 20
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 20 of 71: papers 1,901 to 2,000 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Shallow Cross-Encoders for Low-Latency Retrieval 29 Mar 2024 · 1 repository · arXiv:2403.20222
-
A Review of Multi-Modal Large Language and Vision Models 28 Mar 2024 · 0 repositories · arXiv:2404.01322
-
AlloyBERT: Alloy Property Prediction with Large Language Models 28 Mar 2024 · 0 repositories · arXiv:2403.19783
-
Are Large Language Models Good at Utility Judgments? 28 Mar 2024 · 1 repository · arXiv:2403.19216Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
FACTOID: FACtual enTailment fOr hallucInation Detection 28 Mar 2024 · 0 repositories · arXiv:2403.19113
-
Intelligent Classification and Personalized Recommendation of E-commerce Products Based on Machine Learning 28 Mar 2024 · 0 repositories · arXiv:2403.19345
-
Risk prediction of pathological gambling on social media 28 Mar 2024 · 0 repositories · arXiv:2403.19358
-
A Novel Corpus of Annotated Medical Imaging Reports and Information Extraction Results Using BERT-based Language Models 27 Mar 2024 · 1 repository · arXiv:2403.18975
-
AcTED: Automatic Acquisition of Typical Event Duration for Semi-supervised Temporal Commonsense QA 27 Mar 2024 · 0 repositories · arXiv:2403.18504
-
Boosting Conversational Question Answering with Fine-Grained Retrieval-Augmentation and Self-Check 27 Mar 2024 · 0 repositories · arXiv:2403.18243
-
CPR: Retrieval Augmented Generation for Copyright Protection 27 Mar 2024 · 0 repositories · arXiv:2403.18920
-
Evaluating Large Language Models for Health-Related Text Classification Tasks with Public Social Media Data 27 Mar 2024 · 0 repositories · arXiv:2403.19031
-
Fusion approaches for emotion recognition from speech using acoustic and text-based features 27 Mar 2024 · 0 repositories · arXiv:2403.18635
-
mALBERT: Is a Compact Multilingual BERT Model Still Worth It? 27 Mar 2024 · 0 repositories · arXiv:2403.18338
-
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks 27 Mar 2024 · 1 repository · arXiv:2403.18423Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Are Compressed Language Models Less Subgroup Robust? 26 Mar 2024 · 1 repository · arXiv:2403.17811
-
Fingerprinting web servers through Transformer-encoded HTTP response headers 26 Mar 2024 · 1 repository · arXiv:2404.00056
-
Hierarchical Multi-label Classification for Fine-level Event Extraction from Aviation Accident Reports 26 Mar 2024 · 0 repositories · arXiv:2403.17914
-
Targeted Visualization of the Backbone of Encoder LLMs 26 Mar 2024 · 1 repository · arXiv:2403.18872
-
A Study on How Attention Scores in the BERT Model are Aware of Lexical Categories in Syntactic and Semantic Tasks on the GLUE Benchmark 25 Mar 2024 · 0 repositories · arXiv:2403.16447
-
LexDrafter: Terminology Drafting for Legislative Documents using Retrieval Augmented Generation 24 Mar 2024 · 1 repository · arXiv:2403.16295
-
Fine Tuning LLM for Enterprise: Practical Guidelines and Recommendations 23 Mar 2024 · 0 repositories · arXiv:2404.10779
-
Improving Retrieval for RAG based Question Answering Models on Financial Documents 23 Mar 2024 · 0 repositories · arXiv:2404.07221
-
LlamBERT: Large-scale low-cost data annotation in NLP 23 Mar 2024 · 1 repository · arXiv:2403.15938
-
Towards a RAG-based Summarization Agent for the Electron-Ion Collider 23 Mar 2024 · 1 repository · arXiv:2403.15729
-
Blended RAG: Improving RAG (Retriever-Augmented Generation) Accuracy with Semantic Search and Hybrid Query-Based Retrievers 22 Mar 2024 · 1 repository · arXiv:2404.07220Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MasonTigers at SemEval-2024 Task 1: An Ensemble Approach for Semantic Textual Relatedness 22 Mar 2024 · 0 repositories · arXiv:2403.14990
-
Selecting Query-bag as Pseudo Relevance Feedback for Information-seeking Conversations 22 Mar 2024 · 0 repositories · arXiv:2404.04272
-
Text Clustering with Large Language Model Embeddings 22 Mar 2024 · 0 repositories · arXiv:2403.15112
-
FIT-RAG: Black-Box RAG with Factual Information and Token Reduction 21 Mar 2024 · 0 repositories · arXiv:2403.14374
-
LLM-based Extraction of Contradictions from Patents 21 Mar 2024 · 0 repositories · arXiv:2403.14258
-
Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection 20 Mar 2024 · 1 repository · arXiv:2403.14037
-
Efficient argument classification with compact language models and ChatGPT-4 refinements 20 Mar 2024 · 0 repositories · arXiv:2403.15473
-
Fine-Tuning Pre-trained Language Models to Detect In-Game Trash Talks 19 Mar 2024 · 0 repositories · arXiv:2403.15458
-
Pipelined Biomedical Event Extraction Rivaling Joint Learning 19 Mar 2024 · 0 repositories · arXiv:2403.12386
-
TT-BLIP: Enhancing Fake News Detection Using BLIP and Tri-Transformer 19 Mar 2024 · 0 repositories · arXiv:2403.12481
-
A Disease Labeler for Chinese Chest X-Ray Report Generation 18 Mar 2024 · 0 repositories · arXiv:2404.16852
-
CICLe: Conformal In-Context Learning for Largescale Multi-Class Food Risk Classification 18 Mar 2024 · 1 repository · arXiv:2403.11904
-
Narrative Feature or Structured Feature? A Study of Large Language Models to Identify Cancer Patients at Risk of Heart Failure 18 Mar 2024 · 1 repository · arXiv:2403.11425
-
JORA: JAX Tensor-Parallel LoRA Library for Retrieval Augmented Fine-Tuning 17 Mar 2024 · 1 repository · arXiv:2403.11366
-
DRAGIN: Dynamic Retrieval Augmented Generation based on the Information Needs of Large Language Models 15 Mar 2024 · 1 repository · arXiv:2403.10081Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases 15 Mar 2024 · 1 repository · arXiv:2403.10446Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 13 harvested samples)
-
RAFT: Adapting Language Model to Domain Specific RAG 15 Mar 2024 · 1 repository · arXiv:2403.10131
-
Repoformer: Selective Retrieval for Repository-Level Code Completion 15 Mar 2024 · 0 repositories · arXiv:2403.10059
-
Fisher Mask Nodes for Language Model Merging 14 Mar 2024 · 1 repository · arXiv:2403.09891
-
Incorporating Graph Attention Mechanism into Geometric Problem Solving Based on Deep Reinforcement Learning 14 Mar 2024 · 1 repository · arXiv:2403.14690
-
RAGGED: Towards Informed Design of Retrieval Augmented Generation Systems 14 Mar 2024 · 1 repository · arXiv:2403.09040Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Retrieval augmented text-to-SQL generation for epidemiological question answering using electronic health records 14 Mar 2024 · 1 repository · arXiv:2403.09226
-
Autoregressive Score Generation for Multi-trait Essay Scoring 13 Mar 2024 · 1 repository · arXiv:2403.08332
-
Distilling Named Entity Recognition Models for Endangered Species from Large Language Models 13 Mar 2024 · 0 repositories · arXiv:2403.15430
-
Embedded Translations for Low-resource Automated Glossing 13 Mar 2024 · 0 repositories · arXiv:2403.08189
-
Research on the Application of Deep Learning-based BERT Model in Sentiment Analysis 13 Mar 2024 · 0 repositories · arXiv:2403.08217
-
Rich Semantic Knowledge Enhanced Large Language Models for Few-shot Chinese Spell Checking 13 Mar 2024 · 0 repositories · arXiv:2403.08492
-
Enhancing Readmission Prediction with Deep Learning: Extracting Biomedical Concepts from Clinical Texts 12 Mar 2024 · 0 repositories · arXiv:2403.09722
-
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems 12 Mar 2024 · 1 repository · arXiv:2403.09727
-
LookupFFN: Making Transformers Compute-lite for CPU inference 12 Mar 2024 · 1 repository · arXiv:2403.07221Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions 12 Mar 2024 · 1 repository · arXiv:2403.07678
-
A multi-cohort study on prediction of acute brain dysfunction states using selective state space models 11 Mar 2024 · 0 repositories · arXiv:2403.07201
-
Development of a Reliable and Accessible Caregiving Language Model (CaLM) 11 Mar 2024 · 0 repositories · arXiv:2403.06857
-
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data 9 Mar 2024 · 0 repositories · arXiv:2403.05820
-
Enhancing Multi-Hop Knowledge Graph Reasoning through Reward Shaping Techniques 9 Mar 2024 · 0 repositories · arXiv:2403.05801
-
TokenMark: A Modality-Agnostic Watermark for Pre-trained Transformers 9 Mar 2024 · 0 repositories · arXiv:2403.05842
-
PipeRAG: Fast Retrieval-Augmented Generation via Algorithm-System Co-design 8 Mar 2024 · 0 repositories · arXiv:2403.05676
-
The Impact of Quantization on the Robustness of Transformer-based Text Classifiers 8 Mar 2024 · 0 repositories · arXiv:2403.05365
-
Automating the Information Extraction from Semi-Structured Interview Transcripts 7 Mar 2024 · 1 repository · arXiv:2403.04819
-
Federated Recommendation via Hybrid Retrieval Augmented Generation 7 Mar 2024 · 1 repository · arXiv:2403.04256
-
Enhancing ASD detection accuracy: a combined approach of machine learning and deep learning models with natural language processing 6 Mar 2024 · 1 repository · arXiv:2403.03581
-
FaaF: Facts as a Function for the evaluation of generated text 6 Mar 2024 · 1 repository · arXiv:2403.03888
-
GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection 6 Mar 2024 · 3 repositories · arXiv:2403.03507
-
Japanese-English Sentence Translation Exercises Dataset for Automatic Grading 6 Mar 2024 · 0 repositories · arXiv:2403.03396
-
Exploring Naive Approaches to Tell Apart LLMs Productions from Human-written Text 5 Mar 2024 · 1 repository
-
EEE-QA: Exploring Effective and Efficient Question-Answer Representations 4 Mar 2024 · 1 repository · arXiv:2403.02176
-
NoteLLM: A Retrievable Large Language Model for Note Recommendation 4 Mar 2024 · 0 repositories · arXiv:2403.01744
-
Vanilla Transformers are Transfer Capability Teachers 4 Mar 2024 · 0 repositories · arXiv:2403.01994
-
Fine Tuning vs. Retrieval Augmented Generation for Less Popular Knowledge 3 Mar 2024 · 1 repository · arXiv:2403.01432Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
Multi-level Product Category Prediction through Text Classification 3 Mar 2024 · 1 repository · arXiv:2403.01638
-
Analysis of Privacy Leakage in Federated Large Language Models 2 Mar 2024 · 1 repository · arXiv:2403.04784Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Greed is All You Need: An Evaluation of Tokenizer Inference Methods 2 Mar 2024 · 1 repository · arXiv:2403.01289Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
RAGged Edges: The Double-Edged Sword of Retrieval-Augmented Chatbots 2 Mar 2024 · 0 repositories · arXiv:2403.01193
-
ATP: Enabling Fast LLM Serving via Attention on Top Principal Keys 1 Mar 2024 · 0 repositories · arXiv:2403.02352
-
Large Language Models for Simultaneous Named Entity Extraction and Spelling Correction 1 Mar 2024 · 0 repositories · arXiv:2403.00528
-
Crafting Knowledge: Exploring the Creative Mechanisms of Chat-Based Search Engines 29 Feb 2024 · 0 repositories · arXiv:2402.19421
-
Leveraging pre-trained language models for code generation 29 Feb 2024 · 1 repository
-
PaECTER: Patent-level Representation Learning using Citation-informed Transformers 29 Feb 2024 · 0 repositories · arXiv:2402.19411
-
PeLLE: Encoder-based language models for Brazilian Portuguese based on open data 29 Feb 2024 · 0 repositories · arXiv:2402.19204
-
Retrieval-Augmented Generation for AI-Generated Content: A Survey 29 Feb 2024 · 3 repositories · arXiv:2402.19473
-
Few-Shot Fairness: Unveiling LLM's Potential for Fairness-Aware Classification 28 Feb 2024 · 0 repositories · arXiv:2402.18502
-
Orchid: Flexible and Data-Dependent Convolution for Sequence Modeling 28 Feb 2024 · 0 repositories · arXiv:2402.18508
-
WIKIGENBENCH: Exploring Full-length Wikipedia Generation under Real-World Scenario 28 Feb 2024 · 1 repository · arXiv:2402.18264
-
Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation 28 Feb 2024 · 1 repository · arXiv:2402.18150Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
A Language Model based Framework for New Concept Placement in Ontologies 27 Feb 2024 · 1 repository · arXiv:2402.17897
-
Deep Learning Detection Method for Large Language Models-Generated Scientific Content 27 Feb 2024 · 0 repositories · arXiv:2403.00828
-
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages 27 Feb 2024 · 0 repositories · arXiv:2402.17496
-
Evaluating Very Long-Term Conversational Memory of LLM Agents 27 Feb 2024 · 1 repository · arXiv:2402.17753Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 7 harvested samples) · 7 pointer-only (licence)
-
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems 27 Feb 2024 · 1 repository · arXiv:2402.17840Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability 27 Feb 2024 · 1 repository · arXiv:2402.17887
-
MAGPIE: Multi-Task Media-Bias Analysis Generalization for Pre-Trained Identification of Expressions 27 Feb 2024 · 1 repository · arXiv:2403.07910
-
REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question Answering 27 Feb 2024 · 1 repository · arXiv:2402.17497Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
Adaptation of Biomedical and Clinical Pretrained Models to French Long Documents: A Comparative Study 26 Feb 2024 · 1 repository · arXiv:2402.16689
-
Asymmetry in Low-Rank Adapters of Foundation Models 26 Feb 2024 · 1 repository · arXiv:2402.16842Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)