Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 21
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 21 of 71: papers 2,001 to 2,100 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
From RAGs to riches: Using large language models to write documents for clinical trials 26 Feb 2024 · 0 repositories · arXiv:2402.16406
-
Retrieval Augmented Generation Systems: Automatic Dataset Creation, Evaluation and Boolean Agent Setup 26 Feb 2024 · 1 repository · arXiv:2403.00820
-
Deep Learning Approaches for Improving Question Answering Systems in Hepatocellular Carcinoma Research 25 Feb 2024 · 0 repositories · arXiv:2402.16038
-
Emotion Classification in Short English Texts using Deep Learning Techniques 25 Feb 2024 · 0 repositories · arXiv:2402.16034
-
Hitting "Probe"rty with Non-Linearity, and More 25 Feb 2024 · 0 repositories · arXiv:2402.16168
-
MultiContrievers: Analysis of Dense Retrieval Representations 24 Feb 2024 · 1 repository · arXiv:2402.15925
-
SemEval-2024 Task 8: Weighted Layer Averaging RoBERTa for Black-Box Machine-Generated Text Detection 24 Feb 2024 · 1 repository · arXiv:2402.15873
-
Advancing Parameter Efficiency in Fine-tuning via Representation Editing 23 Feb 2024 · 2 repositories · arXiv:2402.15179Syntology official (archive's flag): 1 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 1 pointer-only (licence)
-
Dual Encoder: Exploiting the Potential of Syntactic and Semantic for Aspect Sentiment Triplet Extraction 23 Feb 2024 · 0 repositories · arXiv:2402.15370
-
Evaluating the Performance of ChatGPT for Spam Email Detection 23 Feb 2024 · 0 repositories · arXiv:2402.15537
-
The Good and The Bad: Exploring Privacy Issues in Retrieval-Augmented Generation (RAG) 23 Feb 2024 · 1 repository · arXiv:2402.16893Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards Efficient Active Learning in NLP via Pretrained Representations 23 Feb 2024 · 0 repositories · arXiv:2402.15613
-
2D Matryoshka Sentence Embeddings 22 Feb 2024 · 1 repository · arXiv:2402.14776
-
Assessing generalization capability of text ranking models in Polish 22 Feb 2024 · 0 repositories · arXiv:2402.14318
-
How Important Is Tokenization in French Medical Masked Language Models? 22 Feb 2024 · 0 repositories · arXiv:2402.15010
-
Transferring BERT Capabilities from High-Resource to Low-Resource Languages Using Vocabulary Matching 22 Feb 2024 · 0 repositories · arXiv:2402.14408
-
ActiveRAG: Autonomously Knowledge Assimilation and Accommodation through Retrieval-Augmented Agents 21 Feb 2024 · 1 repository · arXiv:2402.13547
-
An Explainable Transformer-based Model for Phishing Email Detection: A Large Language Model Approach 21 Feb 2024 · 0 repositories · arXiv:2402.13871
-
Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures 21 Feb 2024 · 0 repositories · arXiv:2402.13640
-
Improving Language Understanding from Screenshots 21 Feb 2024 · 1 repository · arXiv:2402.14073Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Are ELECTRA's Sentence Embeddings Beyond Repair? The Case of Semantic Textual Similarity 20 Feb 2024 · 1 repository · arXiv:2402.13130
-
Benchmarking Retrieval-Augmented Generation for Medicine 20 Feb 2024 · 2 repositories · arXiv:2402.13178Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Can GNN be Good Adapter for LLMs? 20 Feb 2024 · 2 repositories · arXiv:2402.12984Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Exploring the Impact of Table-to-Text Methods on Augmenting LLM-based Question Answering with Domain Hybrid Data 20 Feb 2024 · 0 repositories · arXiv:2402.12869
-
Acquiring Clean Language Models from Backdoor Poisoned Datasets by Downscaling Frequency Space 19 Feb 2024 · 1 repository · arXiv:2402.12026Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CodeArt: Better Code Models by Attention Regularization When Symbols Are Lacking 19 Feb 2024 · 1 repository · arXiv:2402.11842
-
FeB4RAG: Evaluating Federated Search in the Context of Retrieval Augmented Generation 19 Feb 2024 · 0 repositories · arXiv:2402.11891
-
Graph-Based Retriever Captures the Long Tail of Biomedical Knowledge 19 Feb 2024 · 0 repositories · arXiv:2402.12352
-
Head-wise Shareable Attention for Large Language Models 19 Feb 2024 · 2 repositories · arXiv:2402.11819
-
KARL: Knowledge-Aware Retrieval and Representations aid Retention and Learning in Students 19 Feb 2024 · 0 repositories · arXiv:2402.12291
-
Language Model Adaptation to Specialized Domains through Selective Masking based on Genre and Topical Characteristics 19 Feb 2024 · 1 repository · arXiv:2402.12036
-
Mafin: Enhancing Black-Box Embeddings with Model Augmented Fine-Tuning 19 Feb 2024 · 0 repositories · arXiv:2402.12177
-
Ontology Enhanced Claim Detection 19 Feb 2024 · 0 repositories · arXiv:2402.12282
-
What Evidence Do Language Models Find Convincing? 19 Feb 2024 · 1 repository · arXiv:2402.11782Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 14 harvested samples)
-
A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models 18 Feb 2024 · 1 repository · arXiv:2402.11469
-
Metric-Learning Encoding Models Identify Processing Profiles of Linguistic Features in BERT's Representations 18 Feb 2024 · 1 repository · arXiv:2402.11608
-
Utilizing BERT for Information Retrieval: Survey, Applications, Resources, and Challenges 18 Feb 2024 · 0 repositories · arXiv:2403.00784
-
Detecting a Proxy for Potential Comorbid ADHD in People Reporting Anxiety Symptoms from Social Media Data 17 Feb 2024 · 0 repositories · arXiv:2403.05561
-
GenDec: A robust generative Question-decomposition method for Multi-hop reasoning 17 Feb 2024 · 0 repositories · arXiv:2402.11166
-
Emoji Driven Crypto Assets Market Reactions 16 Feb 2024 · 0 repositories · arXiv:2402.10481
-
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss 16 Feb 2024 · 2 repositories · arXiv:2402.10790
-
LogELECTRA: Self-supervised Anomaly Detection for Unstructured Logs 16 Feb 2024 · 0 repositories · arXiv:2402.10397
-
Where is the answer? Investigating Positional Bias in Language Model Knowledge Extraction 16 Feb 2024 · 1 repository · arXiv:2402.12170
-
COVIDHealth: A Benchmark Twitter Dataset and Machine Learning based Web Application for Classifying COVID-19 Discussions 15 Feb 2024 · 1 repository · arXiv:2402.09897
-
Grounding Language Model with Chunking-Free In-Context Retrieval 15 Feb 2024 · 0 repositories · arXiv:2402.09760
-
A Language Model for Particle Tracking 14 Feb 2024 · 0 repositories · arXiv:2402.10239
-
Leveraging Large Language Models for Enhanced NLP Task Performance through Knowledge Distillation and Optimized Training Strategies 14 Feb 2024 · 0 repositories · arXiv:2402.09282
-
Regional inflation analysis using social network data 14 Feb 2024 · 0 repositories · arXiv:2403.00774
-
ScamSpot: Fighting Financial Fraud in Instagram Comments 14 Feb 2024 · 0 repositories · arXiv:2402.08869
-
BERT4FCA: A Method for Bipartite Link Prediction using Formal Concept Analysis and BERT 13 Feb 2024 · 0 repositories · arXiv:2402.08236
-
Improving Black-box Robustness with In-Context Rewriting 13 Feb 2024 · 1 repository · arXiv:2402.08225
-
Developing a Multi-variate Prediction Model For COVID-19 From Crowd-sourced Respiratory Voice Data 12 Feb 2024 · 0 repositories · arXiv:2402.07619
-
G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering 12 Feb 2024 · 2 repositories · arXiv:2402.07630Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Leveraging AI to Advance Science and Computing Education across Africa: Challenges, Progress and Opportunities 12 Feb 2024 · 0 repositories · arXiv:2402.07397
-
PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models 12 Feb 2024 · 2 repositories · arXiv:2402.07867Syntology official (archive's flag): 10 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 17 harvested samples) · 7 pointer-only (licence)
-
T-RAG: Lessons from the LLM Trenches 12 Feb 2024 · 0 repositories · arXiv:2402.07483
-
HyperBERT: Mixing Hypergraph-Aware Layers with Language Models for Node Classification on Text-Attributed Hypergraphs 11 Feb 2024 · 1 repository · arXiv:2402.07309
-
Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models 11 Feb 2024 · 0 repositories · arXiv:2402.07179
-
Understanding the Training Speedup from Sampling with Approximate Losses 10 Feb 2024 · 0 repositories · arXiv:2402.07052
-
FaBERT: Pre-training BERT on Persian Blogs 9 Feb 2024 · 0 repositories · arXiv:2402.06617
-
G-SciEdBERT: A Contextualized LLM for Science Assessment Tasks in German 9 Feb 2024 · 1 repository · arXiv:2402.06584
-
Efficient Models for the Detection of Hate, Abuse and Profanity 8 Feb 2024 · 0 repositories · arXiv:2402.05624
-
Efficient Stagewise Pretraining via Progressive Subnetworks 8 Feb 2024 · 0 repositories · arXiv:2402.05913
-
Neural Models for Source Code Synthesis and Completion 8 Feb 2024 · 0 repositories · arXiv:2402.06690
-
Aspect-Based Sentiment Analysis for Open-Ended HR Survey Responses 7 Feb 2024 · 0 repositories · arXiv:2402.04812
-
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models 7 Feb 2024 · 0 repositories · arXiv:2402.05034
-
Enhancing Retrieval Processes for Language Generation with Augmented Queries 6 Feb 2024 · 0 repositories · arXiv:2402.16874
-
LegalLens: Leveraging LLMs for Legal Violation Identification in Unstructured Text 6 Feb 2024 · 1 repository · arXiv:2402.04335
-
Stanceosaurus 2.0: Classifying Stance Towards Russian and Spanish Misinformation 6 Feb 2024 · 0 repositories · arXiv:2402.03642
-
The Use of a Large Language Model for Cyberbullying Detection 6 Feb 2024 · 0 repositories · arXiv:2402.04088
-
Accurate and Well-Calibrated ICD Code Assignment Through Attention Over Diverse Label Embeddings 5 Feb 2024 · 1 repository · arXiv:2402.03172Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Arabic Synonym BERT-based Adversarial Examples for Text Classification 5 Feb 2024 · 1 repository · arXiv:2402.03477
-
C-RAG: Certified Generation Risks for Retrieval-Augmented Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03181Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Enhancing textual textbook question answering with large language models and retrieval augmented generation 5 Feb 2024 · 1 repository · arXiv:2402.05128Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Financial Report Chunking for Effective Retrieval Augmented Generation 5 Feb 2024 · 1 repository · arXiv:2402.05131
-
LB-KBQA: Large-language-model and BERT based Knowledge-Based Question and Answering System 5 Feb 2024 · 0 repositories · arXiv:2402.05130
-
Multi-Lingual Malaysian Embedding: Leveraging Large Language Models for Semantic Representations 5 Feb 2024 · 0 repositories · arXiv:2402.03053
-
UniMem: Towards a Unified View of Long-Context Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.03009
-
Breaking MLPerf Training: A Case Study on Optimizing BERT 4 Feb 2024 · 0 repositories · arXiv:2402.02447
-
Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation 4 Feb 2024 · 0 repositories · arXiv:2402.14594
-
Data Quality Matters: Suicide Intention Detection on Social Media Posts Using RoBERTa-CNN 3 Feb 2024 · 0 repositories · arXiv:2402.02262
-
DE³-BERT: Distance-Enhanced Early Exiting for BERT based on Prototypical Networks 3 Feb 2024 · 0 repositories · arXiv:2402.05948
-
Clarifying the Path to User Satisfaction: An Investigation into Clarification Usefulness 2 Feb 2024 · 1 repository · arXiv:2402.01934
-
LLM-Detector: Improving AI-Generated Chinese Text Detection with Open-Source LLM Instruction Tuning 2 Feb 2024 · 1 repository · arXiv:2402.01158
-
Predicting ATP binding sites in protein sequences using Deep Learning and Natural Language Processing 2 Feb 2024 · 0 repositories · arXiv:2402.01829
-
Retrieval Augmented End-to-End Spoken Dialog Models 2 Feb 2024 · 0 repositories · arXiv:2402.01828
-
CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks 2 Feb 2024 · 0 repositories · arXiv:2402.01176
-
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA 1 Feb 2024 · 0 repositories · arXiv:2402.01767
-
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models 1 Feb 2024 · 1 repository · arXiv:2402.00794
-
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes 1 Feb 2024 · 0 repositories · arXiv:2402.00987
-
RAG-Fusion: a New Take on Retrieval-Augmented Generation 31 Jan 2024 · 0 repositories · arXiv:2402.03367
-
Arabic Tweet Act: A Weighted Ensemble Pre-Trained Transformer Model for Classifying Arabic Speech Acts on Twitter 30 Jan 2024 · 0 repositories · arXiv:2401.17373
-
Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs 30 Jan 2024 · 1 repository · arXiv:2401.16638
-
CRUD-RAG: A Comprehensive Chinese Benchmark for Retrieval-Augmented Generation of Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.17043Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Detecting mental disorder on social media: a ChatGPT-augmented explainable approach 30 Jan 2024 · 1 repository · arXiv:2401.17477
-
Detecting Racist Text in Bengali: An Ensemble Deep Learning Framework 30 Jan 2024 · 0 repositories · arXiv:2401.16748
-
Fine-tuning Transformer-based Encoder for Turkish Language Understanding Tasks 30 Jan 2024 · 0 repositories · arXiv:2401.17396
-
Large Multi-Modal Models (LMMs) as Universal Foundation Models for AI-Native Wireless Systems 30 Jan 2024 · 0 repositories · arXiv:2402.01748
-
Single Word Change is All You Need: Designing Attacks and Defenses for Text Classifiers 30 Jan 2024 · 0 repositories · arXiv:2401.17196
-
Towards Generating Informative Textual Description for Neurons in Language Models 30 Jan 2024 · 0 repositories · arXiv:2401.16731