Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 10
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 10 of 71: papers 901 to 1,000 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Scaling BERT Models for Turkish Automatic Punctuation and Capitalization Correction 3 Dec 2024 · 0 repositories · arXiv:2412.02698
-
Semantic Tokens in Retrieval Augmented Generation 3 Dec 2024 · 0 repositories · arXiv:2412.02563
-
MBA-RAG: a Bandit Approach for Adaptive Retrieval-Augmented Generation through Question Complexity 2 Dec 2024 · 1 repository · arXiv:2412.01572Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Su-RoBERTa: A Semi-supervised Approach to Predicting Suicide Risk through Social Media using Base Language Models 2 Dec 2024 · 0 repositories · arXiv:2412.01353
-
A Comprehensive Guide to Explainable AI: From Classical Models to LLMs 1 Dec 2024 · 1 repository · arXiv:2412.00800
-
Lightweight Contenders: Navigating Semi-Supervised Text Mining through Peer Collaboration and Self Transcendence 1 Dec 2024 · 1 repository · arXiv:2412.00883
-
Does Self-Attention Need Separate Weights in Transformers? 30 Nov 2024 · 0 repositories · arXiv:2412.00359
-
Fairness at Every Intersection: Uncovering and Mitigating Intersectional Biases in Multimodal Clinical Predictions 30 Nov 2024 · 0 repositories · arXiv:2412.00606
-
Advanced System Integration: Analyzing OpenAPI Chunking for Retrieval-Augmented Generation 29 Nov 2024 · 0 repositories · arXiv:2411.19804
-
Generating a Low-code Complete Workflow via Task Decomposition and RAG 29 Nov 2024 · 0 repositories · arXiv:2412.00239
-
Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems 29 Nov 2024 · 0 repositories · arXiv:2411.19710
-
Knowledge Management for Automobile Failure Analysis Using Graph RAG 29 Nov 2024 · 0 repositories · arXiv:2411.19539
-
RAGDiffusion: Faithful Cloth Generation via External Knowledge Assimilation 29 Nov 2024 · 0 repositories · arXiv:2411.19528
-
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation 29 Nov 2024 · 0 repositories · arXiv:2411.19921
-
Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems 29 Nov 2024 · 0 repositories · arXiv:2411.19463
-
Efficient Learning Content Retrieval with Knowledge Injection 28 Nov 2024 · 0 repositories · arXiv:2412.00125
-
RevPRAG: Revealing Poisoning Attacks in Retrieval-Augmented Generation through LLM Activation Analysis 28 Nov 2024 · 0 repositories · arXiv:2411.18948
-
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation 28 Nov 2024 · 1 repository · arXiv:2411.19067
-
Can bidirectional encoder become the ultimate winner for downstream applications of foundation models? 27 Nov 2024 · 0 repositories · arXiv:2411.18021
-
Evaluating and Improving the Robustness of Security Attack Detectors Generated by LLMs 27 Nov 2024 · 1 repository · arXiv:2411.18216
-
Fine-Tuning Large Language Models for Scientific Text Classification: A Comparative Study 27 Nov 2024 · 0 repositories · arXiv:2412.00098
-
Fine-Tuning Small Embeddings for Elevated Performance 27 Nov 2024 · 0 repositories · arXiv:2411.18099
-
On Importance of Code-Mixed Embeddings for Hate Speech Identification 27 Nov 2024 · 0 repositories · arXiv:2411.18577
-
BERT or FastText? A Comparative Analysis of Contextual as well as Non-Contextual Embeddings 26 Nov 2024 · 1 repository · arXiv:2411.17661
-
Fairness And Performance In Harmony: Data Debiasing Is All You Need 26 Nov 2024 · 0 repositories · arXiv:2411.17374
-
Linguistic Laws Meet Protein Sequences: A Comparative Analysis of Subword Tokenization Methods 26 Nov 2024 · 1 repository · arXiv:2411.17669
-
What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics 26 Nov 2024 · 0 repositories · arXiv:2411.17593
-
AtomR: Atomic Operator-Empowered Large Language Models for Heterogeneous Knowledge Reasoning 25 Nov 2024 · 1 repository · arXiv:2411.16495Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models 25 Nov 2024 · 0 repositories · arXiv:2411.16991
-
Human-Calibrated Automated Testing and Validation of Generative Language Models 25 Nov 2024 · 0 repositories · arXiv:2411.16391
-
LaB-RAG: Label Boosted Retrieval Augmented Generation for Radiology Report Generation 25 Nov 2024 · 1 repository · arXiv:2411.16523
-
Predictive Power of LLMs in Financial Markets 25 Nov 2024 · 0 repositories · arXiv:2411.16569
-
StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training 25 Nov 2024 · 0 repositories · arXiv:2411.16618
-
Development of Pre-Trained Transformer-based Models for the Nepali Language 24 Nov 2024 · 0 repositories · arXiv:2411.15734
-
RAMIE: Retrieval-Augmented Multi-task Information Extraction with Large Language Models on Dietary Supplements 24 Nov 2024 · 0 repositories · arXiv:2411.15700
-
A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit 23 Nov 2024 · 1 repository · arXiv:2411.15404
-
Improving Next Tokens via Second-Last Predictions with Generate and Refine 23 Nov 2024 · 0 repositories · arXiv:2411.15661
-
Inducing Human-like Biases in Moral Reasoning Language Models 23 Nov 2024 · 0 repositories · arXiv:2411.15386
-
Traditional Chinese Medicine Case Analysis System for High-Level Semantic Abstraction: Optimized with Prompt and RAG 23 Nov 2024 · 0 repositories · arXiv:2411.15491
-
Astro-HEP-BERT: A bidirectional language model for studying the meanings of concepts in astrophysics and high energy physics 22 Nov 2024 · 0 repositories · arXiv:2411.14877
-
Comparative Analysis of Pooling Mechanisms in LLMs: A Sentiment Analysis Perspective 22 Nov 2024 · 0 repositories · arXiv:2411.14654
-
KBAlign: Efficient Self Adaptation on Specific Knowledge Bases 22 Nov 2024 · 1 repository · arXiv:2411.14790
-
An Experimental Study on Data Augmentation Techniques for Named Entity Recognition on Low-Resource Domains 21 Nov 2024 · 0 repositories · arXiv:2411.14551
-
BERT-Based Approach for Automating Course Articulation Matrix Construction with Explainable AI 21 Nov 2024 · 1 repository · arXiv:2411.14254
-
FastRAG: Retrieval Augmented Generation for Semi-structured Data 21 Nov 2024 · 0 repositories · arXiv:2411.13773
-
G-RAG: Knowledge Expansion in Material Science 21 Nov 2024 · 1 repository · arXiv:2411.14592
-
POS-tagging to highlight the skeletal structure of sentences 21 Nov 2024 · 2 repositories · arXiv:2411.14393
-
Towards Knowledge Checking in Retrieval-augmented Generation: A Representation Perspective 21 Nov 2024 · 0 repositories · arXiv:2411.14572
-
Combining Autoregressive and Autoencoder Language Models for Text Classification 20 Nov 2024 · 1 repository · arXiv:2411.13282
-
DMQR-RAG: Diverse Multi-Query Rewriting for RAG 20 Nov 2024 · 0 repositories · arXiv:2411.13154
-
Multimodal large language model for wheat breeding: a new exploration of smart breeding 20 Nov 2024 · 0 repositories · arXiv:2411.15203
-
On the Way to LLM Personalization: Learning to Remember User Conversations 20 Nov 2024 · 0 repositories · arXiv:2411.13405
-
Retrieval-Augmented Generation for Domain-Specific Question Answering: A Case Study on Pittsburgh and CMU 20 Nov 2024 · 0 repositories · arXiv:2411.13691
-
Unlocking Historical Clinical Trial Data with ALIGN: A Compositional Large Language Model System for Medical Coding 20 Nov 2024 · 0 repositories · arXiv:2411.13163
-
DLBacktrace: A Model Agnostic Explainability for any Deep Learning Models 19 Nov 2024 · 1 repository · arXiv:2411.12643
-
Enhancing Multi-Class Disease Classification: Neoplasms, Cardiovascular, Nervous System, and Digestive Disorders Using Advanced LLMs 19 Nov 2024 · 0 repositories · arXiv:2411.12712
-
Strengthening Fake News Detection: Leveraging SVM and Sophisticated Text Vectorization Techniques. Defying BERT? 19 Nov 2024 · 0 repositories · arXiv:2411.12703
-
CNMBERT: A Model for Converting Hanyu Pinyin Abbreviations to Chinese Characters 18 Nov 2024 · 1 repository · arXiv:2411.11770
-
Suicide Risk Assessment on Social Media with Semi-Supervised Learning 18 Nov 2024 · 0 repositories · arXiv:2411.12767
-
Understanding Student Sentiment on Mental Health Support in Colleges Using Large Language Models 18 Nov 2024 · 0 repositories · arXiv:2412.04326
-
A Novel Approach to Eliminating Hallucinations in Large Language Model-Assisted Causal Discovery 16 Nov 2024 · 0 repositories · arXiv:2411.12759
-
Debias-CLR: A Contrastive Learning Based Debiasing Method for Algorithmic Fairness in Healthcare Applications 15 Nov 2024 · 0 repositories · arXiv:2411.10544
-
Hysteresis Activation Function for Efficient Inference 15 Nov 2024 · 1 repository · arXiv:2411.10573Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Information Extraction from Clinical Notes: Are We Ready to Switch to Large Language Models? 15 Nov 2024 · 1 repository · arXiv:2411.10020
-
SoftLMs: Efficient Adaptive Low-Rank Approximation of Language Models using Soft-Thresholding Mechanism 15 Nov 2024 · 0 repositories · arXiv:2411.10543
-
Adopting RAG for LLM-Aided Future Vehicle Design 14 Nov 2024 · 0 repositories · arXiv:2411.09590
-
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering 14 Nov 2024 · 0 repositories · arXiv:2411.09213
-
Initial Nugget Evaluation Results for the TREC 2024 RAG Track with the AutoNuggetizer Framework 14 Nov 2024 · 2 repositories · arXiv:2411.09607
-
A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look 13 Nov 2024 · 1 repository · arXiv:2411.08275
-
Analyst Reports and Stock Performance: Evidence from the Chinese Market 13 Nov 2024 · 0 repositories · arXiv:2411.08726
-
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection 13 Nov 2024 · 0 repositories · arXiv:2411.08868
-
LogLLM: Log-based Anomaly Detection Using Large Language Models 13 Nov 2024 · 1 repository · arXiv:2411.08561
-
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data 13 Nov 2024 · 0 repositories · arXiv:2411.08438
-
Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models 12 Nov 2024 · 1 repository · arXiv:2411.07474
-
Query Optimization for Parametric Knowledge Refinement in Retrieval-Augmented Large Language Models 12 Nov 2024 · 0 repositories · arXiv:2411.07820
-
Retrieval Augmented Time Series Forecasting 12 Nov 2024 · 1 repository · arXiv:2411.08249Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Trustful LLMs: Customizing and Grounding Text Generation with Knowledge Bases and Dual Decoders 12 Nov 2024 · 0 repositories · arXiv:2411.07870
-
A Unified Multi-Task Learning Architecture for Hate Detection Leveraging User-Based Information 11 Nov 2024 · 0 repositories · arXiv:2411.06855
-
AssistRAG: Boosting the Potential of Large Language Models with an Intelligent Information Assistant 11 Nov 2024 · 1 repository · arXiv:2411.06805Syntology official (archive's flag): 10 ran · 10 ran (of which 1 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Autonomous Droplet Microfluidic Design Framework with Large Language Models 11 Nov 2024 · 1 repository · arXiv:2411.06691
-
Evaluating Large Language Models on Financial Report Summarization: An Empirical Study 11 Nov 2024 · 0 repositories · arXiv:2411.06852
-
Invar-RAG: Invariant LLM-aligned Retrieval for Better Generation 11 Nov 2024 · 0 repositories · arXiv:2411.07021
-
LA4SR: illuminating the dark proteome with generative AI 11 Nov 2024 · 0 repositories · arXiv:2411.06798
-
TempCharBERT: Keystroke Dynamics for Continuous Access Control Based on Pre-trained Language Models 11 Nov 2024 · 0 repositories · arXiv:2411.07224
-
The Backpropagation of the Wave Network 11 Nov 2024 · 0 repositories · arXiv:2411.06989
-
Toward Optimal Search and Retrieval for RAG 11 Nov 2024 · 1 repository · arXiv:2411.07396
-
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement 10 Nov 2024 · 1 repository · arXiv:2411.06558Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Clustering Algorithms and RAG Enhancing Semi-Supervised Text Classification with Large LLMs 9 Nov 2024 · 0 repositories · arXiv:2411.06175
-
Exploring Knowledge Boundaries in Large Language Models for Retrieval Judgment 9 Nov 2024 · 0 repositories · arXiv:2411.06207
-
Improved intent classification based on context information using a windows-based approach 9 Nov 2024 · 0 repositories · arXiv:2411.06022
-
Leveraging Retrieval-Augmented Generation for Persian University Knowledge Retrieval 9 Nov 2024 · 0 repositories · arXiv:2411.06237
-
Robust Detection of LLM-Generated Text: A Comparative Analysis 9 Nov 2024 · 0 repositories · arXiv:2411.06248
-
Sufficient Context: A New Lens on Retrieval Augmented Generation Systems 9 Nov 2024 · 0 repositories · arXiv:2411.06037
-
AgentOps: Enabling Observability of LLM Agents 8 Nov 2024 · 1 repository · arXiv:2411.05285
-
FinDVer: Explainable Claim Verification over Long and Hybrid-Content Financial Documents 8 Nov 2024 · 1 repository · arXiv:2411.05764
-
HeartBERT: A Self-Supervised ECG Embedding Model for Efficient and Effective Medical Signal Analysis 8 Nov 2024 · 1 repository · arXiv:2411.11896
-
IntellBot: Retrieval Augmented LLM Chatbot for Cyber Threat Knowledge Delivery 8 Nov 2024 · 1 repository · arXiv:2411.05442
-
Multi-Document Financial Question Answering using LLMs 8 Nov 2024 · 0 repositories · arXiv:2411.07264
-
Qwen2.5-32B: Leveraging Self-Consistent Tool-Integrated Reasoning for Bengali Mathematical Olympiad Problem Solving 8 Nov 2024 · 0 repositories · arXiv:2411.05934
-
Sentiment Analysis of Cyberbullying Data in Social Media 8 Nov 2024 · 1 repository · arXiv:2411.05958