Methods › Natural Language Processing › Subword Segmentation › WordPiece › Papers, page 22
WordPiece
Papers archive 2025-07-28
archive papers tagged: 7,063 · with a code link: 2,910 · where Syntology ran a sample: 650 (529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (650 of 7,063 tagged: 529 with a run with no instrument failure, 121 where every run was a failure of Syntology's instrument)
Page 22 of 71: papers 2,101 to 2,200 of 7,063, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Credit Risk Meets Large Language Models: Building a Risk Indicator from Loan Descriptions in P2P Lending 29 Jan 2024 · 0 repositories · arXiv:2401.16458
-
Development and Testing of a Novel Large Language Model-Based Clinical Decision Support Systems for Medication Safety in 12 Clinical Specialties 29 Jan 2024 · 0 repositories · arXiv:2402.01741
-
Development and Testing of Retrieval Augmented Generation in Large Language Models -- A Case Study Report 29 Jan 2024 · 0 repositories · arXiv:2402.01733
-
BPDec: Unveiling the Potential of Masked Language Modeling Decoder in BERT pretraining 29 Jan 2024 · 0 repositories · arXiv:2401.15861
-
Multi-class Regret Detection in Hindi Devanagari Script 29 Jan 2024 · 0 repositories · arXiv:2401.16561
-
Contrastive Learning and Mixture of Experts Enables Precise Vector Embeddings 28 Jan 2024 · 1 repository · arXiv:2401.15713
-
UnMASKed: Quantifying Gender Biases in Masked Language Models through Linguistically Informed Job Market Prompts 28 Jan 2024 · 0 repositories · arXiv:2401.15798
-
Enhancing Large Language Model Performance To Answer Questions and Extract Information More Accurately 27 Jan 2024 · 0 repositories · arXiv:2402.01722
-
MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries 27 Jan 2024 · 2 repositories · arXiv:2401.15391Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
From RAG to QA-RAG: Integrating Generative AI for Pharmaceutical Regulatory Compliance Process 26 Jan 2024 · 1 repository · arXiv:2402.01717
-
The Power of Noise: Redefining Retrieval for RAG Systems 26 Jan 2024 · 3 repositories · arXiv:2401.14887Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
(Chat)GPT v BERT: Dawn of Justice for Semantic Change Detection 25 Jan 2024 · 1 repository · arXiv:2401.14040
-
Enhanced Labeling Technique for Reddit Text and Fine-Tuned Longformer Models for Classifying Depression Severity in English and Luganda 25 Jan 2024 · 0 repositories · arXiv:2401.14240
-
Socially Aware Synthetic Data Generation for Suicidal Ideation Detection Using Large Language Models 25 Jan 2024 · 0 repositories · arXiv:2402.01712
-
Proactive Emotion Tracker: AI-Driven Continuous Mood and Emotion Monitoring 24 Jan 2024 · 0 repositories · arXiv:2401.13722
-
Segment Any Cell: A SAM-based Auto-prompting Fine-tuning Framework for Nuclei Segmentation 24 Jan 2024 · 0 repositories · arXiv:2401.13220
-
Contrastive Learning in Distilled Models 23 Jan 2024 · 1 repository · arXiv:2401.12472
-
Fast Adversarial Training against Textual Adversarial Attacks 23 Jan 2024 · 0 repositories · arXiv:2401.12461
-
Revolutionizing Retrieval-Augmented Generation with Enhanced PDF Structure Recognition 23 Jan 2024 · 0 repositories · arXiv:2401.12599
-
APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and Inference 22 Jan 2024 · 1 repository · arXiv:2401.12200Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers 22 Jan 2024 · 0 repositories · arXiv:2401.11700
-
Zero-Space Cost Fault Tolerance for Transformer-based Language Models on ReRAM 22 Jan 2024 · 0 repositories · arXiv:2401.11664
-
Confidence Preservation Property in Knowledge Distillation Abstractions 21 Jan 2024 · 0 repositories · arXiv:2401.11365
-
SEBERTNets: Sequence Enhanced BERT Networks for Event Entity Extraction Tasks Oriented to the Finance Field 21 Jan 2024 · 1 repository · arXiv:2401.11408
-
Drop your Decoder: Pre-training with Bag-of-Word Prediction for Dense Passage Retrieval 20 Jan 2024 · 3 repositories · arXiv:2401.11248Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Prompt-RAG: Pioneering Vector Embedding-Free Retrieval-Augmented Generation in Niche Domains, Exemplified by Korean Medicine 20 Jan 2024 · 0 repositories · arXiv:2401.11246
-
Unfair TOS: An Automated Approach using Customized BERT 20 Jan 2024 · 0 repositories · arXiv:2401.11207
-
Mining experimental data from Materials Science literature with Large Language Models: an evaluation study 19 Jan 2024 · 1 repository · arXiv:2401.11052
-
ChatQA: Surpassing GPT-4 on Conversational QA and RAG 18 Jan 2024 · 0 repositories · arXiv:2401.10225
-
BERTologyNavigator: Advanced Question Answering with BERT-based Semantics 17 Jan 2024 · 0 repositories · arXiv:2401.09553
-
Efficient slot labelling 17 Jan 2024 · 0 repositories · arXiv:2401.09343
-
Improving Classification Performance With Human Feedback: Label a few, we label the rest 17 Jan 2024 · 0 repositories · arXiv:2401.09555
-
A Reproducibility Study of Goldilocks: Just-Right Tuning of BERT for TAR 16 Jan 2024 · 1 repository · arXiv:2401.08104
-
RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture 16 Jan 2024 · 0 repositories · arXiv:2401.08406
-
A character-based steganography using masked language modeling 15 Jan 2024 · 1 repository
-
Graph database while computationally efficient filters out quickly the ESG integrated equities in investment management 15 Jan 2024 · 0 repositories · arXiv:2401.07483
-
Towards Efficient Methods in Medical Question Answering using Knowledge Graph Embeddings 15 Jan 2024 · 1 repository · arXiv:2401.07977Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Leveraging the power of transformers for guilt detection in text 15 Jan 2024 · 0 repositories · arXiv:2401.07414
-
SemEval-2017 Task 4: Sentiment Analysis in Twitter using BERT 15 Jan 2024 · 1 repository · arXiv:2401.07944
-
The Chronicles of RAG: The Retriever, the Chunk and the Generator 15 Jan 2024 · 0 repositories · arXiv:2401.07883
-
Understanding YTHDF2-mediated mRNA Degradation By m6A-BERT-Deg 15 Jan 2024 · 1 repository · arXiv:2401.08004
-
Promptformer: Prompted Conformer Transducer for ASR 14 Jan 2024 · 0 repositories · arXiv:2401.07360
-
Bridging the Preference Gap between Retrievers and LLMs 13 Jan 2024 · 0 repositories · arXiv:2401.06954
-
An investigation of structures responsible for gender bias in BERT and DistilBERT 12 Jan 2024 · 0 repositories · arXiv:2401.06495
-
Improved Learned Sparse Retrieval with Corpus-Specific Vocabularies 12 Jan 2024 · 1 repository · arXiv:2401.06703
-
Mapping Transformer Leveraged Embeddings for Cross-Lingual Document Representation 12 Jan 2024 · 1 repository · arXiv:2401.06583
-
Analyzing Regional Impacts of Climate Change using Natural Language Processing Techniques 11 Jan 2024 · 0 repositories · arXiv:2401.06817
-
Prompt-based mental health screening from social media text 11 Jan 2024 · 0 repositories · arXiv:2401.05912
-
Reinforcement Learning for Optimizing RAG for Domain Chatbots 10 Jan 2024 · 0 repositories · arXiv:2401.06800
-
An Assessment on Comprehending Mental Health through Large Language Models 9 Jan 2024 · 0 repositories · arXiv:2401.04592
-
DepressionEmo: A novel dataset for multilabel classification of depression emotions 9 Jan 2024 · 1 repository · arXiv:2401.04655
-
Language Detection for Transliterated Content 9 Jan 2024 · 0 repositories · arXiv:2401.04619
-
Phishing Website Detection through Multi-Model Analysis of HTML Content 9 Jan 2024 · 0 repositories · arXiv:2401.04820
-
An Exploratory Study on Automatic Identification of Assumptions in the Development of Deep Learning Frameworks 8 Jan 2024 · 1 repository · arXiv:2401.03653
-
Anatomy of Neural Language Models 8 Jan 2024 · 1 repository · arXiv:2401.03797
-
Advancing bioinformatics with large language models: components, applications and perspectives 8 Jan 2024 · 0 repositories · arXiv:2401.04155
-
PEneo: Unifying Line Extraction, Line Grouping, and Entity Linking for End-to-end Document Pair Extraction 7 Jan 2024 · 1 repository · arXiv:2401.03472
-
RoBERTurk: Adjusting RoBERTa for Turkish 7 Jan 2024 · 0 repositories · arXiv:2401.03515
-
PIXAR: Auto-Regressive Language Modeling in Pixel Space 6 Jan 2024 · 0 repositories · arXiv:2401.03321
-
Natural Language Programming in Medicine: Administering Evidence Based Clinical Workflows with Autonomous Agents Powered by Generative Large Language Models 5 Jan 2024 · 0 repositories · arXiv:2401.02851
-
German Text Embedding Clustering Benchmark 5 Jan 2024 · 1 repository · arXiv:2401.02709
-
Beyond Extraction: Contextualising Tabular Data for Efficient Summarisation by Language Models 4 Jan 2024 · 0 repositories · arXiv:2401.02333
-
L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages 4 Jan 2024 · 1 repository · arXiv:2401.02254
-
Studying and Recommending Information Highlighting in Stack Overflow Answers 3 Jan 2024 · 1 repository · arXiv:2401.01472
-
Enhancing Multilingual Information Retrieval in Mixed Human Resources Environments: A RAG Model Implementation for Multicultural Enterprise 3 Jan 2024 · 0 repositories · arXiv:2401.01511
-
Iterative Mask Filling: An Effective Text Augmentation Method Using Masked Language Modeling 3 Jan 2024 · 0 repositories · arXiv:2401.01830
-
MLPs Compass: What is learned when MLPs are combined with PLMs? 3 Jan 2024 · 0 repositories · arXiv:2401.01667
-
Natural Language Processing and Multimodal Stock Price Prediction 3 Jan 2024 · 0 repositories · arXiv:2401.01487
-
Revisiting Counterfactual Problems in Referring Expression Comprehension 1 Jan 2024 · 1 repository
-
An Analysis of Embedding Layers and Similarity Scores using Siamese Neural Networks 31 Dec 2023 · 0 repositories · arXiv:2401.00582
-
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models 31 Dec 2023 · 3 repositories · arXiv:2401.00396Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Advancing TTP Analysis: Harnessing the Power of Large Language Models with Retrieval Augmented Generation 30 Dec 2023 · 1 repository · arXiv:2401.00280
-
Why is the User Interface a Dark Pattern? : Explainable Auto-Detection and its Analysis 30 Dec 2023 · 1 repository · arXiv:2401.04119
-
MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining 29 Dec 2023 · 1 repository · arXiv:2312.17482
-
TuPy-E: detecting hate speech in Brazilian Portuguese social media with a novel dataset and comprehensive analysis of models 29 Dec 2023 · 1 repository · arXiv:2312.17704
-
Language Model as an Annotator: Unsupervised Context-aware Quality Phrase Generation 28 Dec 2023 · 0 repositories · arXiv:2312.17349
-
SentinelLMs: Encrypted Input Adaptation and Fine-tuning of Language Models for Private and Secure Inference 28 Dec 2023 · 1 repository · arXiv:2312.17342
-
Relationship between auditory and semantic entrainment using Deep Neural Networks (DNN) 27 Dec 2023 · 0 repositories · arXiv:2312.16599
-
Compositional Generalization in Spoken Language Understanding 25 Dec 2023 · 0 repositories · arXiv:2312.15815
-
Multi-level biomedical NER through multi-granularity embeddings and enhanced labeling 24 Dec 2023 · 0 repositories · arXiv:2312.15550
-
Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference 23 Dec 2023 · 1 repository · arXiv:2312.15159
-
Efficacy of Machine-Generated Instructions 22 Dec 2023 · 0 repositories · arXiv:2312.14423
-
Towards Detecting Cascades of Biased Medical Claims on Twitter 22 Dec 2023 · 0 repositories · arXiv:2312.15040
-
How Smooth Is Attention? 22 Dec 2023 · 0 repositories · arXiv:2312.14820
-
Unsupervised Auditory and Semantic Entrainment Models with Deep Neural Networks 22 Dec 2023 · 0 repositories · arXiv:2312.15098
-
ChatGPT as a commenter to the news: can LLMs generate human-like opinions? 21 Dec 2023 · 1 repository · arXiv:2312.13961
-
How to Prune Your Language Model: Recovering Accuracy on the "Sparsity May Cry'' Benchmark 21 Dec 2023 · 0 repositories · arXiv:2312.13547
-
Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy 20 Dec 2023 · 1 repository · arXiv:2312.12728Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Self-Admitted Technical Debt Detection Approaches: A Decade Systematic Review 19 Dec 2023 · 1 repository · arXiv:2312.15020
-
Efficient Title Reranker for Fast and Improved Knowledge-Intense NLP 19 Dec 2023 · 0 repositories · arXiv:2312.12430
-
"Knowing When You Don't Know": A Multilingual Relevance Assessment Dataset for Robust Retrieval-Augmented Generation 18 Dec 2023 · 1 repository · arXiv:2312.11361Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Retrieval-Augmented Generation for Large Language Models: A Survey 18 Dec 2023 · 4 repositories · arXiv:2312.10997
-
Bengali Intent Classification with Generative Adversarial BERT 17 Dec 2023 · 1 repository · arXiv:2312.10679
-
Can persistent homology whiten Transformer-based black-box models? A case study on BERT compression 17 Dec 2023 · 0 repositories · arXiv:2312.10702
-
Decoding Concerns: Multi-label Classification of Vaccine Sentiments in Social Media 17 Dec 2023 · 1 repository · arXiv:2312.10626
-
Investigating salient representations and label Variance in Dimensional Speech Emotion Analysis 17 Dec 2023 · 0 repositories · arXiv:2312.16180
-
Cross-Linguistic Offensive Language Detection: BERT-Based Analysis of Bengali, Assamese, & Bodo Conversational Hateful Content from Social Media 16 Dec 2023 · 0 repositories · arXiv:2312.10528
-
Investigating Shallow and Deep Learning Techniques for Emotion Classification in Short Persian Texts 16 Dec 2023 · 1 repository
-
SPT: Fine-Tuning Transformer-based Language Models Efficiently with Sparsification 16 Dec 2023 · 1 repository · arXiv:2312.10365
-
Algorithms for automatic intents extraction and utterances classification for goal-oriented dialogue systems 15 Dec 2023 · 0 repositories · arXiv:2312.09658