Methods › Natural Language Processing › Static Word Embeddings › fastText › Papers, page 2
fastText
Papers archive 2025-07-28
archive papers tagged: 240 · with a code link: 76 · where Syntology ran a sample: 10 (8 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (10 of 240 tagged: 8 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument)
Page 2 of 3: papers 101 to 200 of 240, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Evaluating Document Representations for Content-based Legal Literature Recommendations 28 Apr 2021 · 1 repository · arXiv:2104.13841
-
Are Word Embedding Methods Stable and Should We Care About It? 17 Apr 2021 · 0 repositories · arXiv:2104.08433
-
SuperSim: a test set for word similarity and relatedness in Swedish 12 Apr 2021 · 0 repositories · arXiv:2104.05228
-
Text2Chart: A Multi-Staged Chart Generator from Natural Language Text 9 Apr 2021 · 1 repository · arXiv:2104.04584
-
Semi-Supervised Classification of Social Media Posts: Identifying Sex-Industry Posts to Enable Better Support for Those Experiencing Sex-Trafficking 7 Apr 2021 · 0 repositories · arXiv:2104.03233
-
Sentiment Analysis of Code-Mixed Social Media Text (Hinglish) 24 Feb 2021 · 0 repositories · arXiv:2102.12149
-
Co-occurrences using Fasttext embeddings for word similarity tasks in Urdu 22 Feb 2021 · 1 repository · arXiv:2102.10957
-
Towards Emotion Recognition in Hindi-English Code-Mixed Data: A Transformer Based Approach 19 Feb 2021 · 1 repository · arXiv:2102.09943
-
An AutoML-based Approach to Multimodal Image Sentiment Analysis 16 Feb 2021 · 0 repositories · arXiv:2102.08092
-
One Size Does Not Fit All: Finding the Optimal Subword Sizes for FastText Models across Languages 4 Feb 2021 · 0 repositories · arXiv:2102.02585
-
Does a Hybrid Neural Network based Feature Selection Model Improve Text Classification? 22 Jan 2021 · 0 repositories · arXiv:2101.09009
-
HinFlair: pre-trained contextual string embeddings for pos tagging and text classification in the Hindi language 18 Jan 2021 · 0 repositories · arXiv:2101.06949
-
Hostility Detection and Covid-19 Fake News Detection in Social Media 15 Jan 2021 · 0 repositories · arXiv:2101.05953
-
Experimental Evaluation of Deep Learning models for Marathi Text Classification 13 Jan 2021 · 0 repositories · arXiv:2101.04899
-
Evaluation of Deep Learning Models for Hostility Detection in Hindi Text 11 Jan 2021 · 0 repositories · arXiv:2101.04144
-
Graph-of-Tweets: A Graph Merging Approach to Sub-event Identification 8 Jan 2021 · 1 repository · arXiv:2101.03208
-
Language Detection Engine for Multilingual Texting on Mobile Devices 7 Jan 2021 · 0 repositories · arXiv:2101.03963
-
Faster Training of Word Embeddings 1 Jan 2021 · 0 repositories
-
Hate Speech detection in the Bengali language: A dataset and its baseline evaluation 17 Dec 2020 · 0 repositories · arXiv:2012.09686
-
DeftPunk at SemEval-2020 Task 6: Using RNN-ensemble for the Sentence Classification. 1 Dec 2020 · 0 repositories
-
Go Simple and Pre-Train on Domain-Specific Corpora: On the Role of Training Data for Text Classification 1 Dec 2020 · 0 repositories
-
IIITG-ADBU at SemEval-2020 Task 8: A Multimodal Approach to Detect Offensive, Sarcastic and Humorous Memes 1 Dec 2020 · 0 repositories
-
NLP_Passau at SemEval-2020 Task 12: Multilingual Neural Network for Offensive Language Detection in English, Danish and Turkish 1 Dec 2020 · 0 repositories
-
Word Embedding Binarization with Semantic Information Preservation 1 Dec 2020 · 0 repositories
-
Blind signal decomposition of various word embeddings based on join and individual variance explained 30 Nov 2020 · 0 repositories · arXiv:2011.14496
-
Improving Clinical Outcome Predictions Using Convolution over Medical Entities with Multimodal Learning 24 Nov 2020 · 1 repository · arXiv:2011.12349
-
STEPs-RL: Speech-Text Entanglement for Phonetically Sound Representation Learning 23 Nov 2020 · 0 repositories · arXiv:2011.11387
-
DebateSum: A large-scale argument mining and summarization dataset 14 Nov 2020 · 3 repositories · arXiv:2011.07251Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
HyperText: Endowing FastText with Hyperbolic Geometry 30 Oct 2020 · 3 repositories · arXiv:2010.16143Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Gender Prediction Based on Vietnamese Names with Machine Learning Techniques 21 Oct 2020 · 0 repositories · arXiv:2010.10852
-
LT3 at SemEval-2020 Task 9: Cross-lingual Embeddings for Sentiment Analysis of Hinglish Social Media Text 21 Oct 2020 · 0 repositories · arXiv:2010.11019
-
WNUT-2020 Task 2: Identification of Informative COVID-19 English Tweets 16 Oct 2020 · 1 repository · arXiv:2010.08232
-
Geometry matters: Exploring language examples at the decision boundary 14 Oct 2020 · 0 repositories · arXiv:2010.07212
-
gundapusunil at SemEval-2020 Task 9: Syntactic Semantic LSTM Architecture for SENTIment Analysis of Code-MIXed Data 9 Oct 2020 · 0 repositories · arXiv:2010.04395
-
Intrinsic Probing through Dimension Selection 6 Oct 2020 · 1 repository · arXiv:2010.02812Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Pareto Probing: Trading Off Accuracy for Complexity 5 Oct 2020 · 1 repository · arXiv:2010.02180Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
"Did you really mean what you said?" : Sarcasm Detection in Hindi-English Code-Mixed Data using Bilingual Word Embeddings 1 Oct 2020 · 1 repository · arXiv:2010.00310
-
Development of Word Embeddings for Uzbek Language 30 Sep 2020 · 0 repositories · arXiv:2009.14384
-
FarsTail: A Persian Natural Language Inference Dataset 18 Sep 2020 · 1 repository · arXiv:2009.08820
-
EdinburghNLP at WNUT-2020 Task 2: Leveraging Transformers with Generalized Augmentation for Identifying Informativeness in COVID-19 Tweets 6 Sep 2020 · 0 repositories · arXiv:2009.06375
-
Going Beyond T-SNE: Exposing whatlies in Text Embeddings 4 Sep 2020 · 1 repository · arXiv:2009.02113
-
Beyond Next Item Recommendation: Recommending and Evaluating List of Sequences 30 Aug 2020 · 0 repositories · arXiv:2008.13281
-
A Study of fastText Word Embedding Effects in Document Classification in Bangla Language 30 Jul 2020 · 1 repository
-
COVID-19 therapy target discovery with context-aware literature mining 30 Jul 2020 · 0 repositories · arXiv:2007.15681
-
Word Embeddings: Stability and Semantic Change 23 Jul 2020 · 0 repositories · arXiv:2007.16006
-
Morphological Skip-Gram: Using morphological knowledge to improve word representation 20 Jul 2020 · 0 repositories · arXiv:2007.10055
-
Estimating the effect of COVID-19 on mental health: Linguistic indicators of depression during a global pandemic 1 Jul 2020 · 0 repositories
-
On the Learnability of Concepts: With Applications to Comparing Word Embedding Algorithms 17 Jun 2020 · 0 repositories · arXiv:2006.09896
-
qDKT: Question-centric Deep Knowledge Tracing 25 May 2020 · 0 repositories · arXiv:2005.12442
-
Adversarial Alignment of Multilingual Models for Extracting Temporal Expressions from Text 19 May 2020 · 0 repositories · arXiv:2005.09392
-
Categorical Vector Space Semantics for Lambek Calculus with a Relevant Modality 6 May 2020 · 0 repositories · arXiv:2005.03074
-
A First Dataset for Film Age Appropriateness Investigation 1 May 2020 · 0 repositories
-
AI_ML_NIT_Patna @ TRAC - 2: Deep Learning Approach for Multi-lingual Aggression Identification 1 May 2020 · 0 repositories
-
CBOW-tag: a Modified CBOW Algorithm for Generating Embedding Models from Annotated Corpora 1 May 2020 · 0 repositories
-
Czech Historical Named Entity Corpus v 1.0 1 May 2020 · 0 repositories
-
Evaluating the Impact of Sub-word Information and Cross-lingual Word Embeddings on Mi'kmaq Language Modelling 1 May 2020 · 0 repositories
-
Facilitating Corpus Usage: Making Icelandic Corpora More Accessible for Researchers and Language Users 1 May 2020 · 0 repositories
-
High Quality ELMo Embeddings for Seven Less-Resourced Languages 1 May 2020 · 0 repositories
-
Identifying Cognates in English-Dutch and French-Dutch by means of Orthographic Information and Cross-lingual Word Embeddings 1 May 2020 · 0 repositories
-
KLEJ: Comprehensive Benchmark for Polish Language Understanding 1 May 2020 · 1 repository · arXiv:2005.00630Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
TALN/LS2N Participation at the BUCC Shared Task: Bilingual Dictionary Induction from Comparable Corpora 1 May 2020 · 0 repositories
-
UNIOR NLP at MWSA Task - GlobaLex 2020: Siamese LSTM with Attention for Word Sense Alignment 1 May 2020 · 0 repositories
-
Word Embedding Evaluation for Sinhala 1 May 2020 · 0 repositories
-
CrisisBench: Benchmarking Crisis-related Social Media Datasets for Humanitarian Information Processing 14 Apr 2020 · 0 repositories · arXiv:2004.06774
-
DeepSentiPers: Novel Deep Learning Models Trained Over Proposed Augmented Persian Sentiment Corpus 11 Apr 2020 · 1 repository · arXiv:2004.05328
-
Word Sense Disambiguation for 158 Languages using Word Embeddings Only 14 Mar 2020 · 0 repositories · arXiv:2003.06651
-
Multi-SimLex: A Large-Scale Evaluation of Multilingual and Cross-Lingual Lexical Semantic Similarity 10 Mar 2020 · 0 repositories · arXiv:2003.04866
-
Discovering linguistic (ir)regularities in word embeddings through max-margin separating hyperplanes 7 Mar 2020 · 0 repositories · arXiv:2003.03654
-
Comparison of Turkish Word Representations Trained on Different Morphological Forms 13 Feb 2020 · 0 repositories · arXiv:2002.05417
-
Generating Sense Embeddings for Syntactic and Semantic Analogy for Portuguese 21 Jan 2020 · 1 repository · arXiv:2001.07574
-
ExEm: Expert Embedding using dominating set theory with deep learning approaches 16 Jan 2020 · 2 repositories · arXiv:2001.08503
-
A Unified System for Aggression Identification in English Code-Mixed and Uni-Lingual Texts 15 Jan 2020 · 0 repositories · arXiv:2001.05493
-
Character 3-gram Mover's Distance: An Effective Method for Detecting Near-duplicate Japanese-language Recipes 11 Dec 2019 · 0 repositories · arXiv:1912.05171
-
Word Embedding based New Corpus for Low-resourced Language: Sindhi 28 Nov 2019 · 0 repositories · arXiv:1911.12579
-
hauWE: Hausa Words Embedding for Natural Language Processing 25 Nov 2019 · 0 repositories · arXiv:1911.10708
-
High Quality ELMo Embeddings for Seven Less-Resourced Languages 22 Nov 2019 · 0 repositories · arXiv:1911.10049
-
Multilingual Culture-Independent Word Analogy Datasets 22 Nov 2019 · 0 repositories · arXiv:1911.10038
-
Event detection in Colombian security Twitter news using fine-grained latent topic analysis 19 Nov 2019 · 0 repositories · arXiv:1911.08370
-
Towards non-toxic landscapes: Automatic toxic comment detection using DNN 19 Nov 2019 · 0 repositories · arXiv:1911.08395
-
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data 1 Nov 2019 · 2 repositories · arXiv:1911.00359
-
Sentence Embeddings for Russian NLU 29 Oct 2019 · 1 repository · arXiv:1910.13291
-
Using machine learning and information visualisation for discovering latent topics in Twitter news 21 Oct 2019 · 0 repositories · arXiv:1910.09114
-
Privacy- and Utility-Preserving Textual Analysis via Calibrated Multivariate Perturbations 20 Oct 2019 · 1 repository · arXiv:1910.08902
-
Uncovering Flaming Events on News Media in Social Media 16 Sep 2019 · 0 repositories · arXiv:1909.07181
-
Dialogue Act Classification in Team Communication for Robot Assisted Disaster Response 1 Sep 2019 · 0 repositories
-
Evaluation of Stacked Embeddings for Bulgarian on the Downstream Tasks POS and NERC 1 Sep 2019 · 0 repositories
-
Evaluation of vector embedding models in clustering of text documents 1 Sep 2019 · 0 repositories
-
Question Similarity in Community Question Answering: A Systematic Exploration of Preprocessing Methods and Models 1 Sep 2019 · 1 repository
-
Sparse Victory -- A Large Scale Systematic Comparison of count-based and prediction-based vectorizers for text classification 1 Sep 2019 · 1 repository
-
Tagger for Polish Computer Mediated Communication Texts 1 Sep 2019 · 0 repositories
-
Improving Word Embeddings Using Kernel PCA 1 Aug 2019 · 0 repositories
-
Learning Word Embeddings without Context Vectors 1 Aug 2019 · 0 repositories
-
RNN Embeddings for Identifying Difficult to Understand Medical Words 1 Aug 2019 · 1 repository
-
Robust to Noise Models in Natural Language Processing Tasks 1 Jul 2019 · 2 repositories
-
Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion 27 Jun 2019 · 0 repositories · arXiv:1906.11604
-
Word Embeddings for the Armenian Language: Intrinsic and Extrinsic Evaluation 7 Jun 2019 · 1 repository · arXiv:1906.03134
-
Sequence Tagging with Contextual and Non-Contextual Subword Representations: A Multilingual Evaluation 4 Jun 2019 · 1 repository · arXiv:1906.01569
-
Beyond Context: A New Perspective for Word Embeddings 1 Jun 2019 · 0 repositories
-
Deep Learning Techniques for Humor Detection in Hindi-English Code-Mixed Tweets 1 Jun 2019 · 0 repositories
-
How Well Do Embedding Models Capture Non-compositionality? A View from Multiword Expressions 1 Jun 2019 · 0 repositories