Methods › Natural Language Processing › Static Word Embeddings › fastText
fastText
Introduced by Piotr Bojanowski et al. in Enriching Word Vectors with Subword Information
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
fastText embeddings exploit subword information to construct word embeddings. Representations are learnt of character n-grams, and words represented as the sum of the n-gram vectors. This extends the word2vec type models with subword information. This helps the embeddings understand suffixes and prefixes. Once a word is represented using character n-grams, a skipgram model is trained to learn the embeddings.
Papers archive 2025-07-28
30 shown of 240, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Speak2Sign3D: A Multi-modal Pipeline for English Speech to American Sign Language Animation 9 Jul 2025 · 0 repositories · arXiv:2507.06530
-
Guarded Query Routing for Large Language Models 20 May 2025 · 1 repository · arXiv:2505.14524
-
Ultra-FineWeb: Efficient Data Filtering and Verification for High-Quality LLM Training Data 8 May 2025 · 0 repositories · arXiv:2505.05427
-
myNER: Contextualized Burmese Named Entity Recognition with Bidirectional LSTM and fastText Embeddings via Joint Training with POS Tagging 5 Apr 2025 · 1 repository · arXiv:2504.04038
-
A Data-driven Investigation of Euphemistic Language: Comparing the usage of "slave" and "servant" in 19th century US newspapers 19 Mar 2025 · 1 repository · arXiv:2503.15057
-
Text classification using machine learning methods 27 Feb 2025 · 0 repositories · arXiv:2502.19801
-
Poster: Long PHP webshell files detection based on sliding window attention 26 Feb 2025 · 1 repository · arXiv:2502.19257
-
A Multi-tiered Solution for Personalized Baggage Item Recommendations using FastText and Association Rule Mining 16 Jan 2025 · 0 repositories · arXiv:2501.09359
-
Sentiment Analysis in Twitter Social Network Centered on Cryptocurrencies Using Machine Learning 16 Jan 2025 · 0 repositories · arXiv:2501.09777
-
Research on Violent Text Detection System Based on BERT-fasttext Model 21 Dec 2024 · 0 repositories · arXiv:2412.16455
-
UnMA-CapSumT: Unified and Multi-Head Attention-driven Caption Summarization Transformer 16 Dec 2024 · 0 repositories · arXiv:2412.11836
-
On Importance of Code-Mixed Embeddings for Hate Speech Identification 27 Nov 2024 · 0 repositories · arXiv:2411.18577
-
BERT or FastText? A Comparative Analysis of Contextual as well as Non-Contextual Embeddings 26 Nov 2024 · 1 repository · arXiv:2411.17661
-
From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models 6 Nov 2024 · 0 repositories · arXiv:2411.05036
-
Generic Embedding-Based Lexicons for Transparent and Reproducible Text Scoring 1 Nov 2024 · 0 repositories · arXiv:2411.00964
-
LightFusionRec: Lightweight Transformers-Based Cross-Domain Recommendation Model 21 Oct 2024 · 0 repositories · arXiv:2410.15656
-
Stress Detection on Code-Mixed Texts in Dravidian Languages using Machine Learning 8 Oct 2024 · 0 repositories · arXiv:2410.06428
-
Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary? 4 Oct 2024 · 1 repository · arXiv:2410.03240
-
Individuation in Neural Models with and without Visual Grounding 27 Sep 2024 · 0 repositories · arXiv:2409.18868
-
An Evaluation of Sindhi Word Embedding in Semantic Analogies and Downstream Tasks 28 Aug 2024 · 0 repositories · arXiv:2408.15720
-
Representation Bias of Adolescents in AI: A Bilingual, Bicultural Study 4 Aug 2024 · 1 repository · arXiv:2408.01961
-
Constructing the CORD-19 Vaccine Dataset 26 Jul 2024 · 0 repositories · arXiv:2407.18471
-
Enhancing Depressive Post Detection in Bangla: A Comparative Study of TF-IDF, BERT and FastText Embeddings 12 Jul 2024 · 0 repositories · arXiv:2407.09187
-
MaskLID: Code-Switching Language Identification through Iterative Masking 10 Jun 2024 · 1 repository · arXiv:2406.06263
-
RICo: Reddit ideological communities 5 Jun 2024 · 1 repository
-
E2Vec: Feature Embedding with Temporal Information for Analyzing Student Actions in E-Book Systems 24 May 2024 · 1 repository · arXiv:2407.13053
-
RE-GrievanceAssist: Enhancing Customer Experience through ML-Powered Complaint Management 29 Apr 2024 · 0 repositories · arXiv:2404.18963
-
Deep CNN with late fusion for realtime multimodal emotion recognition 15 Apr 2024 · 0 repositories
-
FastSpell: the LangId Magic Spell 12 Apr 2024 · 1 repository · arXiv:2404.08345
-
Breaking the Silence Detecting and Mitigating Gendered Abuse in Hindi, Tamil, and Indian English Online Spaces 2 Apr 2024 · 1 repository · arXiv:2404.02013
Tasks archive 2025-07-28
20 shown of 214 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections