Browse State-of-the-Art › Language Identification › Papers, page 2
Language Identification
Papers archive 2025-07-28
archive papers tagged: 794 · with a code link: 143 · where Syntology ran a sample: 14 (10 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (14 of 794 tagged: 10 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument)
Page 2 of 8: papers 101 to 200 of 794, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
1 Dec 2020 1 repository listed
-
27 Oct 2020 1 repository listed
-
25 Oct 2020 1 repository listed
-
11 Oct 2020 1 repository listed Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
2 Oct 2020 1 repository listed
-
22 Sep 2020 1 repository listed
-
4 Aug 2020 1 repository listed
-
2 Aug 2020 1 repository listed
-
26 Jul 2020 1 repository listed Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
21 Jul 2020 1 repository listed
-
9 Apr 2020 1 repository listed Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
16 Mar 2020 1 repository listed
-
18 Nov 2019 1 repository listed
-
11 Sep 2019 1 repository listed
-
1 Sep 2019 1 repository listed
-
29 Aug 2019 1 repository listed
-
Embeddia at SemEval-2019 Task 6: Detecting Hate with Neural Network and Transfer Learning Approaches1 Jun 2019 1 repository listed
-
1 Jun 2019 1 repository listed
-
16 Mar 2019 1 repository listed
-
25 Feb 2019 1 repository listed
-
10 Dec 2018 1 repository listed
-
29 Nov 2018 1 repository listed
-
1 Nov 2018 1 repository listed
-
9 Sep 2018 1 repository listed
-
1 Aug 2018 1 repository listed
-
1 Aug 2018 1 repository listed
-
1 Jun 2018 1 repository listed
-
16 May 2018 1 repository listed
-
1 May 2018 1 repository listed
-
22 Apr 2018 1 repository listed
-
1 Nov 2017 1 repository listed
-
1 Sep 2017 1 repository listed
-
16 Aug 2017 1 repository listed
-
1 May 2017 1 repository listed
-
1 Apr 2017 1 repository listed
-
12 Jan 2017 1 repository listed
-
1 Dec 2016 1 repository listed
-
10 Aug 2016 1 repository listed
-
1 Apr 2016 1 repository listed
-
23 Sep 2015 1 repository listed
-
1 May 2014 1 repository listed
-
1 Aug 2013 1 repository listed
-
8 Jul 2012 1 repository listed
-
mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks10 Jun 2025 0 repositories listed
-
Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?10 Jun 2025 0 repositories listed
-
Recursive Semantic Anchoring in ISO 639:2023: A Structural Extension to ISO/TC 37 Frameworks7 Jun 2025 0 repositories listed
-
TalTech Systems for the Interspeech 2025 ML-SUPERB 2.0 Challenge2 Jun 2025 0 repositories listed
-
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC30 May 2025 0 repositories listed
-
Token Masking Improves Transformer-Based Text Classification16 May 2025 0 repositories listed
-
Improving Informally Romanized Language Identification30 Apr 2025 0 repositories listed
-
(Im)possibility of Automated Hallucination Detection in Large Language Models23 Apr 2025 0 repositories listed
-
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing27 Mar 2025 0 repositories listed
-
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts25 Feb 2025 0 repositories listed
-
On the use of Performer and Agent Attention for Spoken Language Identification9 Feb 2025 0 repositories listed
-
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance7 Feb 2025 0 repositories listed
-
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID26 Dec 2024 0 repositories listed
-
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection26 Nov 2024 0 repositories listed
-
Exploring Facets of Language Generation in the Limit22 Nov 2024 0 repositories listed
-
Can adversarial attacks by large language models be attributed?12 Nov 2024 0 repositories listed
-
Prompt Engineering Using GPT for Word-Level Code-Mixed Language Identification in Low-Resource Dravidian Languages6 Nov 2024 0 repositories listed
-
Computational Approaches to Arabic-English Code-Switching17 Oct 2024 0 repositories listed
-
Generation through the lens of learning theory17 Oct 2024 0 repositories listed
-
A Multi-Task Text Classification Pipeline with Natural Language Explanations: A User-Centric Evaluation in Sentiment Analysis and Offensive Language Identification in Greek Tweets14 Oct 2024 0 repositories listed
-
Efficiently Identifying Low-Quality Language Subsets in Multilingual Datasets: A Case Study on a Large-Scale Multilingual Audio Dataset5 Oct 2024 0 repositories listed
-
Leveraging Open-Source Large Language Models for Native Language Identification15 Sep 2024 0 repositories listed
-
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model3 Sep 2024 0 repositories listed
-
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features27 Aug 2024 0 repositories listed
-
Towards Generalized Offensive Language Identification26 Jul 2024 0 repositories listed
-
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models29 Jun 2024 0 repositories listed
-
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR26 Jun 2024 0 repositories listed
-
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting18 Jun 2024 0 repositories listed
-
Exploring Spoken Language Identification Strategies for Automatic Transcription of Multilingual Broadcast and Institutional Speech13 Jun 2024 0 repositories listed
-
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets12 Jun 2024 0 repositories listed
-
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation12 Jun 2024 0 repositories listed
-
Malayalam Sign Language Identification using Finetuned YOLOv8 and Computer Vision Techniques8 May 2024 0 repositories listed
-
Whispy: Adapting STT Whisper Models to Real-Time Environments6 May 2024 0 repositories listed
-
A Federated Learning Approach to Privacy Preserving Offensive Language Identification17 Apr 2024 0 repositories listed
-
More than words: Advancements and challenges in speech recognition for singing14 Mar 2024 0 repositories listed
-
Validating and Exploring Large Geographic Corpora13 Mar 2024 0 repositories listed
-
Aligning Speech to Languages to Enhance Code-switching Speech Recognition9 Mar 2024 0 repositories listed
-
Detecting Structured Language Alternations in Historical Documents by Combining Language Identification with Fourier Analysis25 Jan 2024 0 repositories listed
-
Acoustic characterization of speech rhythm: going beyond metrics with recurrent neural networks22 Jan 2024 0 repositories listed
-
Language Detection for Transliterated Content9 Jan 2024 0 repositories listed
-
Generative linguistic representation for spoken language identification18 Dec 2023 0 repositories listed
-
Cross-Linguistic Offensive Language Detection: BERT-Based Analysis of Bengali, Assamese, & Bodo Conversational Hateful Content from Social Media16 Dec 2023 0 repositories listed
-
Leveraging Language ID to Calculate Intermediate CTC Loss for Enhanced Code-Switching Speech Recognition15 Dec 2023 0 repositories listed
-
Attention-Guided Adaptation for Code-Switching Speech Recognition14 Dec 2023 0 repositories listed
-
Native Language Identification with Large Language Models13 Dec 2023 0 repositories listed
-
Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification12 Dec 2023 0 repositories listed
-
A Text-to-Text Model for Multilingual Offensive Language Identification6 Dec 2023 0 repositories listed
-
Offensive Language Identification in Transliterated and Code-Mixed Bangla25 Nov 2023 0 repositories listed
-
The Obscure Limitation of Modular Multilingual Language Models21 Nov 2023 0 repositories listed
-
Fumbling in Babel: An Investigation into ChatGPT's Language Identification Ability16 Nov 2023 0 repositories listed
-
Advanced accent/dialect identification and accentedness assessment with multi-embedding models and automatic speech recognition17 Oct 2023 0 repositories listed
-
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond9 Oct 2023 0 repositories listed
-
Wavelet Scattering Transform for Improving Generalization in Low-Resourced Spoken Language Identification1 Oct 2023 0 repositories listed
-
Multimodal Modeling For Spoken Language Identification19 Sep 2023 0 repositories listed
-
17 Sep 2023 0 repositories listed
-
Robust Open-Set Spoken Language Identification and the CU MultiLang Dataset29 Aug 2023 0 repositories listed
-
Fine-Tuning Llama 2 Large Language Models for Detecting Online Sexual Predatory Chats and Abusive Texts28 Aug 2023 0 repositories listed
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.