{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-character-level-convolutional-neural","title":"A Character-level Convolutional Neural Network for Distinguishing Similar Languages and Dialects","arxiv_id":"1609.07568","date":"2016-09-24","proceeding":"WS 2016 12","authors":["Yonatan Belinkov","James Glass"],"abstract":"Discriminating between closely-related language varieties is considered a\nchallenging and important task. This paper describes our submission to the DSL\n2016 shared-task, which included two sub-tasks: one on discriminating similar\nlanguages and one on identifying Arabic dialects. We developed a\ncharacter-level neural network for this task. Given a sequence of characters,\nour model embeds each character in vector space, runs the sequence through\nmultiple convolutions with different filter widths, and pools the convolutional\nrepresentations to obtain a hidden vector representation of the text that is\nused for predicting the language or dialect. We primarily focused on the Arabic\ndialect identification task and obtained an F1 score of 0.4834, ranking 6th out\nof 18 participants. We also analyze errors made by our system on the Arabic\ndata in some detail, and point to challenges such an approach is faced with.","url_abs":"http://arxiv.org/abs/1609.07568v1","url_pdf":"http://arxiv.org/pdf/1609.07568v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-character-level-convolutional-neural","repo_url":"https://github.com/boknilev/dsl-char-cnn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"dialect-identification","task_name":"Dialect Identification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1609.07568","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}