{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/automated-essay-scoring-with-string-kernels","title":"Automated essay scoring with string kernels and word embeddings","arxiv_id":"1804.07954","date":"2018-04-21","proceeding":"ACL 2018 7","authors":["Mădălina Cozma","Andrei M. Butnaru","Radu Tudor Ionescu"],"abstract":"In this work, we present an approach based on combining string kernels and\nword embeddings for automatic essay scoring. String kernels capture the\nsimilarity among strings based on counting common character n-grams, which are\na low-level yet powerful type of feature, demonstrating state-of-the-art\nresults in various text classification tasks such as Arabic dialect\nidentification or native language identification. To our best knowledge, we are\nthe first to apply string kernels to automatically score essays. We are also\nthe first to combine them with a high-level semantic feature representation,\nnamely the bag-of-super-word-embeddings. We report the best performance on the\nAutomated Student Assessment Prize data set, in both in-domain and cross-domain\nsettings, surpassing recent state-of-the-art deep learning approaches.","url_abs":"http://arxiv.org/abs/1804.07954v2","url_pdf":"http://arxiv.org/pdf/1804.07954v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"automated-essay-scoring","task_name":"Automated Essay Scoring"},{"task_slug":"dialect-identification","task_name":"Dialect Identification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"native-language-identification","task_name":"Native Language Identification"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/automated-essay-scoring-on-asap","task":"Automated Essay Scoring","dataset":"ASAP-AES","model":"HISK+BOSWE","rank_in_archive_order":4,"of":8,"metrics":{"Quadratic Weighted Kappa":"0.785"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.07954","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}