{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/preparation-of-improved-turkish-dataset-for","title":"Preparation of Improved Turkish DataSet for Sentiment Analysis in Social Media","arxiv_id":"1801.09975","date":"2018-01-30","proceeding":null,"authors":["Semiha Makinist","Ibrahim Riza Hallac","Betul Ay Karakus","Galip Aydin"],"abstract":"A public dataset, with a variety of properties suitable for sentiment\nanalysis [1], event prediction, trend detection and other text mining\napplications, is needed in order to be able to successfully perform analysis\nstudies. The vast majority of data on social media is text-based and it is not\npossible to directly apply machine learning processes into these raw data,\nsince several different processes are required to prepare the data before the\nimplementation of the algorithms. For example, different misspellings of same\nword enlarge the word vector space unnecessarily, thereby it leads to reduce\nthe success of the algorithm and increase the computational power requirement.\nThis paper presents an improved Turkish dataset with an effective spelling\ncorrection algorithm based on Hadoop [2]. The collected data is recorded on the\nHadoop Distributed File System and the text based data is processed by\nMapReduce programming model. This method is suitable for the storage and\nprocessing of large sized text based social media data. In this study, movie\nreviews have been automatically recorded with Apache ManifoldCF (MCF) [3] and\ndata clusters have been created. Various methods compared such as Levenshtein\nand Fuzzy String Matching have been proposed to create a public dataset from\ncollected data. Experimental results show that the proposed algorithm, which\ncan be used as an open source dataset in sentiment analysis studies, have been\nperformed successfully to the detection and correction of spelling errors.","url_abs":"http://arxiv.org/abs/1801.09975v2","url_pdf":"http://arxiv.org/pdf/1801.09975v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"preparation-of-improved-turkish-dataset-for","repo_url":"https://github.com/sevvalckc/Turkish-SAD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"},{"task_slug":"spelling-correction","task_name":"Spelling Correction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}