{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/interpreting-the-syntactic-and-social","title":"Interpreting the Syntactic and Social Elements of the Tweet Representations via Elementary Property Prediction Tasks","arxiv_id":"1611.04887","date":"2016-11-15","proceeding":null,"authors":["J Ganesh","Manish Gupta","Vasudeva Varma"],"abstract":"Research in social media analysis is experiencing a recent surge with a large\nnumber of works applying representation learning models to solve high-level\nsyntactico-semantic tasks such as sentiment analysis, semantic textual\nsimilarity computation, hashtag prediction and so on. Although the performance\nof the representation learning models are better than the traditional baselines\nfor the tasks, little is known about the core properties of a tweet encoded\nwithin the representations. Understanding these core properties would empower\nus in making generalizable conclusions about the quality of representations.\nOur work presented here constitutes the first step in opening the black-box of\nvector embedding for social media posts, with emphasis on tweets in particular.\n  In order to understand the core properties encoded in a tweet representation,\nwe evaluate the representations to estimate the extent to which it can model\neach of those properties such as tweet length, presence of words, hashtags,\nmentions, capitalization, and so on. This is done with the help of multiple\nclassifiers which take the representation as input. Essentially, each\nclassifier evaluates one of the syntactic or social properties which are\narguably salient for a tweet. This is also the first holistic study on\nextensively analysing the ability to encode these properties for a wide variety\nof tweet representation models including the traditional unsupervised methods\n(BOW, LDA), unsupervised representation learning methods (Siamese CBOW,\nTweet2Vec) as well as supervised methods (CNN, BLSTM).","url_abs":"http://arxiv.org/abs/1611.04887v1","url_pdf":"http://arxiv.org/pdf/1611.04887v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"interpreting-the-syntactic-and-social","repo_url":"https://github.com/ganeshjawahar/fine-tweet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"property-prediction","task_name":"Property Prediction"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}