{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cats-and-captions-vs-creators-and-the-clock","title":"Cats and Captions vs. Creators and the Clock: Comparing Multimodal Content to Context in Predicting Relative Popularity","arxiv_id":"1703.01725","date":"2017-03-06","proceeding":null,"authors":["Jack Hessel","Lillian Lee","David Mimno"],"abstract":"The content of today's social media is becoming more and more rich,\nincreasingly mixing text, images, videos, and audio. It is an intriguing\nresearch question to model the interplay between these different modes in\nattracting user attention and engagement. But in order to pursue this study of\nmultimodal content, we must also account for context: timing effects, community\npreferences, and social factors (e.g., which authors are already popular) also\naffect the amount of feedback and reaction that social-media posts receive. In\nthis work, we separate out the influence of these non-content factors in\nseveral ways. First, we focus on ranking pairs of submissions posted to the\nsame community in quick succession, e.g., within 30 seconds, this framing\nencourages models to focus on time-agnostic and community-specific content\nfeatures. Within that setting, we determine the relative performance of author\nvs. content features. We find that victory usually belongs to \"cats and\ncaptions,\" as visual and textual features together tend to outperform\nidentity-based features. Moreover, our experiments show that when considered in\nisolation, simple unigram text features and deep neural network visual features\nyield the highest accuracy individually, and that the combination of the two\nmodalities generally leads to the best accuracies overall.","url_abs":"http://arxiv.org/abs/1703.01725v1","url_pdf":"http://arxiv.org/pdf/1703.01725v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cats-and-captions-vs-creators-and-the-clock","repo_url":"https://github.com/jmhessel/catrank","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.01725","atlas_url":"https://app.syntology.ai/?focus=1703.01725","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}