{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/twitter100k-a-real-world-dataset-for-weakly","title":"Twitter100k: A Real-world Dataset for Weakly Supervised Cross-Media Retrieval","arxiv_id":"1703.06618","date":"2017-03-20","proceeding":null,"authors":["Yuting Hu","Liang Zheng","Yi Yang","Yongfeng Huang"],"abstract":"This paper contributes a new large-scale dataset for weakly supervised\ncross-media retrieval, named Twitter100k. Current datasets, such as Wikipedia,\nNUS Wide and Flickr30k, have two major limitations. First, these datasets are\nlacking in content diversity, i.e., only some pre-defined classes are covered.\nSecond, texts in these datasets are written in well-organized language, leading\nto inconsistency with realistic applications. To overcome these drawbacks, the\nproposed Twitter100k dataset is characterized by two aspects: 1) it has 100,000\nimage-text pairs randomly crawled from Twitter and thus has no constraint in\nthe image categories; 2) text in Twitter100k is written in informal language by\nthe users.\n  Since strongly supervised methods leverage the class labels that may be\nmissing in practice, this paper focuses on weakly supervised learning for\ncross-media retrieval, in which only text-image pairs are exploited during\ntraining. We extensively benchmark the performance of four subspace learning\nmethods and three variants of the Correspondence AutoEncoder, along with\nvarious text features on Wikipedia, Flickr30k and Twitter100k. Novel insights\nare provided. As a minor contribution, inspired by the characteristic of\nTwitter100k, we propose an OCR-based cross-media retrieval method. In\nexperiment, we show that the proposed OCR-based method improves the baseline\nperformance.","url_abs":"http://arxiv.org/abs/1703.06618v1","url_pdf":"http://arxiv.org/pdf/1703.06618v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"optical-character-recognition","task_name":"Optical Character Recognition (OCR)"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"weakly-supervised-learning","task_name":"Weakly-supervised Learning"}],"methods":[],"datasets_introduced":[{"slug":"twitter100k","name":"Twitter100k","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.06618","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}