{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-grained-apparel-classification-and","title":"Fine-grained Apparel Classification and Retrieval without rich annotations","arxiv_id":"1811.02385","date":"2018-11-06","proceeding":null,"authors":["Aniket Bhatnagar","Sanchit Aggarwal"],"abstract":"The ability to correctly classify and retrieve apparel images has a variety\nof applications important to e-commerce, online advertising and internet\nsearch. In this work, we propose a robust framework for fine-grained apparel\nclassification, in-shop and cross-domain retrieval which eliminates the\nrequirement of rich annotations like bounding boxes and human-joints or\nclothing landmarks, and training of bounding box/ key-landmark detector for the\nsame. Factors such as subtle appearance differences, variations in human poses,\ndifferent shooting angles, apparel deformations, and self-occlusion add to the\nchallenges in classification and retrieval of apparel items. Cross-domain\nretrieval is even harder due to the presence of large variation between online\nshopping images, usually taken in ideal lighting, pose, positive angle and\nclean background as compared with street photos captured by users in\ncomplicated conditions with poor lighting and cluttered scenes. Our framework\nuses compact bilinear CNN with tensor sketch algorithm to generate embeddings\nthat capture local pairwise feature interactions in a translationally invariant\nmanner. For apparel classification, we pass the feature embeddings through a\nsoftmax classifier, while, the in-shop and cross-domain retrieval pipelines use\na triplet-loss based optimization approach, such that squared Euclidean\ndistance between embeddings measures the dissimilarity between the images.\nUnlike previous works that relied on bounding box, key clothing landmarks or\nhuman joint detectors to assist the final deep classifier, proposed framework\ncan be trained directly on the provided category labels or generated triplets\nfor triplet loss optimization. Lastly, Experimental results on the DeepFashion\nfine-grained categorization, and in-shop and consumer-to-shop retrieval\ndatasets provide a comparative analysis with previous work performed in the\ndomain.","url_abs":"http://arxiv.org/abs/1811.02385v1","url_pdf":"http://arxiv.org/pdf/1811.02385v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fine-grained-apparel-classification-and","repo_url":"https://github.com/aniket03/keras_compact_bilnear_CNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"fine-grained-apparel-classification-and","repo_url":"https://github.com/sunn-e/DeepFashion-retrieval-2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":null,"task_name":"Triplet"}],"methods":[{"method_slug":"triplet-loss","method_name":"Triplet Loss"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}