{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/encoding-video-and-label-priors-for-multi","title":"Encoding Video and Label Priors for Multi-label Video Classification on YouTube-8M dataset","arxiv_id":"1706.07960","date":"2017-06-24","proceeding":null,"authors":["Seil Na","Youngjae Yu","Sang-ho Lee","Ji-Sung Kim","Gunhee Kim"],"abstract":"YouTube-8M is the largest video dataset for multi-label video classification.\nIn order to tackle the multi-label classification on this challenging dataset,\nit is necessary to solve several issues such as temporal modeling of videos,\nlabel imbalances, and correlations between labels. We develop a deep neural\nnetwork model, which consists of four components: the frame encoder, the\nclassification layer, the label processing layer, and the loss function. We\nintroduce our newly proposed methods and discusses how existing models operate\nin the YouTube-8M Classification Task, what insights they have, and why they\nsucceed (or fail) to achieve good performance. Most of the models we proposed\nare very high compared to the baseline models, and the ensemble of the models\nwe used is 8th in the Kaggle Competition.","url_abs":"http://arxiv.org/abs/1706.07960v2","url_pdf":"http://arxiv.org/pdf/1706.07960v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"encoding-video-and-label-priors-for-multi","repo_url":"https://github.com/seilna/youtube-8m","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1706.07960","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}