{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/youtube-8m-a-large-scale-video-classification","title":"YouTube-8M: A Large-Scale Video Classification Benchmark","arxiv_id":"1609.08675","date":"2016-09-27","proceeding":null,"authors":["Sami Abu-El-Haija","Nisarg Kothari","Joonseok Lee","Paul Natsev","George Toderici","Balakrishnan Varadarajan","Sudheendra Vijayanarasimhan"],"abstract":"Many recent advancements in Computer Vision are attributed to large datasets.\nOpen-source software packages for Machine Learning and inexpensive commodity\nhardware have reduced the barrier of entry for exploring novel approaches at\nscale. It is possible to train models over millions of examples within a few\ndays. Although large-scale datasets exist for image understanding, such as\nImageNet, there are no comparable size video classification datasets.\n  In this paper, we introduce YouTube-8M, the largest multi-label video\nclassification dataset, composed of ~8 million videos (500K hours of video),\nannotated with a vocabulary of 4800 visual entities. To get the videos and\ntheir labels, we used a YouTube video annotation system, which labels videos\nwith their main topics. While the labels are machine-generated, they have\nhigh-precision and are derived from a variety of human-based signals including\nmetadata and query click signals. We filtered the video labels (Knowledge Graph\nentities) using both automated and manual curation strategies, including asking\nhuman raters if the labels are visually recognizable. Then, we decoded each\nvideo at one-frame-per-second, and used a Deep CNN pre-trained on ImageNet to\nextract the hidden representation immediately prior to the classification\nlayer. Finally, we compressed the frame features and make both the features and\nvideo-level labels available for download.\n  We trained various (modest) classification models on the dataset, evaluated\nthem using popular evaluation metrics, and report them as baselines. Despite\nthe size of the dataset, some of our models train to convergence in less than a\nday on a single machine using TensorFlow. We plan to release code for training\na TensorFlow model and for computing metrics.","url_abs":"http://arxiv.org/abs/1609.08675v1","url_pdf":"http://arxiv.org/pdf/1609.08675v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/AKASH2907/Content-based-Video-Recommendation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/AKASH2907/Content-based-Video-Relevance-Prediction","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/Cloud-Computing-IoT/Cloud-Enabled-Smart-Speaker","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/Cloud-Computing-IoT/speakEasy","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/boseaslcohort/youtube-8m","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/google/youtube-8m","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"youtube-8m-a-large-scale-video-classification","repo_url":"https://github.com/taufikxu/youtube","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"3d-face-reconstruction","task_name":"3D Face Reconstruction"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"action-recognition-in-videos-2","task_name":"Action Recognition In Videos"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[{"slug":"youtube-8m","name":"YouTube-8M","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-activitynet-1","task":"Action Recognition In Videos","dataset":"ActivityNet","model":"LSTM + Pretrained on YT-8M","rank_in_archive_order":1,"of":1,"metrics":{"mAP":"75.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-sports-1m-1","task":"Action Recognition In Videos","dataset":"Sports-1M","model":"LSTM +Pretrained on YT-8M","rank_in_archive_order":2,"of":2,"metrics":{"Video hit@1":"65.7","Video hit@5":"86.2"},"uses_additional_data":false},{"leaderboard":"/sota/video-classification-on-youtube-8m","task":"Video Classification","dataset":"YouTube-8M","model":"Mixture-of-2-Experts","rank_in_archive_order":3,"of":3,"metrics":{"Hit@1":"70.1","Hit@5":"84.8","PERR":"29.1"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1609.08675","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1609.08675"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AKASH2907/Content-based-Video-Recommendation","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/taufikxu/youtube","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/google/youtube-8m","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/boseaslcohort/youtube-8m","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Cloud-Computing-IoT/Cloud-Enabled-Smart-Speaker","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Cloud-Computing-IoT/speakEasy","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AKASH2907/Content-based-Video-Relevance-Prediction","reach":{"status":"ok"}}],"summary":{"ran_fixture":1,"unverified":8},"by_repo_kind":{"listed":{"samples":9,"ran":1,"repositories":3}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"2ef42071ff785910","entry":"get_segments","repo":"google/youtube-8m","repo_kind":"listed","path":"inference.py","file_url":"https://github.com/google/youtube-8m/blob/HEAD/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2ef42071ff785910"}},{"code_sha256_prefix":"2f3219b9915649d3","entry":"FramePooling","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"model_utils.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/model_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"2f3219b9915649d3"}},{"code_sha256_prefix":"c7c286e6493b9cf6","entry":"SampleRandomFrames","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"model_utils.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/model_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c7c286e6493b9cf6"}},{"code_sha256_prefix":"604001c4ca602f10","entry":"SampleRandomSequence","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"model_utils.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/model_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"604001c4ca602f10"}},{"code_sha256_prefix":"427f933cdb30be37","entry":"calculate_hit_at_one","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"eval_util.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/eval_util.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"427f933cdb30be37"}},{"code_sha256_prefix":"a3d2406f6ce64e29","entry":"calculate_precision_at_equal_recall_rate","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"eval_util.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/eval_util.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a3d2406f6ce64e29"}},{"code_sha256_prefix":"557deb04a3b96763","entry":"flatten","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"eval_util.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/eval_util.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"557deb04a3b96763"}},{"code_sha256_prefix":"a8936fa247d02451","entry":"resize_axis","repo":"taufikxu/youtube","repo_kind":"listed","path":"readers.py","file_url":"https://github.com/taufikxu/youtube/blob/HEAD/readers.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"a8936fa247d02451"}},{"code_sha256_prefix":"03117d2e243f0f99","entry":"to_csv_row","repo":"boseaslcohort/youtube-8m","repo_kind":"listed","path":"convert_prediction_from_json_to_csv.py","file_url":"https://github.com/boseaslcohort/youtube-8m/blob/HEAD/convert_prediction_from_json_to_csv.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"03117d2e243f0f99"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}