{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/weakly-supervised-visual-instrument-playing","title":"Weakly-supervised Visual Instrument-playing Action Detection in Videos","arxiv_id":"1805.02031","date":"2018-05-05","proceeding":null,"authors":["Jen-Yu Liu","Yi-Hsuan Yang","Shyh-Kang Jeng"],"abstract":"Instrument playing is among the most common scenes in music-related videos,\nwhich represent nowadays one of the largest sources of online videos. In order\nto understand the instrument-playing scenes in the videos, it is important to\nknow what instruments are played, when they are played, and where the playing\nactions occur in the scene. While audio-based recognition of instruments has\nbeen widely studied, the visual aspect of the music instrument playing remains\nlargely unaddressed in the literature. One of the main obstacles is the\ndifficulty in collecting annotated data of the action locations for\ntraining-based methods. To address this issue, we propose a weakly-supervised\nframework to find when and where the instruments are played in the videos. We\npropose to use two auxiliary models, a sound model and an object model, to\nprovide supervisions for training the instrument-playing action model. The\nsound model provides temporal supervisions, while the object model provides\nspatial supervisions. They together can simultaneously provide temporal and\nspatial supervisions. The resulted model only needs to analyze the visual part\nof a music video to deduce which, when and where instruments are played. We\nfound that the proposed method significantly improves the localization\naccuracy. We evaluate the result of the proposed method temporally and\nspatially on a small dataset (totally 5,400 frames) that we manually annotated.","url_abs":"http://arxiv.org/abs/1805.02031v1","url_pdf":"http://arxiv.org/pdf/1805.02031v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"weakly-supervised-visual-instrument-playing","repo_url":"https://github.com/ciaua/InstrumentPlayingDetection","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}