{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/d3tw-discriminative-differentiable-dynamic","title":"D3TW: Discriminative Differentiable Dynamic Time Warping for Weakly Supervised Action Alignment and Segmentation","arxiv_id":"1901.02598","date":"2019-01-09","proceeding":"CVPR 2019 6","authors":["Chien-Yi Chang","De-An Huang","Yanan Sui","Li Fei-Fei","Juan Carlos Niebles"],"abstract":"We address weakly supervised action alignment and segmentation in videos,\nwhere only the order of occurring actions is available during training. We\npropose Discriminative Differentiable Dynamic Time Warping (D3TW), the first\ndiscriminative model using weak ordering supervision. The key technical\nchallenge for discriminative modeling with weak supervision is that the loss\nfunction of the ordering supervision is usually formulated using dynamic\nprogramming and is thus not differentiable. We address this challenge with a\ncontinuous relaxation of the min-operator in dynamic programming and extend the\nalignment loss to be differentiable. The proposed D3TW innovatively solves\nsequence alignment with discriminative modeling and end-to-end training, which\nsubstantially improves the performance in weakly supervised action alignment\nand segmentation tasks. We show that our model is able to bypass the\ndegenerated sequence problem usually encountered in previous work and\noutperform the current state-of-the-art across three evaluation metrics in two\nchallenging datasets.","url_abs":"http://arxiv.org/abs/1901.02598v2","url_pdf":"http://arxiv.org/pdf/1901.02598v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"dynamic-time-warping","task_name":"Dynamic Time Warping"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"weakly-supervised-action-segmentation","task_name":"Weakly Supervised Action Segmentation (Transcript)"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/weakly-supervised-action-segmentation","task":"Weakly Supervised Action Segmentation (Transcript)","dataset":"Breakfast","model":"D3TW","rank_in_archive_order":6,"of":7,"metrics":{"Acc":"45.7"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1901.02598","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}