{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cross-modal-and-hierarchical-modeling-of","title":"Cross-Modal and Hierarchical Modeling of Video and Text","arxiv_id":"1810.07212","date":"2018-10-16","proceeding":"ECCV 2018 9","authors":["Bowen Zhang","Hexiang Hu","Fei Sha"],"abstract":"Visual data and text data are composed of information at multiple\ngranularities. A video can describe a complex scene that is composed of\nmultiple clips or shots, where each depicts a semantically coherent event or\naction. Similarly, a paragraph may contain sentences with different topics,\nwhich collectively conveys a coherent message or story. In this paper, we\ninvestigate the modeling techniques for such hierarchical sequential data where\nthere are correspondences across multiple modalities. Specifically, we\nintroduce hierarchical sequence embedding (HSE), a generic model for embedding\nsequential data of different modalities into hierarchically semantic spaces,\nwith either explicit or implicit correspondence information. We perform\nempirical studies on large-scale video and paragraph retrieval datasets and\ndemonstrated superior performance by the proposed methods. Furthermore, we\nexamine the effectiveness of our learned embeddings when applied to downstream\ntasks. We show its utility in zero-shot action recognition and video\ncaptioning.","url_abs":"http://arxiv.org/abs/1810.07212v1","url_pdf":"http://arxiv.org/pdf/1810.07212v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cross-modal-and-hierarchical-modeling-of","repo_url":"https://github.com/Sha-Lab/CMHSE","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"zero-shot-action-recognition","task_name":"Zero-Shot Action Recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.07212","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}