{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/s3d-single-shot-multi-span-detector-via-fully","title":"S3D: Single Shot multi-Span Detector via Fully 3D Convolutional Networks","arxiv_id":"1807.08069","date":"2018-07-21","proceeding":null,"authors":["Da Zhang","Xiyang Dai","Xin Wang","Yuan-Fang Wang"],"abstract":"In this paper, we present a novel Single Shot multi-Span Detector for\ntemporal activity detection in long, untrimmed videos using a simple end-to-end\nfully three-dimensional convolutional (Conv3D) network. Our architecture, named\nS3D, encodes the entire video stream and discretizes the output space of\ntemporal activity spans into a set of default spans over different temporal\nlocations and scales. At prediction time, S3D predicts scores for the presence\nof activity categories in each default span and produces temporal adjustments\nrelative to the span location to predict the precise activity duration. Unlike\nmany state-of-the-art systems that require a separate proposal and\nclassification stage, our S3D is intrinsically simple and dedicatedly designed\nfor single-shot, end-to-end temporal activity detection. When evaluating on\nTHUMOS'14 detection benchmark, S3D achieves state-of-the-art performance and is\nvery efficient and can operate at 1271 FPS.","url_abs":"http://arxiv.org/abs/1807.08069v2","url_pdf":"http://arxiv.org/pdf/1807.08069v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"s3d-single-shot-multi-span-detector-via-fully","repo_url":"https://github.com/dazhang-cv/S3D","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"activity-detection","task_name":"Activity Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.08069","atlas_url":"https://app.syntology.ai/?focus=1807.08069","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}