{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/straight-to-shapes-real-time-detection-of","title":"Straight to Shapes: Real-time Detection of Encoded Shapes","arxiv_id":"1611.07932","date":"2016-11-23","proceeding":"CVPR 2017 7","authors":["Saumya Jetley","Michael Sapienza","Stuart Golodetz","Philip H. S. Torr"],"abstract":"Current object detection approaches predict bounding boxes, but these provide\nlittle instance-specific information beyond location, scale and aspect ratio.\nIn this work, we propose to directly regress to objects' shapes in addition to\ntheir bounding boxes and categories. It is crucial to find an appropriate shape\nrepresentation that is compact and decodable, and in which objects can be\ncompared for higher-order concepts such as view similarity, pose variation and\nocclusion. To achieve this, we use a denoising convolutional auto-encoder to\nestablish an embedding space, and place the decoder after a fast end-to-end\nnetwork trained to regress directly to the encoded shape vectors. This yields\nwhat to the best of our knowledge is the first real-time shape prediction\nnetwork, running at ~35 FPS on a high-end desktop. With higher-order shape\nreasoning well-integrated into the network pipeline, the network shows the\nuseful practical quality of generalising to unseen categories similar to the\nones in the training set, something that most existing approaches fail to\nhandle.","url_abs":"http://arxiv.org/abs/1611.07932v2","url_pdf":"http://arxiv.org/pdf/1611.07932v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"straight-to-shapes-real-time-detection-of","repo_url":"https://github.com/torrvision/straighttoshapes","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1611.07932","atlas_url":"https://app.syntology.ai/?focus=1611.07932","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}