{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-mixed-classification-regression-framework","title":"A Mixed Classification-Regression Framework for 3D Pose Estimation from 2D Images","arxiv_id":"1805.03225","date":"2018-05-08","proceeding":null,"authors":["Siddharth Mahendran","Haider Ali","Rene Vidal"],"abstract":"3D pose estimation from a single 2D image is an important and challenging\ntask in computer vision with applications in autonomous driving, robot\nmanipulation and augmented reality. Since 3D pose is a continuous quantity, a\nnatural formulation for this task is to solve a pose regression problem.\nHowever, since pose regression methods return a single estimate of the pose,\nthey have difficulties handling multimodal pose distributions (e.g. in the case\nof symmetric objects). An alternative formulation, which can capture multimodal\npose distributions, is to discretize the pose space into bins and solve a pose\nclassification problem. However, pose classification methods can give large\npose estimation errors depending on the coarseness of the discretization. In\nthis paper, we propose a mixed classification-regression framework that uses a\nclassification network to produce a discrete multimodal pose estimate and a\nregression network to produce a continuous refinement of the discrete estimate.\nThe proposed framework can accommodate different architectures and loss\nfunctions, leading to multiple classification-regression models, some of which\nachieve state-of-the-art performance on the challenging Pascal3D+ dataset.","url_abs":"http://arxiv.org/abs/1805.03225v1","url_pdf":"http://arxiv.org/pdf/1805.03225v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-mixed-classification-regression-framework","repo_url":"https://github.com/JHUVisionLab/multi-modal-regression","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"3d-pose-estimation","task_name":"3D Pose Estimation"},{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"pose-estimation","task_name":"Pose Estimation"},{"task_slug":"robot-manipulation","task_name":"Robot Manipulation"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.03225","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.03225"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/JHUVisionLab/multi-modal-regression","reach":null}],"summary":{"ran_violates":1},"by_repo_kind":{"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a8aa7c896fd915b5","entry":"my_schedule","repo":"JHUVisionLab/multi-modal-regression","repo_kind":"listed","path":"learnJointCatPoseModel_top1.py","file_url":"https://github.com/JHUVisionLab/multi-modal-regression/blob/HEAD/learnJointCatPoseModel_top1.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a8aa7c896fd915b5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}