Browse State-of-the-Art › Camera Pose Estimation
Camera Pose Estimation
134 papers with code · 1 benchmark · 3 datasets archive 2025-07-28
Camera pose estimation is a crucial task in computer vision and robotics that involves determining the position and orientation (pose) of a camera relative to a given reference frame. This task is essential for various applications, such as augmented reality, 3D reconstruction, SLAM, and autonomous navigation.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| KITTI Odometry Benchmark (7 rows) | Manydepth2 | Manydepth2: Motion-Aware Self-Supervised Multi-Frame Monocular... | code | Syntology ran 0 of 9 samples · 9 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 134 papers with code (304 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Nov 2019 19 repositories listed Syntology ran 6 of 22 samples · 16 unverified · 6 pointer-only (licence)This paper introduces SuperGlue, a neural network that matches two sets of local features by jointly finding correspondences and rejecting non-matchable points.
-
4 Jun 2018 15 repositories listed Syntology ran 17 of 24 samples · 7 unverified · 6 pointer-only (licence)Per-pixel ground-truth depth data is challenging to acquire at scale.
-
1 Jun 2018 5 repositories listedObjects can provide long-range geometric and scale constraints to improve camera pose estimation and reduce monocular drift.
-
1 Apr 2021 4 repositories listed Syntology ran 14 of 16 samples · 2 unverifiedWe present a novel method for local image feature matching.
-
30 Jul 2020 4 repositories listedWe present a solution to the problem of visual odometry from the data acquired by a stereo event-based camera rig.
-
19 Oct 2018 4 repositories listedThis paper addresses the challenge of dense pixel correspondence estimation between two images.
-
22 Apr 2019 3 repositories listedThis paper proposes a robust localization system that employs deep learning for better scene representation, and enhances the accuracy of 6-DOF camera pose estimation.
-
6 Mar 2018 3 repositories listed Syntology ran 0 of 8 samples · 8 unverifiedWe propose GeoNet, a jointly unsupervised learning framework for monocular depth, optical flow and ego-motion estimation from videos.
-
4 Dec 2023 2 repositories listedDense simultaneous localization and mapping (SLAM) is crucial for robotics and augmented reality applications.
-
15 Oct 2023 2 repositories listedPerspective-n-Point (P$n$P) stands as a fundamental algorithm for pose estimation in various applications.
-
23 Jun 2023 2 repositories listed Syntology ran 18 of 40 samples · 22 unverifiedWe introduce LightGlue, a deep neural network that learns to match local features across images.
-
17 Apr 2023 2 repositories listedPurpose: Surgical scene understanding plays a critical role in the technology stack of tomorrow's intervention-assisting systems in endoscopic surgeries.
-
12 Apr 2023 2 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedKeypoint detection & descriptors are foundational tech-nologies for computer vision tasks like image matching, 3D reconstruction and visual odometry.
-
6 Dec 2021 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The reprojection loss is then proposed to directly optimize these sub-pixel keypoints, and the dispersity peak loss is presented for accurate keypoints regularization.
-
16 Mar 2021 2 repositories listed Syntology ran 7 of 10 samples · 3 unverified · 10 pointer-only (licence)In this paper, we go Back to the Feature: we argue that deep networks should focus on learning robust and invariant visual features, while the geometric estimation should be left to principled algorithms.
-
4 Mar 2021 2 repositories listedWe present self-supervised geometric perception (SGP), the first general framework to learn a feature descriptor for correspondence matching without any ground-truth geometric model labels (e.
-
3 Jan 2021 2 repositories listed Syntology ran 6 of 11 samples · 5 unverified · 5 pointer-only (licence)Correspondence selection aims to correctly select the consistent matches (inliers) from an initial set of putative correspondences.
-
5 Dec 2019 2 repositories listed Syntology ran 0 of 14 samples · 14 unverifiedIn contrast, generic camera models allow for very accurate calibration due to their flexibility.
-
28 Aug 2019 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)To the best of our knowledge, this is the first work to show that deep networks trained using unlabelled monocular videos can predict globally scale-consistent camera trajectories over a long video sequence.
-
28 Jul 2017 2 repositories listedVisual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds.
-
25 Apr 2017 2 repositories listedWe present an unsupervised learning framework for the task of monocular depth and camera motion estimation from unstructured video sequences.
-
16 Jul 2025 1 repository listedWe present SpatialTrackerV2, a feed-forward 3D point tracking method for monocular videos.
-
10 Jun 2025 1 repository listedWe further propose a new scene scale-aware evaluation metric for SLAM based on the the optical flow induced by the camera pose estimation error.
-
19 May 2025 1 repository listedIn the first stage, we learn to reconstruct the scene implicitly in a latent space without relying on any explicit 3D representation.
-
21 Apr 2025 1 repository listedMulti-view understanding, the ability to reconcile visual information across diverse viewpoints for effective navigation, manipulation, and 3D scene comprehension, is a fundamental challenge in Multi-Modal Large…
-
31 Mar 2025 1 repository listedRecent advances in DUSt3R have enabled robust estimation of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets.
-
27 Mar 2025 1 repository listed Syntology ran 1 of 18 samples · 17 unverifiedThis paper presents a unified approach to understanding dynamic scenes from casual videos.
-
10 Mar 2025 1 repository listedIn this paper, we present a novel framework designed to reconstruct long-sequence 3D human motion in the world coordinates from in-the-wild videos with multiple shot transitions.
-
28 Feb 2025 1 repository listedOur framework follows a paradigm based on the refinement of an initial pose.
-
23 Jan 2025 1 repository listedMulti-view 3D reconstruction remains a core challenge in computer vision, particularly in applications requiring accurate and scalable representations across diverse perspectives.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections