{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evrt-detr-the-surprising-effectiveness-of","title":"EvRT-DETR: Latent Space Adaptation of Image Detectors for Event-based Vision","arxiv_id":"2412.02890","date":"2024-12-03","proceeding":null,"authors":["Dmitrii Torbunov","Yihui Ren","Animesh Ghose","Odera Dim","Yonggang Cui"],"abstract":"Event-based cameras (EBCs) have emerged as a bio-inspired alternative to traditional cameras, offering advantages in power efficiency, temporal resolution, and high dynamic range. However, the development of image analysis methods for EBCs is challenging due to the sparse and asynchronous nature of the data. This work addresses the problem of object detection for EBC cameras. The current approaches to EBC object detection focus on constructing complex data representations and rely on specialized architectures. We introduce I2EvDet (Image-to-Event Detection), a novel adaptation framework that bridges mainstream object detection with temporal event data processing. First, we demonstrate that a Real-Time DEtection TRansformer, or RT-DETR, a state-of-the-art natural image detector, trained on a simple image-like representation of the EBC data achieves performance comparable to specialized EBC methods. Next, as part of our framework, we develop an efficient adaptation technique that transforms image-based detectors into event-based detection models by modifying their frozen latent representation space through minimal architectural additions. The resulting EvRT-DETR model reaches state-of-the-art performance on the standard benchmark datasets Gen1 (mAP $+2.3$) and 1Mpx/Gen4 (mAP $+1.4$). These results demonstrate a fundamentally new approach to EBC object detection through principled adaptation of mainstream architectures, offering an efficient alternative with potential applications to other temporal visual domains. The code is available at: https://github.com/realtime-intelligence/evrt-detr","url_abs":"https://arxiv.org/abs/2412.02890v2","url_pdf":"https://arxiv.org/pdf/2412.02890v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evrt-detr-the-surprising-effectiveness-of","repo_url":"https://github.com/realtime-intelligence/evrt-detr","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"event-detection","task_name":"Event Detection"},{"task_slug":"event-based-vision","task_name":"Event-based vision"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"ebc","method_name":"EBC"},{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2412.02890","atlas_url":"https://app.syntology.ai/?focus=2412.02890","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2412.02890"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/realtime-intelligence/evrt-detr","reach":null}],"summary":{"ran_draft_wrong":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"7108f9b0037300c0","entry":"construct_train_frame_jitter_transform","repo":"realtime-intelligence/evrt-detr","repo_kind":"official","path":"scripts/train/gen1/video_detection_evrtdetr/train_gen1_evrtdetr_presnet18.py","file_url":"https://github.com/realtime-intelligence/evrt-detr/blob/HEAD/scripts/train/gen1/video_detection_evrtdetr/train_gen1_evrtdetr_presnet18.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"7108f9b0037300c0"}},{"code_sha256_prefix":"7b5df82dbf4f9982","entry":"construct_train_transform","repo":"realtime-intelligence/evrt-detr","repo_kind":"official","path":"scripts/train/gen1/frame_detection_rtdetr/train_gen1_rtdetr_presnet18.py","file_url":"https://github.com/realtime-intelligence/evrt-detr/blob/HEAD/scripts/train/gen1/frame_detection_rtdetr/train_gen1_rtdetr_presnet18.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"7b5df82dbf4f9982"}},{"code_sha256_prefix":"d311daa4051256b3","entry":"construct_train_video_transform","repo":"realtime-intelligence/evrt-detr","repo_kind":"official","path":"scripts/train/gen1/video_detection_evrtdetr/train_gen1_evrtdetr_presnet18.py","file_url":"https://github.com/realtime-intelligence/evrt-detr/blob/HEAD/scripts/train/gen1/video_detection_evrtdetr/train_gen1_evrtdetr_presnet18.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"d311daa4051256b3"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}