{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/event-based-video-reconstruction-using","title":"Event-Based Video Reconstruction Using Transformer","arxiv_id":null,"date":"2021-01-01","proceeding":"ICCV 2021 10","authors":["Wenming Weng","Yueyi Zhang","Zhiwei Xiong"],"abstract":"    Event cameras, which output events by detecting spatio-temporal brightness changes, bring a novel paradigm to image sensors with high dynamic range and low latency. Previous works have achieved impressive performances on event-based video reconstruction by introducing convolutional neural networks (CNNs). However, intrinsic locality of convolutional operations is not capable of modeling long-range dependency, which is crucial to many vision tasks. In this paper, we present a hybrid CNN-Transformer network for event-based video reconstruction (ET-Net), which merits the fine local information from CNN and global contexts from Transformer. In addition, we further propose a Token Pyramid Aggregation strategy to implement multi-scale token integration for relating internal and intersected semantic concepts in the token-space. Experimental results demonstrate that our proposed method achieves superior performance over state-of-the-art methods on multiple real-world event datasets. The code is available at https://github.com/WarranWeng/ET-Net    ","url_abs":"http://openaccess.thecvf.com//content/ICCV2021/html/Weng_Event-Based_Video_Reconstruction_Using_Transformer_ICCV_2021_paper.html","url_pdf":"http://openaccess.thecvf.com//content/ICCV2021/papers/Weng_Event-Based_Video_Reconstruction_Using_Transformer_ICCV_2021_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"event-based-video-reconstruction-using","repo_url":"https://github.com/warranweng/et-net","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"event-based-video-reconstruction","task_name":"Event-Based Video Reconstruction"},{"task_slug":"event-based-object-segmentation","task_name":"Event-based Object Segmentation"},{"task_slug":"video-reconstruction","task_name":"Video Reconstruction"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/event-based-object-segmentation-on-ddd17-seg","task":"Event-based Object Segmentation","dataset":"DDD17-SEG","model":"ETNet","rank_in_archive_order":2,"of":2,"metrics":{"mIoU":"0.34"},"uses_additional_data":false},{"leaderboard":"/sota/event-based-object-segmentation-on-dsec-seg","task":"Event-based Object Segmentation","dataset":"DSEC-SEG","model":"ETNet","rank_in_archive_order":2,"of":2,"metrics":{"mIoU":"0.36"},"uses_additional_data":false},{"leaderboard":"/sota/event-based-object-segmentation-on-mvsec-seg","task":"Event-based Object Segmentation","dataset":"MVSEC-SEG","model":"ETNet","rank_in_archive_order":2,"of":8,"metrics":{"mIoU":"0.37"},"uses_additional_data":false},{"leaderboard":"/sota/event-based-object-segmentation-on-rgbe-seg","task":"Event-based Object Segmentation","dataset":"RGBE-SEG","model":"ETNet","rank_in_archive_order":2,"of":8,"metrics":{"mIoU":"0.35"},"uses_additional_data":false},{"leaderboard":"/sota/video-reconstruction-on-event-camera-dataset","task":"Video Reconstruction","dataset":"Event-Camera Dataset","model":"ET-Net","rank_in_archive_order":2,"of":4,"metrics":{"LPIPS":"0.224","Mean Squared Error":"0.047"},"uses_additional_data":false},{"leaderboard":"/sota/video-reconstruction-on-mvsec","task":"Video Reconstruction","dataset":"MVSEC","model":"ET-Net","rank_in_archive_order":2,"of":4,"metrics":{"LPIPS":"0.489","Mean Squared Error":"0.107"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}