Papers › Event-Based Video Reconstruction Using Transformer

Event-Based Video Reconstruction Using Transformer

1 Jan 2021ICCV 2021 10archive 2025-07-28

Wenming Weng, Yueyi Zhang, Zhiwei Xiong

Event cameras, which output events by detecting spatio-temporal brightness changes, bring a novel paradigm to image sensors with high dynamic range and low latency. Previous works have achieved impressive performances on event-based video reconstruction by introducing convolutional neural networks (CNNs). However, intrinsic locality of convolutional operations is not capable of modeling long-range dependency, which is crucial to many vision tasks. In this paper, we present a hybrid CNN-Transformer network for event-based video reconstruction (ET-Net), which merits the fine local information from CNN and global contexts from Transformer. In addition, we further propose a Token Pyramid Aggregation strategy to implement multi-scale token integration for relating internal and intersected semantic concepts in the token-space. Experimental results demonstrate that our proposed method achieves superior performance over state-of-the-art methods on multiple real-world event datasets. The code is available at https://github.com/WarranWeng/ET-Net

PaperPDFCode

Code

warranweng/et-net officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Event-Based Video ReconstructionEvent-based Object SegmentationVideo Reconstruction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Event-based Object Segmentation DDD17-SEG ETNet mIoU 0.34 #2 of 2 Archive leaderboard report
Event-based Object Segmentation DSEC-SEG ETNet mIoU 0.36 #2 of 2 Archive leaderboard report
Event-based Object Segmentation MVSEC-SEG ETNet mIoU 0.37 #2 of 8 Archive leaderboard report
Event-based Object Segmentation RGBE-SEG ETNet mIoU 0.35 #2 of 8 Archive leaderboard report
Video Reconstruction Event-Camera Dataset ET-Net LPIPS 0.224 #2 of 4 Archive leaderboard report
Video Reconstruction Event-Camera Dataset ET-Net Mean Squared Error 0.047 #2 of 4 Archive leaderboard report
Video Reconstruction MVSEC ET-Net LPIPS 0.489 #2 of 4 Archive leaderboard report
Video Reconstruction MVSEC ET-Net Mean Squared Error 0.107 #2 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections