{"url":"/method/3d-resnet-rs","slug":"3d-resnet-rs","name":"3D ResNet-RS","full_name":"3D ResNet-RS","full_name_withheld":false,"description_markdown":"**3D ResNet-RS** is an architecture and scaling strategy for 3D ResNets for video recognition. The key additions are:\r\n\r\n- **3D ResNet-D stem**: The [ResNet-D](https://paperswithcode.com/method/resnet-d) stem is adapted to 3D inputs by using three consecutive [3D convolutional layers](https://paperswithcode.com/method/3d-convolution). The first convolutional layer employs a temporal kernel size of 5 while the remaining two convolutional layers employ a temporal kernel size of 1.\r\n\r\n- **3D Squeeze-and-Excitation**:  [Squeeze-and-Excite](https://paperswithcode.com/method/squeeze-and-excitation-block) is adapted to spatio-temporal inputs by using a 3D [global average pooling](https://paperswithcode.com/method/global-average-pooling) operation for the squeeze operation. A SE ratio of 0.25 is applied in each 3D bottleneck block for all experiments.\r\n\r\n- **Self-gating**: A self-gating module is used in each 3D bottleneck block after the SE module.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Revisiting 3D ResNets for Video Recognition","paper":"/paper/revisiting-3d-resnets-for-video-recognition","first_author":"Xianzhi Du","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/revisiting-3d-resnets-for-video-recognition"},"source":{"url":"https://arxiv.org/abs/2109.01696v1","title":"Revisiting 3D ResNets for Video Recognition","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Video Recognition Models","url":"/methods/category/video-recognition-models","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":3,"papers_newest_first":[{"paper":"/paper/revisiting-3d-resnets-for-video-recognition","title":"Revisiting 3D ResNets for Video Recognition","date":"2021-09-03","arxiv_id":"2109.01696","n_code_links":5,"syntology":null},{"paper":null,"title":"Prediction of Chronic Kidney Disease Using Deep Neural Network","date":"2020-12-22","arxiv_id":"2012.12089","n_code_links":0,"syntology":null},{"paper":"/paper/w-net-a-cnn-based-architecture-for-white","title":"W-Net: A CNN-based Architecture for White Blood Cells Image Classification","date":"2019-10-02","arxiv_id":"1910.01091","n_code_links":1,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/action-classification","name":"Action Classification","papers":1},{"task":"/task/classification-1","name":"Classification","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/classification","name":"General Classification","papers":1},{"task":"/task/image-classification","name":"Image Classification","papers":1},{"task":"/task/video-recognition","name":"Video Recognition","papers":1},{"task":"/task/image-classification","name":"image-classification","papers":1}],"tasks_shown":7,"n_tasks":7,"usage_by_year":[{"year":"2019","papers":1},{"year":"2020","papers":1},{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/3d-resnet-rs"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}