{"url":"/method/esp","slug":"esp","name":"ESP","full_name":"Efficient Spatial Pyramid","full_name_withheld":false,"description_markdown":"An **Efficient Spatial Pyramid (ESP)** is an image model block based on a factorization principle that decomposes a standard [convolution](https://paperswithcode.com/method/convolution) into two steps: (1) point-wise convolutions and (2) spatial pyramid of dilated convolutions. The point-wise convolutions help in reducing the computation, while the spatial pyramid of dilated convolutions re-samples the feature maps to learn the representations from large effective receptive field. This allows for increased efficiency compared to another image blocks like [ResNeXt](https://paperswithcode.com/method/resnext) blocks and Inception modules.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1803.06815v3","title":"ESPNet: Efficient Spatial Pyramid of Dilated Convolutions for Semantic Segmentation","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/sacmehta/ESPNet/blob/afe71c38edaee3514ca44e0adcafdf36109bf437/train/Model.py#L166","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Image Model Blocks","url":"/methods/category/image-model-blocks","pwc_aliases":[]}],"n_papers_tagged":46,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning","date":"2025-05-31","arxiv_id":"2506.00338","n_code_links":0,"syntology":null},{"paper":"/paper/self-paced-learning-strategy-with-easy-sample","title":"Self-Paced Learning Strategy with Easy Sample Prior Based on Confidence for the Flying Bird Object Detection Model Training","date":"2024-12-09","arxiv_id":"2412.06306","n_code_links":1,"syntology":null},{"paper":null,"title":"HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction","date":"2024-11-02","arxiv_id":"2411.01139","n_code_links":0,"syntology":null},{"paper":null,"title":"InstructBioMol: Advancing Biomolecule Understanding and Design Following Human Instructions","date":"2024-10-10","arxiv_id":"2410.07919","n_code_links":0,"syntology":null},{"paper":"/paper/espnet-codec-comprehensive-training-and","title":"ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech","date":"2024-09-24","arxiv_id":"2409.15897","n_code_links":2,"syntology":null},{"paper":null,"title":"Coherence influx is indispensable for quantum reservoir computing","date":"2024-09-19","arxiv_id":"2409.12693","n_code_links":0,"syntology":null},{"paper":null,"title":"ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration","date":"2024-09-14","arxiv_id":"2409.09506","n_code_links":0,"syntology":null},{"paper":null,"title":"The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization","date":"2024-07-23","arxiv_id":"2407.16447","n_code_links":0,"syntology":null},{"paper":null,"title":"ESP: Extro-Spective Prediction for Long-term Behavior Reasoning in Emergency Scenarios","date":"2024-05-07","arxiv_id":"2405.04100","n_code_links":0,"syntology":null},{"paper":null,"title":"A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system","date":"2024-04-29","arxiv_id":"2406.02563","n_code_links":0,"syntology":null},{"paper":"/paper/loongserve-efficiently-serving-long-context","title":"LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence Parallelism","date":"2024-04-15","arxiv_id":"2404.09526","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Extending echo state property for quantum reservoir computing","date":"2024-03-05","arxiv_id":"2403.02686","n_code_links":0,"syntology":null},{"paper":null,"title":"Synthesizing Environment-Specific People in Photographs","date":"2023-12-22","arxiv_id":"2312.14579","n_code_links":0,"syntology":null},{"paper":null,"title":"Learning characteristic parameters and dynamics of centrifugal pumps under multiphase flow using physics-informed neural networks","date":"2023-10-04","arxiv_id":"2310.03001","n_code_links":0,"syntology":null},{"paper":null,"title":"Edge of stability echo state networks","date":"2023-08-05","arxiv_id":"2308.02902","n_code_links":0,"syntology":null},{"paper":"/paper/fusing-pre-trained-language-models-with","title":"Fusing Pre-Trained Language Models With Multimodal Prompts Through Reinforcement Learning","date":"2023-01-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/euro-espnet-unsupervised-asr-open-source","title":"EURO: ESPnet Unsupervised ASR Open-source Toolkit","date":"2022-11-30","arxiv_id":"2211.17196","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":"/paper/espnet-se-speech-enhancement-for-robust","title":"ESPnet-SE++: Speech Enhancement for Robust Speech Recognition, Translation, and Understanding","date":"2022-07-19","arxiv_id":"2207.09514","n_code_links":1,"syntology":null},{"paper":null,"title":"TALCS: An Open-Source Mandarin-English Code-Switching Corpus and a Speech Recognition Baseline","date":"2022-06-27","arxiv_id":"2206.13135","n_code_links":0,"syntology":null},{"paper":null,"title":"NatiQ: An End-to-end Text-to-Speech System for Arabic","date":"2022-06-15","arxiv_id":"2206.07373","n_code_links":0,"syntology":null},{"paper":"/paper/multimodal-knowledge-alignment-with","title":"Multimodal Knowledge Alignment with Reinforcement Learning","date":"2022-05-25","arxiv_id":"2205.12630","n_code_links":1,"syntology":null},{"paper":null,"title":"Enveloped Sinusoid Parseval Frames","date":"2022-04-18","arxiv_id":"2204.08418","n_code_links":0,"syntology":null},{"paper":"/paper/eend-ss-joint-end-to-end-neural-speaker","title":"EEND-SS: Joint End-to-End Neural Speaker Diarization and Speech Separation for Flexible Number of Speakers","date":"2022-03-31","arxiv_id":"2203.17068","n_code_links":1,"syntology":null},{"paper":null,"title":"Ensemble Spectral Prediction (ESP) Model for Metabolite Annotation","date":"2022-03-25","arxiv_id":"2203.13783","n_code_links":0,"syntology":null},{"paper":null,"title":"A Study of Transducer based End-to-End ASR with ESPnet: Architecture, Auxiliary Loss and Decoding Strategies","date":"2022-01-14","arxiv_id":"2201.05420","n_code_links":0,"syntology":null},{"paper":null,"title":"An Investigation of the Impact of COVID-19 Non-Pharmaceutical Interventions and Economic Support Policies on Foreign Exchange Markets with Explainable AI Techniques","date":"2021-11-02","arxiv_id":"2111.14620","n_code_links":0,"syntology":null},{"paper":null,"title":"An Exploration of Self-Supervised Pretrained Representations for End-to-End Speech Recognition","date":"2021-10-09","arxiv_id":"2110.04590","n_code_links":0,"syntology":null},{"paper":"/paper/daain-detection-of-anomalous-and-adversarial","title":"DAAIN: Detection of Anomalous and Adversarial Input using Normalizing Flows","date":"2021-05-30","arxiv_id":"2105.14638","n_code_links":1,"syntology":null},{"paper":null,"title":"Event-driven timeseries analysis and the comparison of public reactions on COVID-19","date":"2021-04-30","arxiv_id":"2104.14777","n_code_links":0,"syntology":null},{"paper":null,"title":"User-friendly automatic transcription of low-resource languages: Plugging ESPnet into Elpis","date":"2020-12-15","arxiv_id":"2101.03027","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/speech-recognition","name":"Speech Recognition","papers":12},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":11},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":9},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":8},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":5},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":4},{"task":"/task/reinforcement-learning-2","name":"reinforcement-learning","papers":4},{"task":null,"name":"GPU","papers":3},{"task":"/task/language-modeling","name":"Language Modeling","papers":3},{"task":"/task/language-modelling","name":"Language Modelling","papers":3},{"task":"/task/reinforcement-learning-1","name":"Reinforcement Learning (RL)","papers":3},{"task":"/task/segmentation","name":"Segmentation","papers":3},{"task":"/task/text-to-speech","name":"Text to Speech","papers":3},{"task":"/task/text-to-speech-1","name":"text-to-speech","papers":3},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/object-detection","name":"Object Detection","papers":2},{"task":"/task/prediction","name":"Prediction","papers":2},{"task":"/task/real-time-semantic-segmentation","name":"Real-Time Semantic Segmentation","papers":2},{"task":"/task/speech-separation","name":"Speech Separation","papers":2},{"task":"/task/object-detection-1","name":"object-detection","papers":2}],"tasks_shown":20,"n_tasks":57,"usage_by_year":[{"year":"2018","papers":5},{"year":"2019","papers":3},{"year":"2020","papers":9},{"year":"2021","papers":4},{"year":"2022","papers":9},{"year":"2023","papers":4},{"year":"2024","papers":11},{"year":"2025","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/esp"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}