{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rethinking-image-super-resolution-from","title":"Rethinking Image Super-Resolution from Training Data Perspectives","arxiv_id":"2409.00768","date":"2024-09-01","proceeding":null,"authors":["Go Ohtani","Ryu Tadokoro","Ryosuke Yamada","Yuki M. Asano","Iro Laina","Christian Rupprecht","Nakamasa Inoue","Rio Yokota","Hirokatsu Kataoka","Yoshimitsu Aoki"],"abstract":"In this work, we investigate the understudied effect of the training data used for image super-resolution (SR). Most commonly, novel SR methods are developed and benchmarked on common training datasets such as DIV2K and DF2K. However, we investigate and rethink the training data from the perspectives of diversity and quality, {thereby addressing the question of ``How important is SR training for SR models?''}. To this end, we propose an automated image evaluation pipeline. With this, we stratify existing high-resolution image datasets and larger-scale image datasets such as ImageNet and PASS to compare their performances. We find that datasets with (i) low compression artifacts, (ii) high within-image diversity as judged by the number of different objects, and (iii) a large number of images from ImageNet or PASS all positively affect SR performance. We hope that the proposed simple-yet-effective dataset curation pipeline will inform the construction of SR datasets in the future and yield overall better models.","url_abs":"https://arxiv.org/abs/2409.00768v1","url_pdf":"https://arxiv.org/pdf/2409.00768v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rethinking-image-super-resolution-from","repo_url":"https://github.com/gohtanii/DiverSeg-dataset","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"image-super-resolution","task_name":"Image Super-Resolution"},{"task_slug":"super-resolution","task_name":"Super-Resolution"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2409.00768","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2409.00768"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/gohtanii/DiverSeg-dataset","reach":null}],"summary":{"ran":2,"ran_honours":1,"ran_fixture":1},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"4f54d60e12da1bbb","entry":"calc_DCT","repo":"gohtanii/DiverSeg-dataset","repo_kind":"official","path":"blockiness/calc_blockiness.py","file_url":"https://github.com/gohtanii/DiverSeg-dataset/blob/HEAD/blockiness/calc_blockiness.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4f54d60e12da1bbb"}},{"code_sha256_prefix":"ae271c1df44dd73e","entry":"calc_V","repo":"gohtanii/DiverSeg-dataset","repo_kind":"official","path":"blockiness/calc_blockiness.py","file_url":"https://github.com/gohtanii/DiverSeg-dataset/blob/HEAD/blockiness/calc_blockiness.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ae271c1df44dd73e"}},{"code_sha256_prefix":"a2d2ec907298953a","entry":"calc_margin","repo":"gohtanii/DiverSeg-dataset","repo_kind":"official","path":"blockiness/calc_blockiness.py","file_url":"https://github.com/gohtanii/DiverSeg-dataset/blob/HEAD/blockiness/calc_blockiness.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a2d2ec907298953a"}},{"code_sha256_prefix":"6f93b531ec14ffeb","entry":"process_image","repo":"gohtanii/DiverSeg-dataset","repo_kind":"official","path":"blockiness/calc_blockiness.py","file_url":"https://github.com/gohtanii/DiverSeg-dataset/blob/HEAD/blockiness/calc_blockiness.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"DEP_MISSING","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6f93b531ec14ffeb"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}