{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/explicitly-increasing-input-information","title":"Explicitly Increasing Input Information Density for Vision Transformers on Small Datasets","arxiv_id":"2210.14319","date":"2022-10-25","proceeding":null,"authors":["Xiangyu Chen","Ying Qin","Wenju Xu","Andrés M. Bur","Cuncong Zhong","Guanghui Wang"],"abstract":"Vision Transformers have attracted a lot of attention recently since the successful implementation of Vision Transformer (ViT) on vision tasks. With vision Transformers, specifically the multi-head self-attention modules, networks can capture long-term dependencies inherently. However, these attention modules normally need to be trained on large datasets, and vision Transformers show inferior performance on small datasets when training from scratch compared with widely dominant backbones like ResNets. Note that the Transformer model was first proposed for natural language processing, which carries denser information than natural images. To boost the performance of vision Transformers on small datasets, this paper proposes to explicitly increase the input information density in the frequency domain. Specifically, we introduce selecting channels by calculating the channel-wise heatmaps in the frequency domain using Discrete Cosine Transform (DCT), reducing the size of input while keeping most information and hence increasing the information density. As a result, 25% fewer channels are kept while better performance is achieved compared with previous work. Extensive experiments demonstrate the effectiveness of the proposed approach on five small-scale datasets, including CIFAR-10/100, SVHN, Flowers-102, and Tiny ImageNet. The accuracy has been boosted up to 17.05% with Swin and Focal Transformers. Codes are available at https://github.com/xiangyu8/DenseVT.","url_abs":"https://arxiv.org/abs/2210.14319v1","url_pdf":"https://arxiv.org/pdf/2210.14319v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"explicitly-increasing-input-information","repo_url":"https://github.com/xiangyu8/densevt","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"discrete-cosine-transform","method_name":"Discrete Cosine Transform"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"focal-transformers","method_name":"Focal Transformers"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"},{"method_slug":"vision-transformer","method_name":"Vision Transformer"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.14319","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.14319"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/xiangyu8/DenseVT","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/xiangyu8/densevt","reach":{"status":"ok"}}],"summary":{"ran":2,"ran_draft_wrong":1,"ran_fixture":1,"unverified":3},"by_repo_kind":{"official":{"samples":7,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"2256995db71523d9","entry":"PatchEmbed","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2256995db71523d9"}},{"code_sha256_prefix":"3ca18b8b2792fa48","entry":"WindowAttention","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3ca18b8b2792fa48"}},{"code_sha256_prefix":"03883eedf849bf6a","entry":"get_topk_closest_indice","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"03883eedf849bf6a"}},{"code_sha256_prefix":"cf5b29b2daec6574","entry":"window_partition_noreshape","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"cf5b29b2daec6574"}},{"code_sha256_prefix":"bd1c47736a9b3030","entry":"BasicLayer","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bd1c47736a9b3030"}},{"code_sha256_prefix":"0ea62000c8198fd3","entry":"FocalTransformer","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"0ea62000c8198fd3"}},{"code_sha256_prefix":"7b04194d50a7ba24","entry":"FocalTransformerBlock","repo":"xiangyu8/densevt","repo_kind":"official","path":"models/focalvit.py","file_url":"https://github.com/xiangyu8/densevt/blob/HEAD/models/focalvit.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7b04194d50a7ba24"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}