{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/language-through-a-prism-a-spectral-approach","title":"Language Through a Prism: A Spectral Approach for Multiscale Language Representations","arxiv_id":"2011.04823","date":"2020-11-09","proceeding":"NeurIPS 2020 12","authors":["Alex Tamkin","Dan Jurafsky","Noah Goodman"],"abstract":"Language exhibits structure at different scales, ranging from subwords to words, sentences, paragraphs, and documents. To what extent do deep models capture information at these scales, and can we force them to better capture structure across this hierarchy? We approach this question by focusing on individual neurons, analyzing the behavior of their activations at different timescales. We show that signal processing provides a natural framework for separating structure across scales, enabling us to 1) disentangle scale-specific information in existing embeddings and 2) train models to learn more about particular scales. Concretely, we apply spectral filters to the activations of a neuron across an input, producing filtered embeddings that perform well on part of speech tagging (word-level), dialog speech acts classification (utterance-level), or topic classification (document-level), while performing poorly on the other tasks. We also present a prism layer for training models, which uses spectral filters to constrain different neurons to model structure at different scales. Our proposed BERT + Prism model can better predict masked tokens using long-range context and produces multiscale representations that perform better at utterance- and document-level tasks. Our methods are general and readily applicable to other domains besides language, such as images, audio, and video.","url_abs":"https://arxiv.org/abs/2011.04823v1","url_pdf":"https://arxiv.org/pdf/2011.04823v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"language-through-a-prism-a-spectral-approach","repo_url":"https://github.com/zh217/torch-dct","is_official":0,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"part-of-speech-tagging","task_name":"Part-Of-Speech Tagging"},{"task_slug":"topic-classification","task_name":"Topic Classification"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"attention-dropout","method_name":"Attention Dropout"},{"method_slug":"bert","method_name":"BERT"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"linear-warmup-with-linear-decay","method_name":"Linear Warmup With Linear Decay"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"weight-decay","method_name":"Weight Decay"},{"method_slug":"wordpiece","method_name":"WordPiece"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2011.04823","atlas_url":"https://app.syntology.ai/?focus=2011.04823","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2011.04823"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zh217/torch-dct","reach":null}],"summary":{"unverified":5},"by_repo_kind":{"named_in_paper":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"8ee2bb43bd28f6ac","entry":"LinearDCT","repo":"zh217/torch-dct","repo_kind":"named_in_paper","path":"torch_dct/_dct.py","file_url":"https://github.com/zh217/torch-dct/blob/HEAD/torch_dct/_dct.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8ee2bb43bd28f6ac"}},{"code_sha256_prefix":"39394813950ed362","entry":"dct","repo":"zh217/torch-dct","repo_kind":"named_in_paper","path":"torch_dct/_dct.py","file_url":"https://github.com/zh217/torch-dct/blob/HEAD/torch_dct/_dct.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"39394813950ed362"}},{"code_sha256_prefix":"f3f999af44b26296","entry":"dct1","repo":"zh217/torch-dct","repo_kind":"named_in_paper","path":"torch_dct/_dct.py","file_url":"https://github.com/zh217/torch-dct/blob/HEAD/torch_dct/_dct.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f3f999af44b26296"}},{"code_sha256_prefix":"fdb926eceb3af796","entry":"idct","repo":"zh217/torch-dct","repo_kind":"named_in_paper","path":"torch_dct/_dct.py","file_url":"https://github.com/zh217/torch-dct/blob/HEAD/torch_dct/_dct.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fdb926eceb3af796"}},{"code_sha256_prefix":"2f88f7130eaa3ed7","entry":"idct1","repo":"zh217/torch-dct","repo_kind":"named_in_paper","path":"torch_dct/_dct.py","file_url":"https://github.com/zh217/torch-dct/blob/HEAD/torch_dct/_dct.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2f88f7130eaa3ed7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}