{"url":"/method/dynabert","slug":"dynabert","name":"DynaBERT","full_name":"DynaBERT","full_name_withheld":false,"description_markdown":"**DynaBERT** is a [BERT](https://paperswithcode.com/method/bert)-variant which can flexibly adjust the size and latency by selecting adaptive width and depth. The training process of DynaBERT includes first training a width-adaptive BERT and then allowing both adaptive width and depth, by distilling knowledge from the full-sized model to small sub-networks. Network rewiring is also used to keep the more important attention heads and neurons shared by more sub-networks. \r\n\r\nA two-stage procedure is used to train DynaBERT. First, using knowledge distillation (dashed lines) to transfer the knowledge from a fixed teacher model to student sub-networks with adaptive width in DynaBERTW. Then, using knowledge distillation (dashed lines) to transfer the knowledge from a trained DynaBERTW to student sub-networks with adaptive width and depth in DynaBERT.","description_state":"present","introduced_year":null,"introduced_by":{"title":"DynaBERT: Dynamic BERT with Adaptive Width and Depth","paper":"/paper/dynabert-dynamic-bert-with-adaptive-width-and","first_author":"Lu Hou","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/dynabert-dynamic-bert-with-adaptive-width-and"},"source":{"url":"https://arxiv.org/abs/2004.04037v2","title":"DynaBERT: Dynamic BERT with Adaptive Width and Depth","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":"/paper/distilling-linguistic-context-for-language","title":"Distilling Linguistic Context for Language Model Compression","date":"2021-09-17","arxiv_id":"2109.08359","n_code_links":1,"syntology":{"ran":2,"of":11,"unverified":9,"pointer_only":0}},{"paper":"/paper/dynabert-dynamic-bert-with-adaptive-width-and","title":"DynaBERT: Dynamic BERT with Adaptive Width and Depth","date":"2020-04-08","arxiv_id":"2004.04037","n_code_links":3,"syntology":{"ran":5,"of":8,"unverified":3,"pointer_only":8}}],"papers_shown":2,"tasks":[{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/language-modelling","name":"Language Modelling","papers":2},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":1},{"task":"/task/model-compression","name":"Model Compression","papers":1},{"task":null,"name":"Relation","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/model","name":"model","papers":1}],"tasks_shown":7,"n_tasks":7,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/dynabert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}