{"url":"/method/dselect-k","slug":"dselect-k","name":"DSelect-k","full_name":"DSelect-k","full_name_withheld":false,"description_markdown":"**DSelect-k** is a continuously differentiable and sparse gate for Mixture-of-experts (MoE), based on a novel binary encoding formulation. Given a user-specified parameter $k$, the gate selects at most $k$ out of the $n$ experts. The gate can be trained using first-order methods, such as stochastic gradient descent, and offers explicit control over the number of experts to select. This explicit control over sparsity leads to a cardinality-constrained optimization problem, which is computationally challenging. To circumvent this challenge, the authors use a unconstrained reformulation that is equivalent to the original problem. The reformulated problem uses a binary encoding scheme to implicitly enforce the cardinality constraint. By carefully smoothing the binary encoding variables, the reformulated problem can be effectively optimized using first-order methods such as [SGD](https://paperswithcode.com/method/sgd).\r\n\r\nThe motivation for this method is that  existing sparse gates, such as Top-k, are not smooth. The lack of smoothness can lead to convergence and statistical performance issues when training with gradient-based methods.","description_state":"present","introduced_year":null,"introduced_by":{"title":"DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning","paper":"/paper/dselect-k-differentiable-selection-in-the","first_author":"Hussein Hazimeh","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/dselect-k-differentiable-selection-in-the"},"source":{"url":"https://arxiv.org/abs/2106.03760v3","title":"DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Mixture-of-Experts","url":"/methods/category/mixture-of-experts","pwc_aliases":[]}],"n_papers_tagged":2,"archive_num_papers":2,"papers_newest_first":[{"paper":"/paper/comet-learning-cardinality-constrained","title":"COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local Search","date":"2023-06-05","arxiv_id":"2306.02824","n_code_links":2,"syntology":null},{"paper":"/paper/dselect-k-differentiable-selection-in-the","title":"DSelect-k: Differentiable Selection in the Mixture of Experts with Applications to Multi-Task Learning","date":"2021-06-07","arxiv_id":"2106.03760","n_code_links":3,"syntology":{"ran":1,"of":9,"unverified":8,"pointer_only":0}}],"papers_shown":2,"tasks":[{"task":"/task/mixture-of-experts","name":"Mixture-of-Experts","papers":2},{"task":"/task/recommendation-systems","name":"Recommendation Systems","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/multi-task-learning","name":"Multi-Task Learning","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2021","papers":1},{"year":"2023","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/dselect-k"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}