Browse State-of-the-Art › Mixture-of-Experts › Papers, page 12
Mixture-of-Experts
Papers archive 2025-07-28
archive papers tagged: 1,312 · with a code link: 516 · where Syntology ran a sample: 216 (184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (216 of 1,312 tagged: 184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument)
Page 12 of 14: papers 1,101 to 1,200 of 1,312, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference10 Mar 2023 0 repositories listed
-
Improving Expert Specialization in Mixture of Experts28 Feb 2023 0 repositories listed
-
Improved Training of Mixture-of-Experts Language GANs23 Feb 2023 0 repositories listed
-
TMoE-P: Towards the Pareto Optimum for Multivariate Soft Sensors21 Feb 2023 0 repositories listed
-
Massively Multilingual Shallow Fusion with Large Language Models17 Feb 2023 0 repositories listed
-
Fast, Differentiable and Sparse Top-k: a Convex Analysis Perspective2 Feb 2023 0 repositories listed
-
Alternating Updates for Efficient Transformers30 Jan 2023 0 repositories listed
-
PRUDEX-Compass: Towards Systematic Evaluation of Reinforcement Learning in Financial Markets14 Jan 2023 0 repositories listed
-
AdaEnsemble: Learning Adaptively Sparse Structured Ensemble Network for Click-Through Rate Prediction6 Jan 2023 0 repositories listed
-
Mod-Squad: Designing Mixtures of Experts As Modular Multi-Task Learners1 Jan 2023 0 repositories listed
-
Semantic-Aware Dynamic Parameter for Video Inpainting Transformer1 Jan 2023 0 repositories listed
-
Generalizing Multimodal Variational Methods to Sets19 Dec 2022 0 repositories listed
-
Memory-efficient NLLB-200: Language-specific Expert Pruning of a Massively Multilingual Machine Translation Model19 Dec 2022 0 repositories listed
-
MultiCoder: Multi-Programming-Lingual Pre-Training for Low-Resource Code Completion19 Dec 2022 0 repositories listed
-
Fixing MoE Over-Fitting on Low-Resource Languages in Multilingual Machine Translation15 Dec 2022 0 repositories listed
-
Mod-Squad: Designing Mixture of Experts As Modular Multi-Task Learners15 Dec 2022 0 repositories listed
-
SMILE: Scaling Mixture-of-Experts with Efficient Bi-level Routing10 Dec 2022 0 repositories listed
-
Incorporating Polar Field Data for Improved Solar Flare Prediction4 Dec 2022 0 repositories listed
-
Automatically Extracting Information in Medical Dialogue: Expert System And Attention for Labelling28 Nov 2022 0 repositories listed
-
Double Deep Q-Learning in Opponent Modeling24 Nov 2022 0 repositories listed
-
Who Says Elephants Can't Run: Bringing Large Scale MoE Models into Cloud Scale Production18 Nov 2022 0 repositories listed
-
HMOE: Hypernetwork-based Mixture of Experts for Domain Generalization15 Nov 2022 0 repositories listed
-
Handling Trade-Offs in Speech Separation with Sparsely-Gated Mixture of Experts11 Nov 2022 0 repositories listed
-
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations8 Nov 2022 0 repositories listed
-
Using Deep Mixture-of-Experts to Detect Word Meaning Shift for TempoWiC7 Nov 2022 0 repositories listed
-
Safe Real-World Autonomous Driving by Learning to Predict and Plan with a Mixture of Experts3 Nov 2022 0 repositories listed
-
Contextual Mixture of Experts: Integrating Knowledge into Predictive Modeling1 Nov 2022 0 repositories listed
-
Prediction Sets for High-Dimensional Mixture of Experts Models30 Oct 2022 0 repositories listed
-
28 Oct 2022 0 repositories listed
-
Coordination with Humans via Strategy Matching27 Oct 2022 0 repositories listed
-
On the Adversarial Robustness of Mixture of Experts19 Oct 2022 0 repositories listed
-
Tiny-Attention Adapter: Contexts Are More Important Than the Number of Parameters18 Oct 2022 0 repositories listed
-
FEAMOE: Fair, Explainable and Adaptive Mixture of Experts10 Oct 2022 0 repositories listed
-
Deep Learning Mixture-of-Experts Approach for Cytotoxic Edema Assessment in Infants and Children6 Oct 2022 0 repositories listed
-
Probabilistic partition of unity networks for high-dimensional regression problems6 Oct 2022 0 repositories listed
-
Parameter-varying neural ordinary differential equations with partition-of-unity networks1 Oct 2022 0 repositories listed
-
Table-based Fact Verification with Self-labeled Keypoint Alignment1 Oct 2022 0 repositories listed
-
Mixture of experts models for multilevel data: modelling framework and approximation theory30 Sep 2022 0 repositories listed
-
Sparsity-Constrained Optimal Transport30 Sep 2022 0 repositories listed
-
Tuning of Mixture-of-Experts Mixed-Precision Neural Networks29 Sep 2022 0 repositories listed
-
Diversified Dynamic Routing for Vision Tasks26 Sep 2022 0 repositories listed
-
Parameter-Efficient Conformers via Sharing Sparsely-Gated Experts for End-to-End Speech Recognition17 Sep 2022 0 repositories listed
-
Sparse Video Representation Using Steered Mixture-of-Experts With Global Motion Compensation13 Sep 2022 0 repositories listed
-
A Review of Sparse Expert Models in Deep Learning4 Sep 2022 0 repositories listed
-
Context-aware Mixture-of-Experts for Unbiased Scene Graph Generation15 Aug 2022 0 repositories listed
-
A Theoretical View on Sparsely Activated Networks8 Aug 2022 0 repositories listed
-
Edge-Aware Autoencoder Design for Real-Time Mixture-of-Experts Image Compression25 Jul 2022 0 repositories listed
-
Adaptive Mixture of Experts Learning for Generalizable Face Anti-Spoofing20 Jul 2022 0 repositories listed
-
MoEC: Mixture of Expert Clusters19 Jul 2022 0 repositories listed
-
Learning Large-scale Universal User Representation with Sparse Mixture of Experts11 Jul 2022 0 repositories listed
-
Scalable Neural Data Server: A Data Recommender for Transfer Learning19 Jun 2022 0 repositories listed
-
Quantitative Stock Investment by Routing Uncertainty-Aware Trading Experts: A Multi-Task Learning Approach7 Jun 2022 0 repositories listed
-
Multimodal Contrastive Learning with LIMoE: the Language-Image Mixture of Experts6 Jun 2022 0 repositories listed
-
Interpretable Mixture of Experts5 Jun 2022 0 repositories listed
-
Task-Specific Expert Pruning for Sparse Mixture-of-Experts1 Jun 2022 0 repositories listed
-
Automatic Expert Selection for Multi-Scenario and Multi-Task Search28 May 2022 0 repositories listed
-
Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers28 May 2022 0 repositories listed
-
Pluralistic Image Completion with Probabilistic Mixture-of-Experts18 May 2022 0 repositories listed
-
Unified Modeling of Multi-Domain Multi-Device ASR Systems13 May 2022 0 repositories listed
-
ST-ExpertNet: A Deep Expert Framework for Traffic Prediction5 May 2022 0 repositories listed
-
Optimizing Mixture of Experts using Dynamic Recompilations4 May 2022 0 repositories listed
-
How Can Cross-lingual Knowledge Contribute Better to Fine-Grained Entity Typing?1 May 2022 0 repositories listed
-
Residual Mixture of Experts20 Apr 2022 0 repositories listed
-
Towards Efficient Single Image Dehazing and Desnowing19 Apr 2022 0 repositories listed
-
Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners16 Apr 2022 0 repositories listed
-
Mixture of Experts for Biomedical Question Answering15 Apr 2022 0 repositories listed
-
Mixture-of-experts VAEs can disregard variation in surjective multimodal data11 Apr 2022 0 repositories listed
-
On the Adaptation to Concept Drift for CTR Prediction1 Apr 2022 0 repositories listed
-
Efficient Reflectance Capture with a Deep Gated Mixture-of-Experts29 Mar 2022 0 repositories listed
-
14 Mar 2022 0 repositories listed
-
SkillNet-NLU: A Sparsely Activated Model for General-Purpose Natural Language Understanding7 Mar 2022 0 repositories listed
-
Functional mixture-of-experts for classification28 Feb 2022 0 repositories listed
-
Mixture-of-Experts with Expert Choice Routing18 Feb 2022 0 repositories listed
-
A Survey on Dynamic Neural Networks for Natural Language Processing15 Feb 2022 0 repositories listed
-
Physics-Guided Problem Decomposition for Scaling Deep Learning of High-dimensional Eigen-Solvers: The Case of Schrödinger's Equation12 Feb 2022 0 repositories listed
-
One Student Knows All Experts Know: From Sparse to Dense26 Jan 2022 0 repositories listed
-
MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation16 Jan 2022 0 repositories listed
-
Sparsely Activated Mixture-of-Experts are Robust Multi-Task Learners16 Jan 2022 0 repositories listed
-
Towards Lightweight Neural Animation : Exploration of Neural Network Pruning in Mixture of Experts-based Animation Models11 Jan 2022 0 repositories listed
-
Combinations of Adaptive Filters22 Dec 2021 0 repositories listed
-
Efficient Large Scale Language Modeling with Mixtures of Experts20 Dec 2021 0 repositories listed
-
13 Dec 2021 0 repositories listed
-
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition10 Dec 2021 0 repositories listed
-
Anchoring to Exemplars for Training Mixture-of-Expert Cell Embeddings6 Dec 2021 0 repositories listed
-
A Mixture of Expert Based Deep Neural Network for Improved ASR2 Dec 2021 0 repositories listed
-
TAL: Two-stream Adaptive Learning for Generalizable Person Re-identification29 Nov 2021 0 repositories listed
-
Expert Aggregation for Financial Forecasting25 Nov 2021 0 repositories listed
-
SpeechMoE2: Mixture-of-Experts Model with Improved Routing23 Nov 2021 0 repositories listed
-
M6-T: Exploring Sparse Expert Models and Beyond16 Nov 2021 0 repositories listed
-
MoEfication: Conditional Computation of Transformer Models for Efficient Inference16 Nov 2021 0 repositories listed
-
StableMoE: Stable Routing Strategy for Mixture of Experts16 Nov 2021 0 repositories listed
-
SummaReranker: A Multi-Task Mixture-of-Experts Re-ranking Framework for Abstractive Summarization16 Nov 2021 0 repositories listed
-
Table-based Fact Verification with Self-adaptive Mixture of Experts16 Nov 2021 0 repositories listed
-
RTM Super Learner Results at Quality Estimation Task1 Nov 2021 0 repositories listed
-
Polynomial-Spline Neural Networks with Exact Integrals26 Oct 2021 0 repositories listed
-
Simple or Complex? Complexity-Controllable Question Generation with Soft Templates and Deep Mixture of Experts Model13 Oct 2021 0 repositories listed
-
Continual Learning Using Task Conditional Neural Networks29 Sep 2021 0 repositories listed
-
Full-Precision Free Binary Graph Neural Networks29 Sep 2021 0 repositories listed
-
HydraSum - Disentangling Stylistic Features in Text Summarization using Multi-Decoder Models29 Sep 2021 0 repositories listed
-
MECATS: Mixture-of-Experts for Probabilistic Forecasts of Aggregated Time Series29 Sep 2021 0 repositories listed