Browse State-of-the-Art › Mixture-of-Experts › Papers, page 11
Mixture-of-Experts
Papers archive 2025-07-28
archive papers tagged: 1,312 · with a code link: 516 · where Syntology ran a sample: 216 (184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (216 of 1,312 tagged: 184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument)
Page 11 of 14: papers 1,001 to 1,100 of 1,312, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Half-Space Feature Learning in Neural Networks5 Apr 2024 0 repositories listed
-
Psychometry: An Omnifit Model for Image Reconstruction from Human Brain Activity29 Mar 2024 0 repositories listed
-
Revolutionizing Disease Diagnosis with simultaneous functional PET/MR and Deeply Integrated Brain Metabolic, Hemodynamic, and Perfusion Networks29 Mar 2024 0 repositories listed
-
Generalization Error Analysis for Sparse Mixture-of-Experts: A Preliminary Study26 Mar 2024 0 repositories listed
-
GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot20 Mar 2024 0 repositories listed
-
14 Mar 2024 0 repositories listed
-
Conditional computation in neural networks: principles and research trends12 Mar 2024 0 repositories listed
-
Acquiring Diverse Skills using Curriculum Reinforcement Learning with Mixture of Experts11 Mar 2024 0 repositories listed
-
MMoE: Robust Spoiler Detection with Multi-modal Information and Domain-aware Mixture-of-Experts8 Mar 2024 0 repositories listed
-
ConstitutionalExperts: Training a Mixture of Principle-based Prompts7 Mar 2024 0 repositories listed
-
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers4 Mar 2024 0 repositories listed
-
Hypertext Entity Extraction in Webpage4 Mar 2024 0 repositories listed
-
Vanilla Transformers are Transfer Capability Teachers4 Mar 2024 0 repositories listed
-
Enhancing the "Immunity" of Mixture-of-Experts Networks for Adversarial Defense29 Feb 2024 0 repositories listed
-
An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement27 Feb 2024 0 repositories listed
-
Co-Supervised Learning: Improving Weak-to-Strong Generalization with Hierarchical Mixture of Experts23 Feb 2024 0 repositories listed
-
PEMT: Multi-Task Correlation Guided Mixture-of-Experts Enables Parameter-Efficient Transfer Learning23 Feb 2024 0 repositories listed
-
Denoising OCT Images Using Steered Mixture of Experts with Multi-Model Inference20 Feb 2024 0 repositories listed
-
Towards an empirical understanding of MoE design choices20 Feb 2024 0 repositories listed
-
MoRAL: MoE Augmented LoRA for LLMs' Lifelong Learning17 Feb 2024 0 repositories listed
-
Turn Waste into Worth: Rectifying Top-k Router of MoE17 Feb 2024 0 repositories listed
-
AMEND: A Mixture of Experts Framework for Long-tailed Trajectory Prediction13 Feb 2024 0 repositories listed
-
P-Mamba: Marrying Perona Malik Diffusion with Mamba for Efficient Pediatric Echocardiographic Left Ventricular Segmentation13 Feb 2024 0 repositories listed
-
Differentially Private Training of Mixture of Experts Models11 Feb 2024 0 repositories listed
-
Buffer Overflow in Mixture of Experts8 Feb 2024 0 repositories listed
-
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts8 Feb 2024 0 repositories listed
-
On Parameter Estimation in Deviated Gaussian Mixture of Experts7 Feb 2024 0 repositories listed
-
Approximation Rates and VC-Dimension Bounds for (P)ReLU MLP Mixture of Experts5 Feb 2024 0 repositories listed
-
5 Feb 2024 0 repositories listed Syntology 3 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
On Least Square Estimation in Softmax Gating Mixture of Experts5 Feb 2024 0 repositories listed
-
MoDE: A Mixture-of-Experts Model with Mutual Distillation among the Experts31 Jan 2024 0 repositories listed
-
Explainable data-driven modeling via mixture of experts: towards effective blending of grey and black-box models30 Jan 2024 0 repositories listed
-
LLaVA-MoLE: Sparse Mixture of LoRA Experts for Mitigating Data Conflicts in Instruction Finetuning MLLMs29 Jan 2024 0 repositories listed
-
Routers in Vision Mixture of Experts: An Empirical Study29 Jan 2024 0 repositories listed
-
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?25 Jan 2024 0 repositories listed
-
M³TN: Multi-gate Mixture-of-Experts based Multi-valued Treatment Network for Uplift Modeling24 Jan 2024 0 repositories listed
-
Towards A Better Metric for Text-to-Video Generation15 Jan 2024 0 repositories listed
-
Prompt-based mental health screening from social media text11 Jan 2024 0 repositories listed
-
Robust Calibration For Improved Weather Prediction Under Distributional Shift8 Jan 2024 0 repositories listed
-
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models6 Jan 2024 0 repositories listed
-
Efficient Deweather Mixture-of-Experts with Uncertainty-aware Feature-wise Linear Modulation27 Dec 2023 0 repositories listed
-
Agent4Ranking: Semantic Robust Ranking via Personalized Query Rewriting Using Multi-agent LLM24 Dec 2023 0 repositories listed
-
Generator Assisted Mixture of Experts For Feature Acquisition in Batch19 Dec 2023 0 repositories listed
-
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning19 Dec 2023 0 repositories listed
-
From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape18 Dec 2023 0 repositories listed
-
Training of Neural Networks with Uncertain Data: A Mixture of Experts Approach13 Dec 2023 0 repositories listed
-
MoE-AMC: Enhancing Automatic Modulation Classification Performance Using Mixture-of-Experts4 Dec 2023 0 repositories listed
-
Language-driven All-in-one Adverse Weather Removal3 Dec 2023 0 repositories listed
-
MoEC: Mixture of Experts Implicit Neural Compression3 Dec 2023 0 repositories listed
-
1 Dec 2023 0 repositories listed
-
HOMOE: A Memory-Based and Composition-Aware Framework for Zero-Shot Learning with Hopfield Network and Soft Mixture of Experts23 Nov 2023 0 repositories listed
-
Efficient Model Agnostic Approach for Implicit Neural Representation Based Arbitrary-Scale Image Super-Resolution20 Nov 2023 0 repositories listed
-
Memory Augmented Language Models through Mixture of Word Experts15 Nov 2023 0 repositories listed
-
Intentional Biases in LLM Responses11 Nov 2023 0 repositories listed
-
CAME: Competitively Learning a Mixture-of-Experts Model for First-stage Retrieval6 Nov 2023 0 repositories listed
-
Mixture-of-Experts for Open Set Domain Adaptation: A Dual-Space Detection Approach1 Nov 2023 0 repositories listed
-
A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts22 Oct 2023 0 repositories listed
-
Direct Neural Machine Translation with Task-level Mixture of Experts models18 Oct 2023 0 repositories listed
-
Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer15 Oct 2023 0 repositories listed
-
Adaptive Gating in Mixture-of-Experts based Language Models11 Oct 2023 0 repositories listed
-
Beyond the Typical: Modeling Rare Plausible Patterns in Chemical Reactions by Leveraging Sequential Mixture-of-Experts7 Oct 2023 0 repositories listed
-
Reinforcement Learning-based Mixture of Vision Transformers for Video Violence Recognition4 Oct 2023 0 repositories listed
-
Mixture of Quantized Experts (MoQE): Complementary Effect of Low-bit Quantization and Robustness3 Oct 2023 0 repositories listed
-
Statistical Perspective of Top-K Sparse Softmax Gating Mixture of Experts25 Sep 2023 0 repositories listed
-
Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts8 Sep 2023 0 repositories listed
-
Task-Based MoE for Multitask Multilingual Machine Translation30 Aug 2023 0 repositories listed
-
SwapMoE: Serving Off-the-shelf MoE-based Large Language Models with Tunable Memory Budget29 Aug 2023 0 repositories listed
-
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE23 Aug 2023 0 repositories listed
-
FineQuant: Unlocking Efficiency with Fine-Grained Weight-Only Quantization for LLMs16 Aug 2023 0 repositories listed
-
Experts Weights Averaging: A New General Training Scheme for Vision Transformers11 Aug 2023 0 repositories listed
-
A Novel Temporal Multi-Gate Mixture-of-Experts Approach for Vehicle Trajectory and Driving Intention Prediction1 Aug 2023 0 repositories listed
-
Uncertainty-Encoded Multi-Modal Fusion for Robust Object Detection in Autonomous Driving30 Jul 2023 0 repositories listed
-
An Efficient General-Purpose Modular Vision Model via Multi-Task Heterogeneous Training29 Jun 2023 0 repositories listed
-
SkillNet-X: A Multilingual Multitask Model with Sparsely Activated Skills28 Jun 2023 0 repositories listed
-
JiuZhang 2.0: A Unified Chinese Pre-trained Language Model for Multi-task Mathematical Problem Solving19 Jun 2023 0 repositories listed
-
Learning to Specialize: Joint Gating-Expert Training for Adaptive MoEs in Decentralized Settings14 Jun 2023 0 repositories listed
-
Attention Weighted Mixture of Experts with Contrastive Learning for Personalized Ranking in E-commerce8 Jun 2023 0 repositories listed
-
Divide, Conquer, and Combine: Mixture of Semantic-Independent Experts for Zero-Shot Dialogue State Tracking1 Jun 2023 0 repositories listed
-
Modeling Task Relationships in Multi-variate Soft Sensor with Balanced Mixture-of-Experts25 May 2023 0 repositories listed
-
Mixture-of-Experts Meets Instruction Tuning:A Winning Combination for Large Language Models24 May 2023 0 repositories listed
-
Pre-training Multi-task Contrastive Learning Models for Scientific Literature Understanding23 May 2023 0 repositories listed
-
Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language Model23 May 2023 0 repositories listed
-
To Repeat or Not To Repeat: Insights from Scaling LLM under Token-Crisis22 May 2023 0 repositories listed
-
Lifelong Language Pretraining with Distribution-Specialized Experts20 May 2023 0 repositories listed
-
Locking and Quacking: Stacking Bayesian model predictions by log-pooling and superposition12 May 2023 0 repositories listed
-
Towards Convergence Rates for Parameter Estimation in Gaussian-gated Mixture of Experts12 May 2023 0 repositories listed
-
10 May 2023 0 repositories listed
-
Demystifying Softmax Gating Function in Gaussian Mixture of Experts5 May 2023 0 repositories listed
-
Steered Mixture-of-Experts Autoencoder Design for Real-Time Image Modelling and Denoising5 May 2023 0 repositories listed
-
Pipeline MoE: A Flexible MoE Implementation with Pipeline Parallelism22 Apr 2023 0 repositories listed
-
Revisiting Single-gated Mixtures of Experts11 Apr 2023 0 repositories listed
-
FlexMoE: Scaling Large-scale Sparse Pre-trained Model Training via Dynamic Device Placement8 Apr 2023 0 repositories listed
-
Mixed Regression via Approximate Message Passing5 Apr 2023 0 repositories listed
-
Steered Mixture of Experts Regression for Image Denoising with Multi-Model-Inference30 Mar 2023 0 repositories listed
-
WM-MoE: Weather-aware Multi-scale Mixture-of-Experts for Blind Adverse Weather Removal24 Mar 2023 0 repositories listed
-
Disguise without Disruption: Utility-Preserving Face De-Identification23 Mar 2023 0 repositories listed
-
Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset22 Mar 2023 0 repositories listed
-
HDformer: A Higher Dimensional Transformer for Diabetes Detection Utilizing Long Range Vascular Signals17 Mar 2023 0 repositories listed
-
MCR-DL: Mix-and-Match Communication Runtime for Deep Learning15 Mar 2023 0 repositories listed
-
Scaling Vision-Language Models with Sparse Mixture of Experts13 Mar 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.