Methods › General › Ensembling › MoE › Papers, page 3
Mixture of Experts
MoE
Papers archive 2025-07-28
archive papers tagged: 366 · with a code link: 142 · where Syntology ran a sample: 64 (49 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (64 of 366 tagged: 49 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 3 of 4: papers 201 to 300 of 366, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models 1 Nov 2024 · 1 repository · arXiv:2411.00918
-
MoE-I²: Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition 1 Nov 2024 · 0 repositories · arXiv:2411.01016
-
Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts 31 Oct 2024 · 0 repositories · arXiv:2410.23836
-
Efficient and Effective Weight-Ensembling Mixture of Experts for Multi-Task Model Merging 29 Oct 2024 · 0 repositories · arXiv:2410.21804
-
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields 27 Oct 2024 · 0 repositories · arXiv:2410.20306
-
DMT-HI: MOE-based Hyperbolic Interpretable Deep Manifold Transformation for Unspervised Dimensionality Reduction 25 Oct 2024 · 1 repository · arXiv:2410.19504
-
Hierarchical Mixture of Experts: Generalizable Learning for High-Level Synthesis 25 Oct 2024 · 1 repository · arXiv:2410.19225
-
Read-ME: Refactorizing LLMs as Router-Decoupled Mixture of Experts with System Co-Design 24 Oct 2024 · 1 repository · arXiv:2410.19123Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
ExpertFlow: Optimized Expert Activation and Token Allocation for Efficient Mixture-of-Experts Inference 23 Oct 2024 · 0 repositories · arXiv:2410.17954
-
MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning 23 Oct 2024 · 0 repositories · arXiv:2410.18035
-
Optimizing Mixture-of-Experts Inference Time Combining Model Deployment and Communication Scheduling 22 Oct 2024 · 0 repositories · arXiv:2410.17043
-
CartesianMoE: Boosting Knowledge Sharing among Experts via Cartesian Product Routing in Mixture-of-Experts 21 Oct 2024 · 1 repository · arXiv:2410.16077
-
Generalizing Motion Planners with Mixture of Experts for Autonomous Driving 21 Oct 2024 · 1 repository · arXiv:2410.15774Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ViMoE: An Empirical Study of Designing Vision Mixture-of-Experts 21 Oct 2024 · 0 repositories · arXiv:2410.15732
-
Collaboratively adding new knowledge to an LLM 18 Oct 2024 · 1 repository · arXiv:2410.14753
-
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts 18 Oct 2024 · 1 repository · arXiv:2410.14574Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference 16 Oct 2024 · 0 repositories · arXiv:2410.12247
-
MoE-Pruner: Pruning Mixture-of-Experts Large Language Model using the Hints from Its Router 15 Oct 2024 · 0 repositories · arXiv:2410.12013
-
Quadratic Gating Functions in Mixture of Experts: A Statistical Insight 15 Oct 2024 · 0 repositories · arXiv:2410.11222
-
Ada-K Routing: Boosting the Efficiency of MoE-based LLMs 14 Oct 2024 · 0 repositories · arXiv:2410.10456
-
AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality 14 Oct 2024 · 1 repository · arXiv:2410.10054Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts 14 Oct 2024 · 1 repository · arXiv:2410.10626Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Learning to Ground VLMs without Forgetting 14 Oct 2024 · 0 repositories · arXiv:2410.10491
-
Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts 14 Oct 2024 · 1 repository · arXiv:2410.10469
-
Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free 14 Oct 2024 · 1 repository · arXiv:2410.10814Syntology official (archive's flag): 4 ran · 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
AT-MoE: Adaptive Task-planning Mixture of Experts via LoRA Approach 12 Oct 2024 · 0 repositories · arXiv:2410.10896
-
Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts 10 Oct 2024 · 1 repository · arXiv:2410.08245Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Upcycling Large Language Models into Mixture of Experts 10 Oct 2024 · 0 repositories · arXiv:2410.07524
-
MoE++: Accelerating Mixture-of-Experts Methods with Zero-Computation Experts 9 Oct 2024 · 3 repositories · arXiv:2410.07348Syntology official: no sample here; runs from other or unrecorded repositories · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models 8 Oct 2024 · 0 repositories · arXiv:2410.05661
-
A Dynamic Approach to Stock Price Prediction: Comparing RNN and Mixture of Experts Models Across Different Volatility Profiles 4 Oct 2024 · 0 repositories · arXiv:2410.07234
-
Exploring the Benefit of Activation Sparsity in Pre-training 4 Oct 2024 · 1 repository · arXiv:2410.03440Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping 3 Oct 2024 · 0 repositories · arXiv:2410.02475
-
On Expert Estimation in Hierarchical Mixture of Experts: Beyond Softmax Gating Functions 3 Oct 2024 · 0 repositories · arXiv:2410.02935
-
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices 3 Oct 2024 · 1 repository · arXiv:2410.02117Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
EC-DIT: Scaling Diffusion Transformers with Adaptive Expert-Choice Routing 2 Oct 2024 · 0 repositories · arXiv:2410.02098
-
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging 2 Oct 2024 · 0 repositories · arXiv:2410.01610
-
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models 1 Oct 2024 · 0 repositories · arXiv:2410.10846
-
IDEA: An Inverse Domain Expert Adaptation Based Active DNN IP Protection Method 29 Sep 2024 · 0 repositories · arXiv:2410.00059
-
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling 28 Sep 2024 · 1 repository · arXiv:2409.19291Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM 24 Sep 2024 · 0 repositories · arXiv:2409.15905
-
A Gated Residual Kolmogorov-Arnold Networks for Mixtures of Experts 23 Sep 2024 · 1 repository · arXiv:2409.15161
-
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists 20 Sep 2024 · 1 repository · arXiv:2409.13931
-
Retrieval-Augmented Test Generation: How Far Are We? 19 Sep 2024 · 0 repositories · arXiv:2409.12682
-
GRIN: GRadient-INformed MoE 18 Sep 2024 · 0 repositories · arXiv:2409.12136
-
Mixture of Diverse Size Experts 18 Sep 2024 · 0 repositories · arXiv:2409.12210
-
DA-MoE: Towards Dynamic Expert Allocation for Mixture-of-Experts Models 10 Sep 2024 · 0 repositories · arXiv:2409.06669
-
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning 10 Sep 2024 · 0 repositories · arXiv:2409.06211
-
M3-Jepa: Multimodal Alignment via Multi-directional MoE based on the JEPA framework 9 Sep 2024 · 1 repository · arXiv:2409.05929Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
ChartMoE: Mixture of Expert Connector for Advanced Chart Understanding 5 Sep 2024 · 0 repositories · arXiv:2409.03277
-
Interpretable mixture of experts for time series prediction under recurrent and non-recurrent conditions 5 Sep 2024 · 0 repositories · arXiv:2409.03282
-
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model 3 Sep 2024 · 0 repositories · arXiv:2409.02050
-
OLMoE: Open Mixture-of-Experts Language Models 3 Sep 2024 · 2 repositories · arXiv:2409.02060Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts 2 Sep 2024 · 0 repositories · arXiv:2409.00879
-
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching 2 Sep 2024 · 0 repositories · arXiv:2409.01141
-
Revisiting SMoE Language Models by Evaluating Inefficiencies with Task Specific Expert Pruning 2 Sep 2024 · 0 repositories · arXiv:2409.01483
-
Flexible and Effective Mixing of Large Language Models into a Mixture of Domain Experts 30 Aug 2024 · 0 repositories · arXiv:2408.17280
-
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts 28 Aug 2024 · 0 repositories · arXiv:2408.15664
-
Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models 28 Aug 2024 · 1 repository · arXiv:2408.15915
-
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts 28 Aug 2024 · 0 repositories · arXiv:2408.15901
-
A Survey of Large Language Models for European Languages 27 Aug 2024 · 0 repositories · arXiv:2408.15040
-
La-SoftMoE CLIP for Unified Physical-Digital Face Attack Detection 23 Aug 2024 · 0 repositories · arXiv:2408.12793
-
Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler 23 Aug 2024 · 1 repository · arXiv:2408.13359
-
FedMoE: Personalized Federated Learning via Heterogeneous Mixture of Experts 21 Aug 2024 · 0 repositories · arXiv:2408.11304
-
MoE-LPR: Multilingual Extension of Large Language Models through Mixture-of-Experts with Language Priors Routing 21 Aug 2024 · 2 repositories · arXiv:2408.11396
-
HMoE: Heterogeneous Mixture of Experts for Language Modeling 20 Aug 2024 · 0 repositories · arXiv:2408.10681
-
A Unified Framework for Iris Anti-Spoofing: Introducing IrisGeneral Dataset and Masked-MoE Method 19 Aug 2024 · 0 repositories · arXiv:2408.09752
-
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference 19 Aug 2024 · 1 repository · arXiv:2408.10284Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
SMILE: Zero-Shot Sparse Mixture of Low-Rank Experts Construction From Pre-Trained Foundation Models 19 Aug 2024 · 1 repository · arXiv:2408.10174
-
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts 15 Aug 2024 · 0 repositories · arXiv:2408.08274
-
AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies 13 Aug 2024 · 1 repository · arXiv:2408.06567
-
Layerwise Recurrent Router for Mixture-of-Experts 13 Aug 2024 · 1 repository · arXiv:2408.06793
-
HoME: Hierarchy of Multi-Gate Experts for Multi-Task Learning at Kuaishou 10 Aug 2024 · 0 repositories · arXiv:2408.05430
-
LaDiMo: Layer-wise Distillation Inspired MoEfier 8 Aug 2024 · 0 repositories · arXiv:2408.04278
-
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training 8 Aug 2024 · 0 repositories · arXiv:2408.04307
-
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning 8 Aug 2024 · 1 repository · arXiv:2408.04693
-
MoExtend: Tuning New Experts for Modality and Task Extension 7 Aug 2024 · 1 repository · arXiv:2408.03511Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts 31 Jul 2024 · 0 repositories · arXiv:2407.21770
-
Mixture of Modular Experts: Distilling Knowledge from a Multilingual Teacher into Specialized Modular Language Models 28 Jul 2024 · 1 repository · arXiv:2407.19610
-
Dynamic Language Group-Based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing 26 Jul 2024 · 1 repository · arXiv:2407.18581
-
How Lightweight Can A Vision Transformer Be 25 Jul 2024 · 0 repositories · arXiv:2407.17783
-
EEGMamba: Bidirectional State Space Model with Mixture of Experts for EEG Multi-task Classification 20 Jul 2024 · 0 repositories · arXiv:2407.20254
-
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service 19 Jul 2024 · 0 repositories · arXiv:2407.14417
-
Scaling Diffusion Transformers to 16 Billion Parameters 16 Jul 2024 · 1 repository · arXiv:2407.11633Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 4 honoured, 1 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
Low-Rank Interconnected Adaptation across Layers 13 Jul 2024 · 1 repository · arXiv:2407.09946Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
MaskMoE: Boosting Token-Level Learning via Routing Mask in Mixture-of-Experts 13 Jul 2024 · 1 repository · arXiv:2407.09816
-
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts 12 Jul 2024 · 0 repositories · arXiv:2407.09590
-
Swin SMT: Global Sequential Modeling in 3D Medical Image Segmentation 10 Jul 2024 · 1 repository · arXiv:2407.07514
-
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts 9 Jul 2024 · 0 repositories · arXiv:2407.06718
-
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models with Adaptive Expert Placement 5 Jul 2024 · 0 repositories · arXiv:2407.04656
-
LoCo: Low-Bit Communication Adaptor for Large-scale Model Training 5 Jul 2024 · 1 repository · arXiv:2407.04480
-
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation 5 Jul 2024 · 1 repository · arXiv:2407.04822
-
Mixture of A Million Experts 4 Jul 2024 · 1 repository · arXiv:2407.04153Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Efficient-Empathy: Towards Efficient and Effective Selection of Empathy Data 2 Jul 2024 · 0 repositories · arXiv:2407.01937
-
Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models 2 Jul 2024 · 1 repository · arXiv:2407.01906Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Parm: Efficient Training of Large Sparsely-Activated Models with Dedicated Schedules 30 Jun 2024 · 1 repository · arXiv:2407.00599
-
LEMoE: Advanced Mixture of Experts Adaptor for Lifelong Model Editing of Large Language Models 28 Jun 2024 · 0 repositories · arXiv:2406.20030
-
Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model 28 Jun 2024 · 1 repository · arXiv:2406.19905
-
A Closer Look into Mixture-of-Experts in Large Language Models 26 Jun 2024 · 1 repository · arXiv:2406.18219
-
A Survey on Mixture of Experts 26 Jun 2024 · 1 repository · arXiv:2407.06204