Methods › General › Ensembling › MoE › Papers, page 2
Mixture of Experts
MoE
Papers archive 2025-07-28
archive papers tagged: 366 · with a code link: 142 · where Syntology ran a sample: 64 (49 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (64 of 366 tagged: 49 with a run with no instrument failure, 15 where every run was a failure of Syntology's instrument)
Page 2 of 4: papers 101 to 200 of 366, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs 7 Mar 2025 · 0 repositories · arXiv:2503.05139
-
Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts 7 Mar 2025 · 1 repository · arXiv:2503.05447
-
Continual Pre-training of MoEs: How robust is your router? 6 Mar 2025 · 0 repositories · arXiv:2503.05029
-
Speculative MoE: Communication Efficient Parallel MoE Inference with Speculative Token and Expert Pre-scheduling 6 Mar 2025 · 0 repositories · arXiv:2503.04398
-
Convergence Rates for Softmax Gating Mixture of Experts 5 Mar 2025 · 0 repositories · arXiv:2503.03213
-
Union of Experts: Adapting Hierarchical Routing to Equivalently Decomposed Transformer 4 Mar 2025 · 1 repository · arXiv:2503.02495
-
DeRS: Towards Extremely Efficient Upcycled Mixture-of-Experts Models 3 Mar 2025 · 0 repositories · arXiv:2503.01359
-
How Do Consumers Really Choose: Exposing Hidden Preferences with the Mixture of Experts Model 3 Mar 2025 · 0 repositories · arXiv:2503.05800
-
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering 1 Mar 2025 · 0 repositories · arXiv:2503.00413
-
CoSMoEs: Compact Sparse Mixture of Experts 28 Feb 2025 · 0 repositories · arXiv:2503.00245
-
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts 27 Feb 2025 · 4 repositories · arXiv:2502.19811Syntology official (archive's flag): 6 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 20 harvested samples)
-
Mixture of Experts for Recognizing Depression from Interview and Reading Tasks 27 Feb 2025 · 0 repositories · arXiv:2502.20213
-
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts 27 Feb 2025 · 1 repository · arXiv:2502.20395
-
Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization 26 Feb 2025 · 0 repositories · arXiv:2502.19261Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
BigMac: A Communication-Efficient Mixture-of-Experts Model Structure for Fast Training and Inference 24 Feb 2025 · 0 repositories · arXiv:2502.16927
-
Delta Decompression for MoE-based LLMs Compression 24 Feb 2025 · 1 repository · arXiv:2502.17298Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Evaluating Expert Contributions in a MoE LLM for Quiz-Based Tasks 24 Feb 2025 · 0 repositories · arXiv:2502.17187
-
Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment 24 Feb 2025 · 1 repository · arXiv:2502.16894Syntology official (archive's flag): 6 ran · 6 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
The Empirical Impact of Reducing Symmetries on the Performance of Deep Ensembles and MoE 24 Feb 2025 · 0 repositories · arXiv:2502.17391Syntology 6 ran (of which 3 constructed an object rather than computing a result; 6 with no instrument failure: 3 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Binary-Integer-Programming Based Algorithm for Expert Load Balancing in Mixture-of-Experts Models 21 Feb 2025 · 1 repository · arXiv:2502.15451
-
Tight Clusters Make Specialized Experts 21 Feb 2025 · 1 repository · arXiv:2502.15315
-
DSMoE: Matrix-Partitioned Experts with Dynamic Routing for Computation-Efficient Dense LLMs 18 Feb 2025 · 0 repositories · arXiv:2502.12455
-
Every Expert Matters: Towards Effective Knowledge Distillation for Mixture-of-Experts Language Models 18 Feb 2025 · 0 repositories · arXiv:2502.12947
-
Fate: Fast Edge Inference of Mixture-of-Experts Models via Cross-Layer Gate 17 Feb 2025 · 1 repository · arXiv:2502.12224
-
Probing Semantic Routing in Large Mixture-of-Expert Models 15 Feb 2025 · 0 repositories · arXiv:2502.10928
-
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition 11 Feb 2025 · 0 repositories · arXiv:2502.10447
-
Training Sparse Mixture Of Experts Text Embedding Models 11 Feb 2025 · 1 repository · arXiv:2502.07972Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing 10 Feb 2025 · 0 repositories · arXiv:2502.06643
-
Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline 9 Feb 2025 · 1 repository · arXiv:2502.06888
-
fMoE: Fine-Grained Expert Offloading for Large Mixture-of-Experts Serving 7 Feb 2025 · 0 repositories · arXiv:2502.05370
-
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient 7 Feb 2025 · 0 repositories · arXiv:2502.05172
-
Towards Foundational Models for Dynamical System Reconstruction: Hierarchical Meta-Learning via Mixture of Experts 7 Feb 2025 · 0 repositories · arXiv:2502.05335
-
CMoE: Fast Carving of Mixture-of-Experts for Efficient LLM Inference 6 Feb 2025 · 1 repository · arXiv:2502.04416
-
Mixture of neural operator experts for learning boundary conditions and model selection 6 Feb 2025 · 0 repositories · arXiv:2502.04562
-
Optimizing Robustness and Accuracy in Mixture of Experts: A Dual-Model Approach 5 Feb 2025 · 0 repositories · arXiv:2502.06832
-
CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling 3 Feb 2025 · 0 repositories · arXiv:2502.00965
-
MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs 3 Feb 2025 · 0 repositories · arXiv:2502.00997
-
Omni-Mol: Exploring Universal Convergent Space for Omni-Molecular Tasks 3 Feb 2025 · 0 repositories · arXiv:2502.01074
-
Static Batching of Irregular Workloads on GPUs: Framework and Application to Efficient MoE Model Inference 27 Jan 2025 · 0 repositories · arXiv:2501.16103
-
Each Rank Could be an Expert: Single-Ranked Mixture of Experts LoRA for Multi-Task Learning 25 Jan 2025 · 0 repositories · arXiv:2501.15103
-
Hierarchical Time-Aware Mixture of Experts for Multi-Modal Sequential Recommendation 24 Jan 2025 · 1 repository · arXiv:2501.14269Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Mean-field limit from general mixtures of experts to quantum neural networks 24 Jan 2025 · 0 repositories · arXiv:2501.14660
-
Autonomy-of-Experts Models 22 Jan 2025 · 0 repositories · arXiv:2501.13074
-
BLR-MoE: Boosted Language-Routing Mixture of Experts for Domain-Robust Multilingual E2E ASR 22 Jan 2025 · 0 repositories · arXiv:2501.12602
-
Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models 21 Jan 2025 · 0 repositories · arXiv:2501.11873
-
SCFCRC: Simultaneously Counteract Feature Camouflage and Relation Camouflage for Fraud Detection 21 Jan 2025 · 0 repositories · arXiv:2501.12430
-
FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models 18 Jan 2025 · 0 repositories · arXiv:2501.10714
-
OMoE: Diversifying Mixture of Low-Rank Adaptation by Orthogonal Finetuning 17 Jan 2025 · 0 repositories · arXiv:2501.10062
-
LLM-Based Routing in Mixture of Experts: A Novel Framework for Trading 16 Jan 2025 · 0 repositories · arXiv:2501.09636
-
MoE²: Optimizing Collaborative Inference for Edge Large Language Models 16 Jan 2025 · 0 repositories · arXiv:2501.09410
-
GRAPHMOE: Amplifying Cognitive Depth of Mixture-of-Experts Network via Introducing Self-Rethinking Mechanism 14 Jan 2025 · 0 repositories · arXiv:2501.07890
-
MiniMax-01: Scaling Foundation Models with Lightning Attention 14 Jan 2025 · 1 repository · arXiv:2501.08313Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
A Comprehensive Evaluation of Large Language Models on Mental Illnesses in Arabic Context 12 Jan 2025 · 0 repositories · arXiv:2501.06859
-
Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learning 12 Jan 2025 · 1 repository · arXiv:2501.06884Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
TAMER: A Test-Time Adaptive MoE-Driven Framework for EHR Representation Learning 10 Jan 2025 · 1 repository · arXiv:2501.05661
-
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing 9 Jan 2025 · 0 repositories · arXiv:2501.05313
-
LiMoE: Mixture of LiDAR Representation Learners from Automotive Scenes 7 Jan 2025 · 1 repository · arXiv:2501.04004
-
mFabric: An Efficient and Scalable Fabric for Mixture-of-Experts Training 7 Jan 2025 · 0 repositories · arXiv:2501.03905
-
Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual Learning 4 Jan 2025 · 0 repositories · arXiv:2501.02198
-
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection 30 Dec 2024 · 1 repository · arXiv:2412.20665
-
Multimodal Variational Autoencoder: a Barycentric View 29 Dec 2024 · 0 repositories · arXiv:2412.20487
-
AskChart: Universal Chart Understanding through Textual Enhancement 26 Dec 2024 · 1 repository · arXiv:2412.19146
-
BIG-MoE: Bypass Isolated Gating MoE for Generalized Multimodal Face Anti-Spoofing 24 Dec 2024 · 1 repository · arXiv:2412.18065
-
UME: Upcycling Mixture-of-Experts for Scalable and Efficient Automatic Speech Recognition 23 Dec 2024 · 0 repositories · arXiv:2412.17507
-
Part-Of-Speech Sensitivity of Routers in Mixture of Experts Models 22 Dec 2024 · 0 repositories · arXiv:2412.16971
-
Theory of Mixture-of-Experts for Mobile Edge Computing 20 Dec 2024 · 0 repositories · arXiv:2412.15690
-
ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing 19 Dec 2024 · 1 repository · arXiv:2412.14711Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
A Survey on Inference Optimization Techniques for Mixture of Experts Models 18 Dec 2024 · 1 repository · arXiv:2412.14219
-
GraphLoRA: Empowering LLMs Fine-Tuning via Graph Collaboration of MoE 18 Dec 2024 · 0 repositories · arXiv:2412.16216
-
SEKE: Specialised Experts for Keyword Extraction 18 Dec 2024 · 1 repository · arXiv:2412.14087
-
DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE Inference 16 Dec 2024 · 1 repository · arXiv:2501.10375
-
Enhancing Healthcare Recommendation Systems with a Multimodal LLMs-based MOE Architecture 16 Dec 2024 · 0 repositories · arXiv:2412.11557
-
Investigating Mixture of Experts in Dense Retrieval 16 Dec 2024 · 0 repositories · arXiv:2412.11864
-
Llama 3 Meets MoE: Efficient Upcycling 13 Dec 2024 · 1 repository · arXiv:2412.09952
-
MoSLD: An Extremely Parameter-Efficient Mixture-of-Shared LoRAs for Multi-Task Learning 12 Dec 2024 · 0 repositories · arXiv:2412.08946
-
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine 12 Dec 2024 · 1 repository · arXiv:2412.09278Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems 10 Dec 2024 · 0 repositories · arXiv:2412.07067
-
Post-Training Statistical Calibration for Higher Activation Sparsity 10 Dec 2024 · 1 repository · arXiv:2412.07174Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
Object Detection using Event Camera: A MoE Heat Conduction based Detector and A New Benchmark Dataset 9 Dec 2024 · 1 repository · arXiv:2412.06647
-
Customize Segment Anything Model for Multi-Modal Semantic Segmentation with Mixture of LoRA Experts 5 Dec 2024 · 0 repositories · arXiv:2412.04220
-
Convolutional Neural Networks and Mixture of Experts for Intrusion Detection in 5G Networks and beyond 4 Dec 2024 · 0 repositories · arXiv:2412.03483
-
CA-MoE: Channel-Adapted MoE for Incremental Weather Forecasting 3 Dec 2024 · 0 repositories · arXiv:2412.02503
-
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference 27 Nov 2024 · 0 repositories · arXiv:2412.00099
-
Mixture of Experts in Image Classification: What's the Sweet Spot? 27 Nov 2024 · 0 repositories · arXiv:2411.18322
-
UOE: Unlearning One Expert Is Enough For Mixture-of-experts LLMS 27 Nov 2024 · 0 repositories · arXiv:2411.18797
-
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning 26 Nov 2024 · 1 repository · arXiv:2412.00069
-
MH-MoE: Multi-Head Mixture-of-Experts 25 Nov 2024 · 0 repositories · arXiv:2411.16205
-
LLaMA-MoE v2: Exploring Sparsity of LLaMA from Perspective of Mixture-of-Experts with Post-Training 24 Nov 2024 · 1 repository · arXiv:2411.15708Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
MERLOT: A Distilled LLM-based Mixture-of-Experts Framework for Scalable Encrypted Traffic Classification 20 Nov 2024 · 0 repositories · arXiv:2411.13004
-
MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs 18 Nov 2024 · 0 repositories · arXiv:2411.11217
-
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection 13 Nov 2024 · 0 repositories · arXiv:2411.08982
-
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model 12 Nov 2024 · 0 repositories · arXiv:2411.08212
-
WDMoE: Wireless Distributed Mixture of Experts for Large Language Models 11 Nov 2024 · 0 repositories · arXiv:2411.06681
-
Learning Mixtures of Experts with EM 9 Nov 2024 · 0 repositories · arXiv:2411.06056
-
NeKo: Toward Post Recognition Generative Correction Large Language Models with Task-Oriented Experts 8 Nov 2024 · 0 repositories · arXiv:2411.05945
-
FedMoE-DA: Federated Mixture of Experts via Domain Aware Fine-grained Aggregation 4 Nov 2024 · 0 repositories · arXiv:2411.02115
-
Facet-Aware Multi-Head Mixture-of-Experts Model for Sequential Recommendation 3 Nov 2024 · 0 repositories · arXiv:2411.01457
-
HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference 3 Nov 2024 · 0 repositories · arXiv:2411.01433
-
RS-MoE: Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering 3 Nov 2024 · 0 repositories · arXiv:2411.01595
-
PMoL: Parameter Efficient MoE for Preference Mixing of LLM Alignment 2 Nov 2024 · 0 repositories · arXiv:2411.01245