Browse State-of-the-Art › Mixture-of-Experts › Papers, page 10
Mixture-of-Experts
Papers archive 2025-07-28
archive papers tagged: 1,312 · with a code link: 516 · where Syntology ran a sample: 216 (184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (216 of 1,312 tagged: 184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument)
Page 10 of 14: papers 901 to 1,000 of 1,312, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Interpretable mixture of experts for time series prediction under recurrent and non-recurrent conditions5 Sep 2024 0 repositories listed
-
Configurable Foundation Models: Building LLMs from a Modular Perspective4 Sep 2024 0 repositories listed
-
Pluralistic Salient Object Detection4 Sep 2024 0 repositories listed
-
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model3 Sep 2024 0 repositories listed
-
Beyond Parameter Count: Implicit Bias in Soft Mixture of Experts2 Sep 2024 0 repositories listed
-
Duplex: A Device for Large Language Models with Mixture of Experts, Grouped Query Attention, and Continuous Batching2 Sep 2024 0 repositories listed
-
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts28 Aug 2024 0 repositories listed
-
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts28 Aug 2024 0 repositories listed
-
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis27 Aug 2024 0 repositories listed
-
Advancing Enterprise Spatio-Temporal Forecasting Applications: Data Mining Meets Instruction Tuning of Language Models For Multi-modal Time Series Analysis in Low-Resource Settings24 Aug 2024 0 repositories listed
-
La-SoftMoE CLIP for Unified Physical-Digital Face Attack Detection23 Aug 2024 0 repositories listed
-
Multi-Treatment Multi-Task Uplift Modeling for Enhancing User Growth23 Aug 2024 0 repositories listed
-
The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities23 Aug 2024 0 repositories listed
-
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging22 Aug 2024 0 repositories listed
-
FedMoE: Personalized Federated Learning via Heterogeneous Mixture of Experts21 Aug 2024 0 repositories listed
-
HMoE: Heterogeneous Mixture of Experts for Language Modeling20 Aug 2024 0 repositories listed
-
A Unified Framework for Iris Anti-Spoofing: Introducing IrisGeneral Dataset and Masked-MoE Method19 Aug 2024 0 repositories listed
-
BAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of Experts15 Aug 2024 0 repositories listed
-
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning13 Aug 2024 0 repositories listed
-
HoME: Hierarchy of Multi-Gate Experts for Multi-Task Learning at Kuaishou10 Aug 2024 0 repositories listed
-
LaDiMo: Layer-wise Distillation Inspired MoEfier8 Aug 2024 0 repositories listed
-
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training8 Aug 2024 0 repositories listed
-
Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization5 Aug 2024 0 repositories listed
-
HMDN: Hierarchical Multi-Distribution Network for Click-Through Rate Prediction2 Aug 2024 0 repositories listed
-
Multimodal Fusion and Coherence Modeling for Video Topic Segmentation1 Aug 2024 0 repositories listed
-
PMoE: Progressive Mixture of Experts with Asymmetric Transformer for Continual Learning31 Jul 2024 0 repositories listed
-
MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts31 Jul 2024 0 repositories listed
-
Distribution Learning for Molecular Regression30 Jul 2024 0 repositories listed
-
Time series forecasting with high stakes: A field study of the air cargo industry29 Jul 2024 0 repositories listed
-
Wolf: Captioning Everything with a World Summarization Framework26 Jul 2024 0 repositories listed
-
How Lightweight Can A Vision Transformer Be25 Jul 2024 0 repositories listed
-
Wonderful Matrices: More Efficient and Effective Architecture for Language Modeling Tasks24 Jul 2024 0 repositories listed
-
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism24 Jul 2024 0 repositories listed
-
EEGMamba: Bidirectional State Space Model with Mixture of Experts for EEG Multi-task Classification20 Jul 2024 0 repositories listed
-
EVLM: An Efficient Vision-Language Model for Visual Understanding19 Jul 2024 0 repositories listed
-
Mixture of Experts with Mixture of Precisions for Tuning Quality of Service19 Jul 2024 0 repositories listed
-
Discussion: Effective and Interpretable Outcome Prediction by Training Sparse Mixtures of Linear Experts18 Jul 2024 0 repositories listed
-
Mixture of Experts based Multi-task Supervise Learning from Crowds18 Jul 2024 0 repositories listed
-
Boost Your NeRF: A Model-Agnostic Mixture of Experts Framework for High Quality and Efficient Rendering15 Jul 2024 0 repositories listed
-
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts12 Jul 2024 0 repositories listed
-
An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio11 Jul 2024 0 repositories listed
-
A Simple Architecture for Enterprise Large Language Model Applications based on Role based security and Clearance Levels using Retrieval-Augmented Generation or Mixture of Experts9 Jul 2024 0 repositories listed
-
SAM-Med3D-MoE: Towards a Non-Forgetting Segment Anything Model via Mixture of Experts for 3D Medical Image Segmentation6 Jul 2024 0 repositories listed
-
Lazarus: Resilient and Elastic Training of Mixture-of-Experts Models with Adaptive Expert Placement5 Jul 2024 0 repositories listed
-
MobileFlow: A Multimodal LLM For Mobile GUI Agent5 Jul 2024 0 repositories listed
-
Terminating Differentiable Tree Experts2 Jul 2024 0 repositories listed
-
Investigating the potential of Sparse Mixtures-of-Experts for multi-domain neural machine translation1 Jul 2024 0 repositories listed
-
Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning1 Jul 2024 0 repositories listed
-
LEMoE: Advanced Mixture of Experts Adaptor for Lifelong Model Editing of Large Language Models28 Jun 2024 0 repositories listed
-
Towards Personalized Federated Multi-Scenario Multi-Task Recommendation27 Jun 2024 0 repositories listed
-
Mixture of Experts in a Mixture of RL settings26 Jun 2024 0 repositories listed
-
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR26 Jun 2024 0 repositories listed
-
MoESD: Mixture of Experts Stable Diffusion to Mitigate Gender Bias25 Jun 2024 0 repositories listed
-
Theory on Mixture-of-Experts in Continual Learning24 Jun 2024 0 repositories listed
-
SimSMoE: Solving Representational Collapse via Similarity Measure22 Jun 2024 0 repositories listed
-
Low-Rank Mixture-of-Experts for Continual Medical Image Segmentation19 Jun 2024 0 repositories listed
-
P-Tailor: Customizing Personality Traits for Language Models via Mixture of Specialized LoRA Experts18 Jun 2024 0 repositories listed
-
Variational Distillation of Diffusion Policies into Mixture of Experts18 Jun 2024 0 repositories listed
-
Interpretable Cascading Mixture-of-Experts for Urban Traffic Congestion Prediction14 Jun 2024 0 repositories listed
-
Continual Traffic Forecasting via Mixture of Experts5 Jun 2024 0 repositories listed
-
Filtered not Mixed: Stochastic Filtering-Based Online Gating for Mixture of Large Language Models5 Jun 2024 0 repositories listed
-
Node-wise Filtering in Graph Neural Networks: A Mixture of Experts Approach5 Jun 2024 0 repositories listed
-
Style Mixture of Experts for Expressive Text-To-Speech Synthesis5 Jun 2024 0 repositories listed
-
Optimizing 6G Integrated Sensing and Communications (ISAC) via Expert Networks1 Jun 2024 0 repositories listed
-
Training-efficient density quantum machine learning30 May 2024 0 repositories listed
-
MEMoE: Enhancing Model Editing with Mixture of Experts Adaptors29 May 2024 0 repositories listed
-
MoNDE: Mixture of Near-Data Experts for Large-Scale Sparse Models29 May 2024 0 repositories listed
-
LoRA-Switch: Boosting the Efficiency of Dynamic LLM Adapters via System-Algorithm Co-design28 May 2024 0 repositories listed
-
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts26 May 2024 0 repositories listed
-
Expert-Token Resonance: Redefining MoE Routing through Affinity-Driven Active Selection24 May 2024 0 repositories listed
-
Statistical Advantages of Perturbing Cosine Router in Mixture of Experts23 May 2024 0 repositories listed
-
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts22 May 2024 0 repositories listed
-
Learning More Generalized Experts by Merging Experts in Mixture-of-Experts19 May 2024 0 repositories listed
-
Many Hands Make Light Work: Task-Oriented Dialogue System with Module-Based Mixture-of-Experts16 May 2024 0 repositories listed
-
A Mixture-of-Experts Approach to Few-Shot Task Transfer in Open-Ended Text Worlds9 May 2024 0 repositories listed
-
SUTRA: Scalable Multilingual Language Model Architecture7 May 2024 0 repositories listed
-
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training6 May 2024 0 repositories listed
-
MEET: Mixture of Experts Extra Tree-Based sEMG Hand Gesture Identification6 May 2024 0 repositories listed
-
WDMoE: Wireless Distributed Large Language Models with Mixture of Experts6 May 2024 0 repositories listed
-
Mixture of partially linear experts5 May 2024 0 repositories listed
-
Hierarchical mixture of discriminative Generalized Dirichlet classifiers2 May 2024 0 repositories listed
-
Mixture of insighTful Experts (MoTE): The Synergy of Thought Chains and Expert Mixtures in Self-Alignment1 May 2024 0 repositories listed
-
MoPEFT: A Mixture-of-PEFTs for the Segment Anything Model1 May 2024 0 repositories listed
-
Powering In-Database Dynamic Model Slicing for Structured Data Analytics1 May 2024 0 repositories listed
-
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping30 Apr 2024 0 repositories listed
-
Mix of Experts Language Model for Named Entity Recognition30 Apr 2024 0 repositories listed
-
Towards Incremental Learning in Large Language Models: A Critical Review28 Apr 2024 0 repositories listed
-
Integration of Mixture of Experts and Multimodal Generative AI in Internet of Vehicles: A Survey25 Apr 2024 0 repositories listed
-
U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF25 Apr 2024 0 repositories listed
-
A Novel A.I Enhanced Reservoir Characterization with a Combined Mixture of Experts -- NVIDIA Modulus based Physics Informed Neural Operator Forward Model20 Apr 2024 0 repositories listed
-
A Large-scale Medical Visual Task Adaptation Benchmark19 Apr 2024 0 repositories listed
-
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation17 Apr 2024 0 repositories listed
-
Generative AI Agents with Large Language Model for Satellite Networks via a Mixture of Experts Transmission14 Apr 2024 0 repositories listed
-
Intuition-aware Mixture-of-Rank-1-Experts for Parameter Efficient Finetuning13 Apr 2024 0 repositories listed
-
Mixture of Experts Soften the Curse of Dimensionality in Operator Learning13 Apr 2024 0 repositories listed
-
Identifying Shopping Intent in Product QA for Proactive Recommendations9 Apr 2024 0 repositories listed
-
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models8 Apr 2024 0 repositories listed
-
SEER-MoE: Sparse Expert Efficiency through Regularization for Mixture-of-Experts7 Apr 2024 0 repositories listed
-
Shortcut-connected Expert Parallelism for Accelerating Mixture-of-Experts7 Apr 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.