Browse State-of-the-Art › Mixture-of-Experts › Papers, page 6
Mixture-of-Experts
Papers archive 2025-07-28
archive papers tagged: 1,312 · with a code link: 516 · where Syntology ran a sample: 216 (184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (216 of 1,312 tagged: 184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument)
Page 6 of 14: papers 501 to 600 of 1,312, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
10 Jul 2019 1 repository listed
-
24 May 2019 1 repository listed
-
28 Feb 2019 1 repository listed
-
3 Dec 2018 1 repository listed
-
8 Oct 2018 1 repository listed
-
12 Sep 2018 1 repository listed
-
7 Sep 2018 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
5 Jun 2018 1 repository listed Syntology 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
7 Mar 2018 1 repository listed
-
6 Feb 2018 1 repository listed
-
12 Sep 2017 1 repository listed
-
13 Jul 2017 1 repository listed
-
11 Jul 2017 1 repository listed
-
8 Jul 2017 1 repository listed
-
27 Feb 2017 1 repository listed
-
18 Sep 2016 1 repository listed
-
GEMINUS: Dual-aware Global and Scene-Adaptive Mixture-of-Experts for End-to-End Autonomous Driving19 Jul 2025 0 repositories listed
-
R^2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning17 Jul 2025 0 repositories listed
-
Mixture of Experts in Large Language Models15 Jul 2025 0 repositories listed
-
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive13 Jul 2025 0 repositories listed
-
KAT-V1: Kwai-AutoThink Technical Report11 Jul 2025 0 repositories listed
-
Efficient Training of Large-Scale AI Models Through Federated Mixture-of-Experts: A System-Level Approach8 Jul 2025 0 repositories listed
-
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis8 Jul 2025 0 repositories listed
-
What You Have is What You Track: Adaptive and Robust Multimodal Tracking8 Jul 2025 0 repositories listed
-
UGG-ReID: Uncertainty-Guided Graph Model for Multi-Modal Object Re-Identification7 Jul 2025 0 repositories listed
-
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging29 Jun 2025 0 repositories listed
-
EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning26 Jun 2025 0 repositories listed
-
Little By Little: Continual Learning via Self-Activated Sparse Mixture-of-Rank Adaptive Learning26 Jun 2025 0 repositories listed
-
Opportunistic Osteoporosis Diagnosis via Texture-Preserving Self-Supervision, Mixture of Experts and Multi-Task Integration25 Jun 2025 0 repositories listed
-
An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify23 Jun 2025 0 repositories listed
-
Security Assessment of DeepSeek and GPT Series Models against Jailbreak Attacks23 Jun 2025 0 repositories listed
-
SAFEx: Analyzing Vulnerabilities of MoE-Based LLMs via Stable Safety-critical Expert Identification20 Jun 2025 0 repositories listed
-
Exploring Speaker Diarization with Mixture of Experts17 Jun 2025 0 repositories listed
-
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing17 Jun 2025 0 repositories listed
-
MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models17 Jun 2025 0 repositories listed
-
NeuroMoE: A Transformer-Based Mixture-of-Experts Framework for Multi-Modal Neurological Disorder Classification17 Jun 2025 0 repositories listed
-
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs17 Jun 2025 0 repositories listed
-
Scaling Intelligence: Designing Data Centers for Next-Gen Language Models17 Jun 2025 0 repositories listed
-
Single-Example Learning in a Mixture of GPDMs with Latent Geometries17 Jun 2025 0 repositories listed
-
Utility-Driven Speculative Decoding for Mixture-of-Experts17 Jun 2025 0 repositories listed
-
EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware Optimization16 Jun 2025 0 repositories listed
-
Load Balancing Mixture of Experts with Similarity Preserving Routers16 Jun 2025 0 repositories listed
-
Serving Large Language Models on Huawei CloudMatrix38415 Jun 2025 0 repositories listed
-
Optimus-3: Towards Generalist Multimodal Minecraft Agents with Scalable Task Experts12 Jun 2025 0 repositories listed
-
GigaChat Family: Efficient Russian Language Modeling Through Mixture of Experts Architecture11 Jun 2025 0 repositories listed
-
MedMoE: Modality-Specialized Mixture of Experts for Medical Vision-Language Understanding10 Jun 2025 0 repositories listed
-
M2Restore: Mixture-of-Experts-based Mamba-CNN Fusion Framework for All-in-One Image Restoration9 Jun 2025 0 repositories listed
-
MIRA: Medical Time Series Foundation Model for Real-World Health Data9 Jun 2025 0 repositories listed
-
MoE-GPS: Guidlines for Prediction Strategy for Dynamic Expert Duplication in MoE Load Balancing9 Jun 2025 0 repositories listed
-
Breaking Data Silos: Towards Open and Scalable Mobility Foundation Models via Generative Continual Learning7 Jun 2025 0 repositories listed
-
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities6 Jun 2025 0 repositories listed
-
Lifelong Evolution: Collaborative Learning between Large and Small Language Models for Continuous Emergent Fake News Detection5 Jun 2025 0 repositories listed
-
Brain-Like Processing Pathways Form in Models With Heterogeneous Experts3 Jun 2025 0 repositories listed
-
Enhancing Multimodal Continual Instruction Tuning with BranchLoRA31 May 2025 0 repositories listed
-
Decoding Knowledge Attribution in Mixture-of-Experts: A Framework of Basic-Refinement Collaboration and Efficiency Analysis30 May 2025 0 repositories listed
-
GradPower: Powering Gradients for Faster Language Model Pre-Training30 May 2025 0 repositories listed
-
Mixture-of-Experts for Personalized and Semantic-Aware Next Location Prediction30 May 2025 0 repositories listed
-
On the Expressive Power of Mixture-of-Experts for Structured Complex Tasks30 May 2025 0 repositories listed
-
A Survey of Generative Categories and Techniques in Multimodal Large Language Models29 May 2025 0 repositories listed
-
Noise-Robustness Through Noise: Asymmetric LoRA Adaption with Poisoning Expert29 May 2025 0 repositories listed
-
Point-MoE: Towards Cross-Domain Generalization in 3D Semantic Segmentation via Mixture-of-Experts29 May 2025 0 repositories listed
-
Revisiting Uncertainty Estimation and Calibration of Large Language Models29 May 2025 0 repositories listed
-
Two Is Better Than One: Rotations Scale LoRAs29 May 2025 0 repositories listed
-
Advancing Expert Specialization for Better MoE28 May 2025 0 repositories listed
-
EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models28 May 2025 0 repositories listed
-
ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation28 May 2025 0 repositories listed
-
MoE-Gyro: Self-Supervised Over-Range Reconstruction and Denoising for MEMS Gyroscopes27 May 2025 0 repositories listed
-
MoESD: Unveil Speculative Decoding's Potential for Accelerating Sparse MoE26 May 2025 0 repositories listed
-
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments26 May 2025 0 repositories listed
-
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID26 May 2025 0 repositories listed
-
Integrating Dynamical Systems Learning with Foundational Models: A Meta-Evolutionary AI Framework for Clinical Trials25 May 2025 0 repositories listed
-
μ-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts24 May 2025 0 repositories listed
-
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter24 May 2025 0 repositories listed
-
On Minimax Estimation of Parameters in Softmax-Contaminated Mixture of Experts24 May 2025 0 repositories listed
-
TrajMoE: Spatially-Aware Mixture of Experts for Unified Human Mobility Modeling24 May 2025 0 repositories listed
-
EvidenceMoE: A Physics-Guided Mixture-of-Experts with Evidential Critics for Advancing Fluorescence Light Detection and Ranging in Scattering Media23 May 2025 0 repositories listed
-
22 May 2025 0 repositories listed
-
DualComp: End-to-End Learning of a Unified Dual-Modality Lossless Compressor22 May 2025 0 repositories listed
-
Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks21 May 2025 0 repositories listed
-
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought21 May 2025 0 repositories listed
-
21 May 2025 0 repositories listed Syntology 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 4 harvested samples) · 4 pointer-only (licence)
-
Time Tracker: Mixture-of-Experts-Enhanced Foundation Time Series Forecasting Model with Decoupled Training Pipelines21 May 2025 0 repositories listed
-
Balanced and Elastic End-to-end Training of Dynamic LLMs20 May 2025 0 repositories listed
-
EfficientLLM: Efficiency in Large Language Models20 May 2025 0 repositories listed
-
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation20 May 2025 0 repositories listed
-
Multimodal Mixture of Low-Rank Experts for Sentiment Analysis and Emotion Recognition20 May 2025 0 repositories listed
-
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach20 May 2025 0 repositories listed
-
StPR: Spatiotemporal Preservation and Routing for Exemplar-Free Video Class-Incremental Learning20 May 2025 0 repositories listed
-
THOR-MoE: Hierarchical Task-Guided and Context-Responsive Routing for Neural Machine Translation20 May 2025 0 repositories listed
-
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training20 May 2025 0 repositories listed
-
Model Selection for Gaussian-gated Gaussian Mixture of Experts Using Dendrograms of Mixing Measures19 May 2025 0 repositories listed
-
Seeing the Unseen: How EMoE Unveils Bias in Text-to-Image Diffusion Models19 May 2025 0 repositories listed
-
True Zero-Shot Inference of Dynamical Systems Preserving Long-Term Statistics19 May 2025 0 repositories listed
-
Improving Coverage in Combined Prediction Sets with Weighted p-values17 May 2025 0 repositories listed
-
MINGLE: Mixtures of Null-Space Gated Low-Rank Experts for Test-Time Continual Model Merging17 May 2025 0 repositories listed
-
Model Merging in Pre-training of Large Language Models17 May 2025 0 repositories listed
-
On DeepSeekMoE: Statistical Benefits of Shared Experts and Normalized Sigmoid Gating16 May 2025 0 repositories listed
-
A Fast Kernel-based Conditional Independence test with Application to Causal Discovery16 May 2025 0 repositories listed
-
MoE-CAP: Benchmarking Cost, Accuracy and Performance of Sparse Mixture-of-Experts Systems16 May 2025 0 repositories listed
-
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production16 May 2025 0 repositories listed
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.