Browse State-of-the-Art › Mixture-of-Experts › Papers, page 7
Mixture-of-Experts
Papers archive 2025-07-28
archive papers tagged: 1,312 · with a code link: 516 · where Syntology ran a sample: 216 (184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (216 of 1,312 tagged: 184 with a run with no instrument failure, 32 where every run was a failure of Syntology's instrument)
Page 7 of 14: papers 601 to 700 of 1,312, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures14 May 2025 0 repositories listed
-
AM-Thinking-v1: Advancing the Frontier of Reasoning at 32B Scale13 May 2025 0 repositories listed
-
PWC-MoE: Privacy-Aware Wireless Collaborative Mixture of Experts13 May 2025 0 repositories listed
-
UMoE: Unifying Attention and FFN with Shared Experts12 May 2025 0 repositories listed
-
FreqMoE: Dynamic Frequency Enhancement for Neural PDE Solvers11 May 2025 0 repositories listed
-
11 May 2025 0 repositories listed
-
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts11 May 2025 0 repositories listed
-
QoS-Efficient Serving of Multiple Mixture-of-Expert LLMs Using Partial Runtime Reconfiguration10 May 2025 0 repositories listed
-
FloE: On-the-Fly MoE Inference on Memory-constrained GPU9 May 2025 0 repositories listed
-
Divide-and-Conquer: Cold-Start Bundle Recommendation via Mixture of Diffusion Experts8 May 2025 0 repositories listed
-
Pangu Ultra MoE: How to Train Your Big MoE on Ascend NPUs7 May 2025 0 repositories listed
-
SToLa: Self-Adaptive Touch-Language Framework with Tactile Commonsense Reasoning in Open-Ended Scenarios7 May 2025 0 repositories listed
-
3D Gaussian Splatting Data Compression with Mixture of Priors6 May 2025 0 repositories listed
-
Faster MoE LLM Inference for Extremely Large Models6 May 2025 0 repositories listed
-
STAR-Rec: Making Peace with Length Variance and Pattern Diversity in Sequential Recommendation6 May 2025 0 repositories listed
-
Towards Smart Point-and-Shoot Photography6 May 2025 0 repositories listed
-
Multimodal Deep Learning-Empowered Beam Prediction in Future THz ISAC Systems5 May 2025 0 repositories listed
-
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques5 May 2025 0 repositories listed
-
CoCoAFusE: Beyond Mixtures of Experts via Model Fusion2 May 2025 0 repositories listed
-
Perception-Informed Neural Networks: Beyond Physics-Informed Neural Networks2 May 2025 0 repositories listed
-
CICADA: Cross-Domain Interpretable Coding for Anomaly Detection and Adaptation in Multivariate Time Series1 May 2025 0 repositories listed
-
MoxE: Mixture of xLSTM Experts with Entropy-Aware Routing for Efficient Language Modeling1 May 2025 0 repositories listed
-
Accelerating Mixture-of-Experts Training with Adaptive Expert Replication28 Apr 2025 0 repositories listed
-
PICO: Secure Transformers via Robust Prompt Isolation and Cybersecurity Oversight26 Apr 2025 0 repositories listed
-
BadMoE: Backdooring Mixture-of-Experts LLMs via Optimizing Routing Triggers and Infecting Dormant Experts24 Apr 2025 0 repositories listed
-
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core21 Apr 2025 0 repositories listed
-
HAECcity: Open-Vocabulary Scene Understanding of City-Scale Point Clouds with Superpoint Graph Clustering18 Apr 2025 0 repositories listed
-
Multi-Type Context-Aware Conversational Recommender Systems via Mixture-of-Experts18 Apr 2025 0 repositories listed
-
D²MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving17 Apr 2025 0 repositories listed
-
Trend Filtered Mixture of Experts for Automated Gating of High-Frequency Flow Cytometry Data16 Apr 2025 0 repositories listed
-
Unveiling Hidden Collaboration within Mixture-of-Experts in Large Language Models16 Apr 2025 0 repositories listed
-
14 Apr 2025 0 repositories listed Syntology 4 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Mixture-of-Shape-Experts (MoSE): End-to-End Shape Dictionary Framework to Prompt SAM for Generalizable Medical Segmentation13 Apr 2025 0 repositories listed
-
MoE-Lens: Towards the Hardware Limit of High-Throughput MoE LLM Serving Under Resource Constraints12 Apr 2025 0 repositories listed
-
Regularized infill criteria for multi-objective Bayesian optimization with application to aircraft design11 Apr 2025 0 repositories listed
-
RouterKT: Mixture-of-Experts for Knowledge Tracing11 Apr 2025 0 repositories listed
-
Adaptive Detection of Fast Moving Celestial Objects Using a Mixture of Experts and Physical-Inspired Neural Network10 Apr 2025 0 repositories listed
-
Scaling Laws for Native Multimodal Models Scaling Laws for Native Multimodal Models10 Apr 2025 0 repositories listed
-
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning10 Apr 2025 0 repositories listed
-
FedMerge: Federated Personalization via Model Merging9 Apr 2025 0 repositories listed
-
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models9 Apr 2025 0 repositories listed
-
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations8 Apr 2025 0 repositories listed
-
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs4 Apr 2025 0 repositories listed
-
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation4 Apr 2025 0 repositories listed
-
MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism3 Apr 2025 0 repositories listed
-
Advancing MoE Efficiency: A Collaboration-Constrained Routing (C2R) Strategy for Better Expert Parallelism Design2 Apr 2025 0 repositories listed
-
A Unified Virtual Mixture-of-Experts Framework:Enhanced Inference and Hallucination Mitigation in Single-Model System1 Apr 2025 0 repositories listed
-
Detecting Financial Fraud with Hybrid Deep Learning: A Mix-of-Experts Approach to Sequential and Anomalous Patterns1 Apr 2025 0 repositories listed
-
Unimodal-driven Distillation in Multimodal Emotion Recognition with Dynamic Fusion31 Mar 2025 0 repositories listed
-
Mixture of Routers30 Mar 2025 0 repositories listed
-
Beyond Standard MoE: Mixture of Latent Experts for Resource-Efficient Language Models29 Mar 2025 0 repositories listed
-
S2MoE: Robust Sparse Mixture of Experts via Stochastic Learning29 Mar 2025 0 repositories listed
-
Sparse Mixture of Experts as Unified Competitive Learning29 Mar 2025 0 repositories listed
-
Exploiting Mixture-of-Experts Redundancy Unlocks Multimodal Generative Abilities28 Mar 2025 0 repositories listed
-
iMedImage Technical Report27 Mar 2025 0 repositories listed
-
LLaVA-CMoE: Towards Continual Mixture of Experts for Large Vision-Language Models27 Mar 2025 0 repositories listed
-
RocketPPA: Code-Level Power, Performance, and Area Prediction via LLM and Mixture of Experts27 Mar 2025 0 repositories listed
-
Enhancing Multi-modal Models with Heterogeneous MoE Adapters for Fine-tuning26 Mar 2025 0 repositories listed
-
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation26 Mar 2025 0 repositories listed
-
Optimal Scaling Laws for Efficiency Gains in a Theoretical Transformer-Augmented Sectional MoE Framework26 Mar 2025 0 repositories listed
-
Reasoning Beyond Limits: Advances and Open Problems for LLMs26 Mar 2025 0 repositories listed
-
BiPrompt-SAM: Enhancing Image Segmentation via Explicit Selection between Point and Text Prompts25 Mar 2025 0 repositories listed
-
M²CD: A Unified MultiModal Framework for Optical-SAR Change Detection with Mixture of Experts and Self-Distillation25 Mar 2025 0 repositories listed
-
Resilient Sensor Fusion under Adverse Sensor Failures via Multi-Modal Expert Fusion25 Mar 2025 0 repositories listed
-
Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding24 Mar 2025 0 repositories listed
-
ExpertRAG: Efficient RAG with Mixture of Experts -- Optimizing Context Retrieval for Adaptive LLM Responses23 Mar 2025 0 repositories listed
-
Every Sample Matters: Leveraging Mixture-of-Experts and High-Quality Data for Efficient and Accurate Code LLM22 Mar 2025 0 repositories listed
-
Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts20 Mar 2025 0 repositories listed
-
UniCoRN: Latent Diffusion-based Unified Controllable Image Restoration Network across Multiple Degradations20 Mar 2025 0 repositories listed
-
Leveraging MoE-based Large Language Model for Zero-Shot Multi-Task Semantic Communication19 Mar 2025 0 repositories listed
-
SemEval-2025 Task 1: AdMIRe -- Advancing Multimodal Idiomaticity Representation19 Mar 2025 0 repositories listed
-
Core-Periphery Principle Guided State Space Model for Functional Connectome Classification18 Mar 2025 0 repositories listed
-
MAST-Pro: Dynamic Mixture-of-Experts for Adaptive Segmentation of Pan-Tumors with Knowledge-Driven Prompts18 Mar 2025 0 repositories listed
-
Adaptive Mixture of Low-Rank Experts for Robust Audio Spoofing Detection15 Mar 2025 0 repositories listed
-
A Review of DeepSeek Models' Key Innovative Techniques14 Mar 2025 0 repositories listed
-
dFLMoE: Decentralized Federated Learning via Mixture of Experts for Medical Data Analysis13 Mar 2025 0 repositories listed
-
Ensemble Learning for Large Language Models in Text and Code Generation: A Survey13 Mar 2025 0 repositories listed
-
Astrea: A MOE-based Visual Understanding Model with Progressive Alignment12 Mar 2025 0 repositories listed
-
Automatic Operator-level Parallelism Planning for Distributed Deep Learning -- A Mixed-Integer Programming Approach12 Mar 2025 0 repositories listed
-
Double-Stage Feature-Level Clustering-Based Mixture of Experts Framework12 Mar 2025 0 repositories listed
-
FaVChat: Unlocking Fine-Grained Facail Video Understanding with Multimodal Large Language Models12 Mar 2025 0 repositories listed
-
Priority-Aware Preemptive Scheduling for Mixed-Priority Workloads in MoE Inference12 Mar 2025 0 repositories listed
-
Accelerating MoE Model Inference with Expert Sharding11 Mar 2025 0 repositories listed
-
MoE-Loco: Mixture of Experts for Multitask Locomotion11 Mar 2025 0 repositories listed
-
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models11 Mar 2025 0 repositories listed
-
UniF²ace: Fine-grained Face Understanding and Generation with Unified Multimodal Models11 Mar 2025 0 repositories listed
-
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference10 Mar 2025 0 repositories listed
-
GM-MoE: Low-Light Enhancement with Gated-Mechanism Mixture-of-Experts10 Mar 2025 0 repositories listed
-
MoFE: Mixture of Frozen Experts Architecture9 Mar 2025 0 repositories listed
-
A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts8 Mar 2025 0 repositories listed
-
MANDARIN: Mixture-of-Experts Framework for Dynamic Delirium and Coma Prediction in ICU Patients: Development and Validation of an Acute Brain Dysfunction Prediction Model8 Mar 2025 0 repositories listed
-
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering8 Mar 2025 0 repositories listed
-
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts7 Mar 2025 0 repositories listed
-
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs7 Mar 2025 0 repositories listed
-
FMT:A Multimodal Pneumonia Detection Model Based on Stacking MOE Framework7 Mar 2025 0 repositories listed
-
Symbolic Mixture-of-Experts: Adaptive Skill-based Routing for Heterogeneous Reasoning7 Mar 2025 0 repositories listed
-
A Generalist Cross-Domain Molecular Learning Framework for Structure-Based Drug Discovery6 Mar 2025 0 repositories listed
-
Continual Pre-training of MoEs: How robust is your router?6 Mar 2025 0 repositories listed
-
Predictable Scale: Part I -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining6 Mar 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.