{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2603-13606","title":"NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL","arxiv_id":"2603.13606","date":"2026-03-13","proceeding":null,"authors":["Amos Goldman","Nimrod Boker","Maayan Sheraizin","Nimrod Admoni","Artem Polyakov","Subhadeep Bhattacharya","Fan Yu","Kai Sun","Georgios Theodorakis","Hsin-Chun Yin","Peter-Jan Gootzen","Aamir Shafi","Assaf Ravid","Salvatore Di Girolamo","James Dinan","Xiaofan Li","Manjunath Gorentla Venkata","Gil Bloch"],"abstract":"Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, driving the development of specialized device-initiated communication libraries such as DeepEP, Hybrid-EP, and others. These libraries demonstrate the performance benefits of GPU-initiated RDMA for MoE dispatch and combine operations. This paper presents NCCL EP (Expert Parallelism), a ground-up MoE communication library built entirely on NCCL's Device API. NCCL EP provides unified ncclEpDispatch and ncclEpCombine primitives with both C and Python interfaces, supporting Low-Latency (LL) mode for inference decoding and High-Throughput (HT) mode for training and inference prefill. LL targets small batch sizes (1-128 tokens) using direct all-to-all RDMA+NVLink mesh connectivity with double-buffered communication for overlapping dispatch and combine phases. HT targets large batches (4096+ tokens) using hierarchical communication that aggregates tokens within NVLink domains before inter-node RDMA transmission. Both modes leverage Device API for both intra- and inter-node communications, taking advantage of its topology awareness and optimized GPU-initiated implementation. We evaluate NCCL EP on an H100-based cluster across multi-node configurations, demonstrating competitive LL kernel performance and presenting end-to-end results with vLLM integration. By building MoE communication natively within NCCL, NCCL EP provides a supported path for expert parallelism on current and emerging NVIDIA platforms.","url_abs":"https://arxiv.org/abs/2603.13606","url_pdf":"https://arxiv.org/pdf/2603.13606","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2603.13606","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2603.13606"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/NVIDIA/nccl","reach":null}],"summary":{"unverified":3},"by_repo_kind":{"found_in_text":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":3,"samples":[{"code_sha256_prefix":"39be2f8f0f14883c","entry":"normalize_type","repo":"NVIDIA/nccl","repo_kind":"found_in_text","path":"contrib/nccl_checkpoint/gen_shim.py","file_url":"https://github.com/NVIDIA/nccl/blob/HEAD/contrib/nccl_checkpoint/gen_shim.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"39be2f8f0f14883c"}},{"code_sha256_prefix":"ff2d1ea234f826e5","entry":"parse_param","repo":"NVIDIA/nccl","repo_kind":"found_in_text","path":"contrib/nccl_checkpoint/gen_shim.py","file_url":"https://github.com/NVIDIA/nccl/blob/HEAD/contrib/nccl_checkpoint/gen_shim.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"ff2d1ea234f826e5"}},{"code_sha256_prefix":"502cb39d09d89676","entry":"strip_comments","repo":"NVIDIA/nccl","repo_kind":"found_in_text","path":"contrib/nccl_checkpoint/gen_shim.py","file_url":"https://github.com/NVIDIA/nccl/blob/HEAD/contrib/nccl_checkpoint/gen_shim.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"502cb39d09d89676"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}