{"url":"/method/chimera","slug":"chimera","name":"Chimera","full_name":"Chimera","full_name_withheld":false,"description_markdown":"**Chimera** is a pipeline model parallelism scheme which combines bidirectional pipelines for efficiently training large-scale models. The key idea of Chimera is to combine two pipelines in different directions (down and up pipelines). \r\n\r\nDenote $N$ as the number of micro-batches executed by each worker within a training iteration, and $D$ the number of pipeline stages (depth), and $P$ the number of workers.\r\n\r\nThe Figure shows an example with four pipeline stages (i.e. $D=4$). Here we assume there are $D$ micro-batches executed by each worker within a training iteration, namely $N=D$, which is the minimum to keep all the stages active. \r\n\r\nIn the down pipeline, stage$\\_{0}$∼stage$\\_{3}$ are mapped to $P\\_{0}∼P\\_{3}$ linearly, while in the up pipeline the stages are mapped in a completely opposite order. The $N$ (assuming an even number) micro-batches are equally partitioned among the two pipelines. Each pipeline schedules $N/2$ micro-batches using 1F1B strategy, as shown in the left part of the Figure. Then, by merging these two pipelines together, we obtain the pipeline schedule of Chimera. Given an even number of stages $D$ (which can be easily satisfied in practice), it is guaranteed that there is no conflict (i.e., there is at most one micro-batch occupies the same time slot on each worker) during merging.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines","paper":"/paper/chimera-efficiently-training-large-scale","first_author":"Shigang Li","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/chimera-efficiently-training-large-scale"},"source":{"url":"https://arxiv.org/abs/2107.06925v3","title":"Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Distributed Methods","url":"/methods/category/distributed-methods","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Synchronous Pipeline Parallel","url":"/methods/category/synchronous-pipeline-parallel","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Model Parallel Methods","url":"/methods/category/model-parallel-methods","pwc_aliases":[]}],"n_papers_tagged":16,"archive_num_papers":16,"papers_newest_first":[{"paper":"/paper/chimera-a-knowledge-base-of-idea","title":"CHIMERA: A Knowledge Base of Idea Recombination in Scientific Literature","date":"2025-05-27","arxiv_id":"2505.20779","n_code_links":1,"syntology":null},{"paper":null,"title":"A CMOS Probabilistic Computing Chip With In-situ hardware Aware Learning","date":"2025-04-18","arxiv_id":"2504.14070","n_code_links":0,"syntology":null},{"paper":null,"title":"Chimera: A Block-Based Neural Architecture Search Framework for Event-Based Object Detection","date":"2024-12-27","arxiv_id":"2412.19646","n_code_links":0,"syntology":null},{"paper":null,"title":"Language model driven: a PROTAC generation pipeline with dual constraints of structure and property","date":"2024-12-12","arxiv_id":"2412.09661","n_code_links":0,"syntology":null},{"paper":null,"title":"Chimera: Improving Generalist Model with Domain-Specific Experts","date":"2024-12-08","arxiv_id":"2412.05983","n_code_links":0,"syntology":null},{"paper":null,"title":"Chimera: Accurate retrosynthesis prediction by ensembling models with diverse inductive biases","date":"2024-12-06","arxiv_id":"2412.05269","n_code_links":0,"syntology":null},{"paper":null,"title":"Chimera: Effectively Modeling Multivariate Time Series with 2-Dimensional State Space Models","date":"2024-06-06","arxiv_id":"2406.04320","n_code_links":0,"syntology":null},{"paper":null,"title":"Chimera: A Lossless Decoding Method for Accelerating Large Language Models Inference by Fusing all Tokens","date":"2024-02-24","arxiv_id":"2402.15758","n_code_links":0,"syntology":null},{"paper":null,"title":"Artificial Bee Colony optimization of Deep Convolutional Neural Networks in the context of Biomedical Imaging","date":"2024-02-23","arxiv_id":"2402.15246","n_code_links":0,"syntology":null},{"paper":"/paper/spoofing-resilient-lidar-gps-factor-graph","title":"Spoofing-Resilient LiDAR-GPS Factor Graph Localization with Chimera Authentication","date":"2023-07-10","arxiv_id":"2307.04692","n_code_links":1,"syntology":null},{"paper":null,"title":"Investigating the generative dynamics of energy-based neural networks","date":"2023-05-11","arxiv_id":"2305.06745","n_code_links":0,"syntology":null},{"paper":null,"title":"Chimera: A Hybrid Machine Learning Driven Multi-Objective Design Space Exploration Tool for FPGA High-Level Synthesis","date":"2022-07-03","arxiv_id":"2207.07917","n_code_links":0,"syntology":null},{"paper":null,"title":"Physics-inspired Ising Computing with Ring Oscillator Activated p-bits","date":"2022-05-15","arxiv_id":"2205.07402","n_code_links":0,"syntology":null},{"paper":null,"title":"Complex dynamics of a heterogeneous network of Hindmarsh-Rose neurons","date":"2022-05-03","arxiv_id":"2205.01790","n_code_links":0,"syntology":null},{"paper":null,"title":"Convex Non-negative Matrix Factorization Through Quantum Annealing","date":"2022-03-28","arxiv_id":"2203.15634","n_code_links":0,"syntology":null},{"paper":"/paper/chimera-efficiently-training-large-scale","title":"Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines","date":"2021-07-14","arxiv_id":"2107.06925","n_code_links":1,"syntology":null}],"papers_shown":16,"tasks":[{"task":"/task/state-space-models","name":"State Space Models","papers":2},{"task":"/task/active-learning","name":"Active Learning","papers":1},{"task":"/task/all","name":"All","papers":1},{"task":"/task/anomaly-detection","name":"Anomaly Detection","papers":1},{"task":"/task/descriptive","name":"Descriptive","papers":1},{"task":"/task/drug-discovery","name":"Drug Discovery","papers":1},{"task":null,"name":"GPU","papers":1},{"task":"/task/high-level-synthesis","name":"High-Level Synthesis","papers":1},{"task":"/task/human-detection","name":"Human Detection","papers":1},{"task":"/task/hybrid-machine-learning","name":"Hybrid Machine Learning","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/mamba","name":"Mamba","papers":1},{"task":"/task/math","name":"Math","papers":1},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":null,"name":"Position","papers":1},{"task":"/task/property-prediction","name":"Property Prediction","papers":1},{"task":"/task/retrosynthesis","name":"Retrosynthesis","papers":1},{"task":"/task/scheduling","name":"Scheduling","papers":1}],"tasks_shown":20,"n_tasks":29,"usage_by_year":[{"year":"2021","papers":1},{"year":"2022","papers":4},{"year":"2023","papers":2},{"year":"2024","papers":7},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/chimera"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}