{"url":"/method/i-bert","slug":"i-bert","name":"I-BERT","full_name":"I-BERT","full_name_withheld":false,"description_markdown":"**I-BERT** is a quantized version of [BERT](https://paperswithcode.com/method/bert) that quantizes the entire inference with integer-only arithmetic. Based on lightweight integer only approximation methods for nonlinear operations, e.g., [GELU](https://paperswithcode.com/method/gelu), [Softmax](https://paperswithcode.com/method/softmax), and [Layer Normalization](https://paperswithcode.com/method/layer-normalization), it performs an end-to-end integer-only [BERT](https://paperswithcode.com/method/bert) inference without any floating point calculation.\r\n\r\nIn particular, GELU and Softmax are approximated with lightweight second-order polynomials, which can be evaluated with integer-only arithmetic. For LayerNorm, integer-only computation is performed by leveraging a known algorithm for integer calculation of\r\nsquare root.","description_state":"present","introduced_year":null,"introduced_by":{"title":"I-BERT: Integer-only BERT Quantization","paper":"/paper/i-bert-integer-only-bert-quantization","first_author":"Sehoon Kim","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/i-bert-integer-only-bert-quantization"},"source":{"url":"https://arxiv.org/abs/2101.01321v3","title":"I-BERT: Integer-only BERT Quantization","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":3,"papers_newest_first":[{"paper":"/paper/mixed-non-linear-quantization-for-vision","title":"Mixed Non-linear Quantization for Vision Transformers","date":"2024-07-26","arxiv_id":"2407.18437","n_code_links":1,"syntology":null},{"paper":null,"title":"The Feasibility of Implementing Large-Scale Transformers on Multi-FPGA Platforms","date":"2024-04-24","arxiv_id":"2404.16158","n_code_links":0,"syntology":null},{"paper":"/paper/i-bert-integer-only-bert-quantization","title":"I-BERT: Integer-only BERT Quantization","date":"2021-01-05","arxiv_id":"2101.01321","n_code_links":7,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/quantization","name":"Quantization","papers":2},{"task":null,"name":"GPU","papers":1},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1}],"tasks_shown":4,"n_tasks":4,"usage_by_year":[{"year":"2021","papers":1},{"year":"2024","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/i-bert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}