Methods › Natural Language Processing › Language Models › GPT-NeoX
GPT-NeoX
Introduced by Sid Black et al. in GPT-NeoX-20B: An Open-Source Autoregressive Language Model
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
GPT-NeoX is an autoregressive transformer decoder model whose architecture largely follows that of GPT-3, with a few notable deviations. The model has 20 billion parameters with 44 layers, a hidden dimension size of 6144, and 64 heads. The main difference with GPT-3 is the change in tokenizer, the addition of Rotary Positional Embeddings, the parallel computation of attention and feed-forward layers, and a different initialization scheme and hyperparameters.
Papers archive 2025-07-28
11 shown of 11, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Extending LLMs' Context Window with 100 Samples 13 Jan 2024 · 1 repository · arXiv:2401.07004Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)
-
Efficient LLM Inference on CPUs 1 Nov 2023 · 2 repositories · arXiv:2311.00502Syntology ran 2 of 8 samples · 6 unverified
-
CLEX: Continuous Length Extrapolation for Large Language Models 25 Oct 2023 · 1 repository · arXiv:2310.16450Syntology ran 9 of 10 samples · 1 unverified
-
How well can machine-generated texts be identified and can language models be trained to avoid identification? 25 Oct 2023 · 0 repositories · arXiv:2310.16992
-
H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models 21 Sep 2023 · 1 repository
-
Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data? 16 Sep 2023 · 1 repository · arXiv:2309.08963
-
H₂O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models 24 Jun 2023 · 2 repositories · arXiv:2306.14048Syntology ran 1 of 1 samples · 0 unverified
-
Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks 23 May 2023 · 1 repository · arXiv:2305.14201
-
DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature 26 Jan 2023 · 4 repositories · arXiv:2301.11305Syntology ran 3 of 9 samples · 6 unverified · 2 pointer-only (licence)
-
Mass-Editing Memory in a Transformer 13 Oct 2022 · 2 repositories · arXiv:2210.07229Syntology ran 6 of 8 samples · 2 unverified
-
GPT-NeoX-20B: An Open-Source Autoregressive Language Model 14 Apr 2022 · 11 repositories · arXiv:2204.06745Syntology ran 0 of 3 samples · 3 unverified
Tasks archive 2025-07-28
14 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Language Modelling | 3 |
| GPU | 2 |
| Language Modeling | 2 |
| Position | 2 |
| 4k | 1 |
| Articles | 1 |
| Attribute | 1 |
| Dataset Generation | 1 |
| Hallucination | 1 |
| Linguistic Acceptability | 1 |
| Multi-task Language Understanding | 1 |
| Quantization | 1 |
| Text Detection | 1 |
| Text Generation | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections