SAVRN
Search Contact SAVRN

Open-weight model · Text generation

gpt-neox-20b

by EleutherAI EleutherAI/gpt-neox-20b

GPT-NeoX-20B is a 20 billion parameter autoregressive language model trained on the Pile using the GPT-NeoX library. Its architecture intentionally resembles that of GPT-3, and is almost identical to that of GPT-J- 6B.

Parameters20.7B
Context2,048
Weights82.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads707.1k

Runs On

What it takes to serve gpt-neox-20b (20.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 41.5 GB 49.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 20.7 GB 24.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 10.4 GB 12.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on gpt-neox-20b

20.7 billion parameters, a 2,048-token window, and a training set with its own datasheet. EleutherAI released it on April 7, 2022 as a general-purpose English text generator whose design tracks GPT-3 and sits close to GPT-J-6B. Memory is comfortable: 49.8 GB at 16-bit, 24.9 GB at 8-bit, 12.4 GB at 4-bit, all inside the single MI300X at $1.85 per hour in the cheapest listed setup. The 2,048-token context sets the job: short prompts, completion and fine-tuning experiments rather than long documents.

Apache 2.0 gives you commercial use, modification and redistribution with notice obligations and a patent grant, so a fine-tuned derivative can be redistributed if the notices travel with it and significant changes are stated. Read what you inherit: the Pile is described in arXiv:2101.00027 and its datasheet in arXiv:2201.07311, and the model paper is arXiv:2204.06745, which answer the provenance questions a review board will raise.

Model Card

By EleutherAI, published under apache-2.0, revision c292233c833e.

GPT-NeoX-20B is a 20 billion parameter autoregressive language model trained on the Pile using the GPT-NeoX library. Its architecture intentionally resembles that of GPT-3, and is almost identical to that of GPT-J- 6B. Its training dataset contains a multitude of English-language texts, reflecting the general-purpose nature of this model. See the accompanying paper for details about model architecture (including how it differs from GPT-3), training procedure, and additional evaluations.

Model details

parameters layers model heads head vocab -5

Uses and limitations

Intended use

Read the full model card (877 words)

Configuration

Architecture
GPTNeoXForCausalLM
Context length (tokens)
2,048
Layers
44
Hidden size
6,144
Feed-forward size
24,576
Attention heads
64
Vocabulary size
50,432
Stored precision
float16
Model type
gpt_neox

Identity and Version

Repository
EleutherAI/gpt-neox-20b
Publisher
EleutherAI
Task
Text generation
Modality
Text
Library
transformers
Parameters
20.7B parameters
Languages
en
Revision
c292233c833e336628618a88a648727eb3dff0a7
First published
2022-04-07
Last updated
2024-01-31

Files and Weights

102 files, 82.6 GB in total. The weights are 92 files totalling 82.6 GB in bin, safetensors.

Weights92 files · 82.6 GB
Configuration4 files · 118.8 KB
Tokenizer4 files · 3.6 MB
Documentation1 file · 9.0 KB
Repository1 file · 4.4 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00046.safetensorsWeights926.0 MB 24a2fc18ab96
model-00002-of-00046.safetensorsWeights910.3 MB eb326a6fc184
model-00003-of-00046.safetensorsWeights910.3 MB 95effe912096
model-00004-of-00046.safetensorsWeights910.3 MB a0ac9466a9c1
model-00005-of-00046.safetensorsWeights910.3 MB 9d6613a91548
model-00006-of-00046.safetensorsWeights910.3 MB cddbc8a13a77
model-00007-of-00046.safetensorsWeights910.3 MB 097ac8776465
model-00008-of-00046.safetensorsWeights910.3 MB 47ec7084cea6
model-00009-of-00046.safetensorsWeights910.3 MB 01fdfabdb95d
model-00010-of-00046.safetensorsWeights910.3 MB bcd3060ee2fb
model-00011-of-00046.safetensorsWeights910.3 MB fd1a57dfabaf
model-00012-of-00046.safetensorsWeights910.3 MB c85e72751dff
model-00013-of-00046.safetensorsWeights910.3 MB 179423abc852
model-00014-of-00046.safetensorsWeights910.3 MB 8dcc73bab432
model-00015-of-00046.safetensorsWeights910.3 MB 1df6529dfc08
model-00016-of-00046.safetensorsWeights910.3 MB 2b1a3727b480
model-00017-of-00046.safetensorsWeights910.3 MB 5860f6ef89f5
model-00018-of-00046.safetensorsWeights910.3 MB 62d8b2094835
model-00019-of-00046.safetensorsWeights910.3 MB 06cd4a326cd5
model-00020-of-00046.safetensorsWeights910.3 MB 8d192a4a220b
model-00021-of-00046.safetensorsWeights910.3 MB 24c8388deca5
model-00022-of-00046.safetensorsWeights910.3 MB dd821f6f05df
model-00023-of-00046.safetensorsWeights910.3 MB 21ac34fc590e
model-00024-of-00046.safetensorsWeights910.3 MB 506b22eebc59
model-00025-of-00046.safetensorsWeights910.3 MB eed055bbd7cd
model-00026-of-00046.safetensorsWeights910.3 MB b25b6edd4b9d
model-00027-of-00046.safetensorsWeights910.3 MB cf9c0cf13f55
model-00028-of-00046.safetensorsWeights910.3 MB 793c91ab150a
model-00029-of-00046.safetensorsWeights910.3 MB f4b1bcc99c22
model-00030-of-00046.safetensorsWeights910.3 MB bf342ac358e2
model-00031-of-00046.safetensorsWeights910.3 MB 713341d88368
model-00032-of-00046.safetensorsWeights910.3 MB d26e8d4c5811
model-00033-of-00046.safetensorsWeights910.3 MB 757880b98e70
model-00034-of-00046.safetensorsWeights910.3 MB 808c213e7e6e
model-00035-of-00046.safetensorsWeights910.3 MB 11113c4a99b3
model-00036-of-00046.safetensorsWeights910.3 MB cd5649b0348a
model-00037-of-00046.safetensorsWeights910.3 MB aaf2e9c33501
model-00038-of-00046.safetensorsWeights910.3 MB a815e95afd13
model-00039-of-00046.safetensorsWeights910.3 MB cbcc362c5028
model-00040-of-00046.safetensorsWeights910.3 MB 1a24e2b3a2be
model-00041-of-00046.safetensorsWeights910.3 MB d8e26c191074
model-00042-of-00046.safetensorsWeights910.3 MB 7172ff989ec0
model-00043-of-00046.safetensorsWeights910.3 MB 0266b06e509f
model-00044-of-00046.safetensorsWeights910.3 MB 0ae11994b839
model-00045-of-00046.safetensorsWeights604.1 MB 70ae0ac80dd9
model-00046-of-00046.safetensorsWeights619.7 MB 01499aa37fda
pytorch_model-00001-of-00046.binWeights926.0 MB 91a6926dd4e2
pytorch_model-00002-of-00046.binWeights910.3 MB a96a36c8efc6
pytorch_model-00003-of-00046.binWeights910.3 MB 582e2bbbbd7c
pytorch_model-00004-of-00046.binWeights910.3 MB dad2df6d880f
pytorch_model-00005-of-00046.binWeights910.3 MB 2dd1111e7f20
pytorch_model-00006-of-00046.binWeights910.3 MB c7de215cb446
pytorch_model-00007-of-00046.binWeights910.3 MB 53b30b05245b
pytorch_model-00008-of-00046.binWeights910.3 MB f4b0850c82bc
pytorch_model-00009-of-00046.binWeights910.3 MB 19b9c1d811ef
pytorch_model-00010-of-00046.binWeights910.3 MB 343cfe1a395d
pytorch_model-00011-of-00046.binWeights910.3 MB 089530ad755e
pytorch_model-00012-of-00046.binWeights910.3 MB 7040b3aadd06
pytorch_model-00013-of-00046.binWeights910.3 MB 50d7a7891881
pytorch_model-00014-of-00046.binWeights910.3 MB 4ba306de96a2
pytorch_model-00015-of-00046.binWeights910.3 MB 5105b09e42b2
pytorch_model-00016-of-00046.binWeights910.3 MB 41c5345d12c4
pytorch_model-00017-of-00046.binWeights910.3 MB 73a3f74f1e0c
pytorch_model-00018-of-00046.binWeights910.3 MB c17a90b84c95
pytorch_model-00019-of-00046.binWeights910.3 MB ce0eab7c8271
pytorch_model-00020-of-00046.binWeights910.3 MB 4a3e63161295
pytorch_model-00021-of-00046.binWeights910.3 MB 260f41230c97
pytorch_model-00022-of-00046.binWeights910.3 MB 43392e9228f8
pytorch_model-00023-of-00046.binWeights910.3 MB 11fb9cefd826
pytorch_model-00024-of-00046.binWeights910.3 MB 0ff5af200c26
pytorch_model-00025-of-00046.binWeights910.3 MB e38d80739d11
pytorch_model-00026-of-00046.binWeights910.3 MB ead5458cf979
pytorch_model-00027-of-00046.binWeights910.3 MB 34cfb1757f7e
pytorch_model-00028-of-00046.binWeights910.3 MB 4090aa141fb5
pytorch_model-00029-of-00046.binWeights910.3 MB 55b81d9a2491
pytorch_model-00030-of-00046.binWeights910.3 MB eaab6b1b9b16
pytorch_model-00031-of-00046.binWeights910.3 MB 9ad862519bb2
pytorch_model-00032-of-00046.binWeights910.3 MB 23752ed93b42
pytorch_model-00033-of-00046.binWeights910.3 MB 9d0842b355f9
pytorch_model-00034-of-00046.binWeights910.3 MB 7275a1541dd2
pytorch_model-00035-of-00046.binWeights910.3 MB d7b1e3ead65e
pytorch_model-00036-of-00046.binWeights910.3 MB 92801ef5b1dd
pytorch_model-00037-of-00046.binWeights910.3 MB 080b0ab69fdc
pytorch_model-00038-of-00046.binWeights910.3 MB 603b045908ca
pytorch_model-00039-of-00046.binWeights910.3 MB 9823f13162ca
pytorch_model-00040-of-00046.binWeights910.3 MB a231cae597ce
pytorch_model-00041-of-00046.binWeights910.3 MB bc909def52bb
pytorch_model-00042-of-00046.binWeights910.3 MB 2ac94d1b42f8
pytorch_model-00043-of-00046.binWeights910.3 MB 36259a4252ea
pytorch_model-00044-of-00046.binWeights910.3 MB c2c40045b534
pytorch_model-00045-of-00046.binWeights604.1 MB a7baf3947166
pytorch_model-00046-of-00046.binWeights619.7 MB 925dc7a0c53a
config.jsonConfiguration613 B
model.safetensors.index.jsonConfiguration60.4 KB
pytorch_model.bin.index.jsonConfiguration57.7 KB 044586635a1c
special_tokens_map.jsonConfiguration90 B
README.mdDocumentation9.0 KB
.gitattributesRepository4.4 KB
merges.txtTokenizer456.6 KB
tokenizer.jsonTokenizer2.1 MB
tokenizer_config.jsonTokenizer156 B
vocab.jsonTokenizer1.1 MB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
82.6 GB
Download from EleutherAI

Released by EleutherAI through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2101.00027
  • Described by arXiv:2104.09864
  • Described by arXiv:2201.07311
  • Described by arXiv:2204.06745
  • Trained on (disclosed) EleutherAI/pile

Memory Requirements

PrecisionWeights in memory
As published82.6 GB
16-bit41.5 GB
8-bit20.7 GB
4-bit10.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare gpt-neox-20b

Questions About gpt-neox-20b

How much GPU memory does gpt-neox-20b need?

About 49.8 GB at 16-bit and 12.4 GB at 4-bit: the weights (20.7B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run gpt-neox-20b on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use gpt-neox-20b commercially?

Yes. gpt-neox-20b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is gpt-neox-20b's context length?

2,048 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

Gemma-4-31B-IT-NVFP4

NVIDIA

Gemma 4 31B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding, and multimodal understanding on consumer GPUs and workstations, with a 256K-token context window and support for over 140 languages. The model uses a hybrid attention mechanism that interleaves local sliding-window and full global attention, with unified Keys and Values in global layers and Proportional RoPE (p-RoPE) to support long-context performance. The NVIDIA Gemma 4 31B IT NVFP4 model is quantized with NVIDIA Model…

Open weights other 20.9B parameters 262,144 tokens Model Optimizer

Model · Text generation

gpt-oss-20b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 20.9B parameters 131,072 tokens transformers

Model · Text generation

Ornith-1.5-35B-A3B-NVFP4

Ornith

Chirp Chirp! We are introducing Ornith-1.5, a major step toward building foundation models through end-to-end self-improvement. Ornith-1.5 extends Ornith-1.0 (which was developed on top of Qwen3.5 and Gemma4 with additional continued pretraining, mid-training, and post-training) by expanding the self-improvement loop from scaffold and rollout optimization to jointly optimizing task generation, scaffold construction, and solution rollouts. Rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning. For…

Open weights mit 19.5B parameters 262,144 tokens transformers

A refusal-removed (abliterated) build of Qwen's Qwen3.8-Flash-Next, quantized to EXL3 2.50 bpw so the full model runs on a single 24 GB card (RTX 3090 / 4090) using MoE CPU-offload — with the vision tower, MTP head, native 262,144-token context, and the PLE n-gram table all intact. Requires the same MoE CPU-offload setup as the stock 2.50bpw pack. Needs ~59 GB host RAM for the CPU expert tail and a fast NVMe for the streamed n-gram table. Expected on an RTX 3090: ~38 tok/s decode with MTP on (~28 without), ~20 tok/s at 175K depth, ~664 tok/s prefill. See the upstream repo for the full measured ledger; this quant uses the identical flags and layout, so numbers should track closely. Fired on…

Open weights other 22.3B parameters 262,144 tokens

Model · Text generation

Qwen3.6-35B-A3B-NVFP4

NVIDIA

The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is the quantized version of Alibaba's Qwen3.6-35B-A3B model, which is an auto-regressive language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Qwen3.6-35B-A3B-NVFP4 model is quantized with Model Optimizer. This model is ready for commercial/non-commercial use. This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case; see link to Non-NVIDIA (Qwen3.6-35B-A3B) Model Card from Alibaba. Global Developers looking to take off-the-shelf, pre-quantized models for deployment in AI Agent systems, chatbots…

Open weights apache-2.0 18.7B parameters 262,144 tokens Model Optimizer

Model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4

NVIDIA

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-Nano-3-30B-A3B-NVFP4 is a quantized version of Nemotron-Nano-3-30B-A3B and is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be…

Open weights other 18.2B parameters 262,144 tokens transformers