# cobalt-seeded-rl-base-ram…ncp5-n3nc-base by Alexander Gurung
Source: https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base (4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 7.9 GB | 9.5 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 4.0 GB | 4.8 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 2.0 GB | 2.4 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 9, 2026.

[cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base/gpus)

## Model Card

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 56 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3ncbase. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at the…

Excerpt from the card by Alexander Gurung.

## Configuration

Architecture

NemotronHForCausalLM

Context length (tokens)

262,144

Hidden size

3,136

Feed-forward size

12,544

Attention heads

40

Key/value heads

8

Head dimension

128

Vocabulary size

131,072

Routed experts

8

Experts active per token

2

Model type

nemotron_h

## Identity and Version

Repository

agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base

Publisher

Alexander Gurung

Task

Text generation

Modality

Text

Library

transformers

Parameters

4B parameters

Languages

Not stated by the source

Revision

7faf216b1363a8afad21504c03a39ea6886eb085

First published

2026-10-07

Last updated

2026-10-09

## Files and Weights

7 files, 8.0 GB in total. The weights are 1 file totalling 7.9 GB in safetensors.

Weights1 file · 7.9 GB

Configuration2 files · 2.7 KB

Tokenizer2 files · 17.1 MB

Documentation1 file · 2.6 KB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 7.9 GB | 122cd854fccb |
| config.json | Configuration | 2.5 KB | — |
| generation_config.json | Configuration | 192 B | — |
| README.md | Documentation | 2.6 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| tokenizer.json | Tokenizer | 17.1 MB | 623c34567aeb |
| tokenizer_config.json | Tokenizer | 393 B | — |

## License and Download

License

Not stated by the source

Access

Open weights, no gate

Download size

7.9 GB

[Download from Alexander Gurung](https://huggingface.co/agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base)

Released by Alexander Gurung through its official repository on Hugging Face.

## Built From

- Derived from [Qwen/Qwen3-4B-Instruct-2507](https://savrn.com/models/qwen3-4b-instruct-2507)

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 7.9 GB |
| 16-bit | 7.9 GB |
| 8-bit | 4.0 GB |
| 4-bit | 2.0 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base

### How much GPU memory does cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base need?

About 9.5 GB at 16-bit and 2.4 GB at 4-bit: the weights (4B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### What is cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [NVIDIA-Nemotron-3-Nano-4B-BF16](https://savrn.com/models/nvidia-nemotron-3-nano-4b-bf16)

[NVIDIA](https://savrn.com/model-publishers/nvidia)

Dec 2025 \- Jan 2026 September 2024 The pretraining data has a cutoff date of September 2024\. NVIDIA-Nemotron-3-Nano-4B-BF16 is a small language model (SLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the…

Open weights other 4B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/nvidia-nemotron-3-nano-4b-bf16)

Model · Text generation

### [cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-groot16](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-groot16)

[Alexander Gurung](https://savrn.com/model-publishers/agurung)

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 54 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3ncgroot16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at…

Open weights 4B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-groot16)

Model · Text generation

### [cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nb-groot16](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nb-groot16)

[Alexander Gurung](https://savrn.com/model-publishers/agurung)

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 8 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3nbgroot16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at…

Open weights 4B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nb-groot16)

Model · Text generation

### [cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-iid16](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-iid16)

[Alexander Gurung](https://savrn.com/model-publishers/agurung)

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 24 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3nciid16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at the…

Open weights 4B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-iid16)

Model · Text generation

### [cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-vs16](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-vs16)

[Alexander Gurung](https://savrn.com/model-publishers/agurung)

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 22 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3ncvs16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at the…

Open weights 4B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-vs16)

Model · Text generation

### [cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-iid16-s2](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-iid16-s2)

[Alexander Gurung](https://savrn.com/model-publishers/agurung)

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 12 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3nciid16s2. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at…

Open weights 4B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-iid16-s2)

## Alexander Gurung

[All models and datasets](https://savrn.com/model-publishers/agurung)

## Versions

- [7faf216b1363](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base/versions/7faf216b1363) · current 2026-10-09
- [82a52039f85d](https://savrn.com/models/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base/versions/82a52039f85d) 2026-10-08

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-09.
- [Hugging Face record](https://huggingface.co/agurung/cobalt-seeded-rl-base-ramp25-stoppen-gen4k-ep2-ncp5-n3nc-base)
- [How the hub is built](https://savrn.com/model-hub/methodology)
