Model · Text generation
NVIDIA
Dec 2025 \- Jan 2026 September 2024 The pretraining data has a cutoff date of September 2024\. NVIDIA-Nemotron-3-Nano-4B-BF16 is a small language model (SLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be controlled via a system prompt. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in accuracy for harder prompts that require reasoning. Conversely, allowing the…
Open weights
other
4B parameters
262,144 tokens
transformers
An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 54 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3ncgroot16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at…
Open weights
4B parameters
262,144 tokens
transformers
An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 56 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3ncbase. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at the…
Open weights
4B parameters
262,144 tokens
transformers
An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 8 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3nbgroot16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at…
Open weights
4B parameters
262,144 tokens
transformers
An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 22 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3ncvs16. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at the…
Open weights
4B parameters
262,144 tokens
transformers
An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 12 of RL run seededrlbaseramp25stoppengen4kep2ncp5n3nciid16s2. - This is the best checkpoint by pass@8 so far in this run. Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). This checkpoint is the main revision (git branch) of the repo, with the model at…
Open weights
4B parameters
262,144 tokens
transformers