NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments…
Runs On
What it takes to serve Cosmos3-Edge (3.9B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 7.7 GB | 9.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 3.9 GB | 4.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 1.9 GB | 2.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.
SAVRN's Notes on Cosmos3-Edge
Nine point three gigabytes of memory is the whole ask for Cosmos3-Edge at 16-bit, and that number settles the hardware question. NVIDIA built this 3.9 billion parameter world model for physical AI: text, image, video or an action trajectory goes in, and video, image, audio or action commands come out for robots, vehicles and factory floors. At 8-bit the working set drops to 4.6 GB and at 4-bit to 2.3 GB, so the cheapest card on the Index, one MI300X with 192 GB at $1.85 an hour on-demand, carries it with most of the card idle. It belongs beside larger jobs.
The license reads other, with no summary in our record, so pull the actual terms before any commercial deployment. Check the 131,072-token context against your real input lengths, and watch the dates: released July 1, 2026, updated September 16, 2026, so the files are still moving.
Model Card
NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications. Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy…
Excerpt from the card by NVIDIA, licensed other.
Configuration
- Architecture
- Cosmos3EdgeForConditionalGeneration
- Context length (tokens)
- 131,072
- Layers
- 28
- Hidden size
- 2,048
- Feed-forward size
- 9,216
- Attention heads
- 16
- Key/value heads
- 8
- Head dimension
- 128
- Vocabulary size
- 131,072
- Model type
- cosmos3_edge
Identity and Version
- Repository
- nvidia/Cosmos3-Edge
- Publisher
- NVIDIA
- Task
- Not stated by the source
- Modality
- Other
- Library
- cosmos
- Parameters
- 3.9B parameters
- Languages
- Not stated by the source
- Revision
- 344d602b128d1bbdacb43b08d0a3626f46343e29
- First published
- 2026-07-01
- Last updated
- 2026-09-16
Files and Weights
54 files, 9.2 GB in total. The weights are 4 files totalling 9.1 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| transformer/diffusion_pytorch_model-00001-of-00002.safetensors | Weights | 5.0 GB | 116134e5492c |
| transformer/diffusion_pytorch_model-00002-of-00002.safetensors | Weights | 1.7 GB | 3b4c0aafe270 |
| vae/diffusion_pytorch_model.safetensors | Weights | 1.4 GB | 230496cb59ff |
| vision_encoder/model.safetensors | Weights | 978.7 MB | 2180ad739ecc |
| assets/diffusers_outputs/edge_action_id_av_inverse_0_diffusers.json | Configuration | 15.9 KB | — |
| assets/diffusers_outputs/edge_action_id_av_inverse_1_diffusers.json | Configuration | 15.7 KB | — |
| assets/edge_action_id_av_0_output.json | Configuration | 11.2 KB | — |
| assets/edge_action_id_av_1_output.json | Configuration | 11.0 KB | — |
| assets/example_action_fd_umi_action_chunks.json | Configuration | 10.3 KB | — |
| assets/example_i2v_prompt.json | Configuration | 9.2 KB | — |
| assets/example_reasoning_prompt.json | Configuration | 151 B | — |
| assets/negative_prompt.json | Configuration | 17.6 KB | — |
| config.json | Configuration | 1.7 KB | — |
| generation_config.json | Configuration | 154 B | — |
| model.safetensors.index.json | Configuration | 68.6 KB | — |
| model_index.json | Configuration | 555 B | — |
| modular_model_index.json | Configuration | 1.4 KB | — |
| preprocessor_config.json | Configuration | 334 B | — |
| processor_config.json | Configuration | 875 B | — |
| scheduler/scheduler_config.json | Configuration | 889 B | — |
| special_tokens_map.json | Configuration | 563 B | — |
| text_tokenizer/special_tokens_map.json | Configuration | 563 B | — |
| transformer/config.json | Configuration | 1.1 KB | — |
| transformer/diffusion_pytorch_model.safetensors.index.json | Configuration | 53.6 KB | — |
| vae/config.json | Configuration | 1.8 KB | — |
| video_preprocessor_config.json | Configuration | 367 B | — |
| BIAS.md | Documentation | 4.7 KB | — |
| EXPLAINABILITY.md | Documentation | 3.2 KB | — |
| PRIVACY.md | Documentation | 1.2 KB | — |
| README.md | Documentation | 53.0 KB | — |
| SAFETY.md | Documentation | 3.7 KB | — |
| assets/benchmark-image2video.png | Other | 58.5 KB | — |
| assets/benchmark-overall.png | Other | 81.0 KB | fdc1dbbeda37 |
| assets/diffusers_outputs/edge_action_fd_diffusers_chunk_00.mp4 | Other | 14.8 KB | — |
| assets/diffusers_outputs/edge_action_fd_diffusers_chunk_01.mp4 | Other | 13.2 KB | — |
| assets/diffusers_outputs/edge_action_fd_umi_2chunk_diffusers.mp4 | Other | 23.6 KB | — |
| assets/diffusers_outputs/edge_i2v_diffusers.mp4 | Other | 1.3 MB | 2fe93251039d |
| assets/edge_action_fd_umi_2chunk_output.mp4 | Other | 32.8 KB | 10037bf0c2a2 |
| assets/edge_action_id_av_0_output.png | Other | 113.8 KB | 18b90a143ee0 |
| assets/edge_action_id_av_1_output.png | Other | 102.5 KB | 660085cebabd |
| assets/edge_i2v_output.mp4 | Other | 8.1 MB | d4e87cbc2efe |
| assets/example_action_fd_umi_first_frame.png | Other | 67.0 KB | — |
| assets/example_action_id_av_0_input.mp4 | Other | 1.1 MB | ff205f86ae16 |
| assets/example_action_id_av_1_input.mp4 | Other | 1.6 MB | 169e65cee76e |
| assets/example_i2v_input.jpg | Other | 860.8 KB | 1de51eb5c6d5 |
| assets/example_reasoning_input.png | Other | 230.7 KB | 6686b937bdb2 |
| chat_template.jinja | Other | 12.2 KB | — |
| images/benchmark-reasoning.png | Other | 265.8 KB | 75974a959d0e |
| text_tokenizer/chat_template.jinja | Other | 12.2 KB | — |
| .gitattributes | Repository | 2.9 KB | — |
| text_tokenizer/tokenizer.json | Tokenizer | 17.1 MB | 4dc692a99dca |
| text_tokenizer/tokenizer_config.json | Tokenizer | 393 B | — |
| tokenizer.json | Tokenizer | 17.1 MB | 4dc692a99dca |
| tokenizer_config.json | Tokenizer | 177.3 KB | — |
License and Download
- License
- other
- Access
- Open weights, no gate
- Download size
- 9.1 GB
Released by NVIDIA through its official repository on Hugging Face.
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 9.1 GB |
| 16-bit | 7.7 GB |
| 8-bit | 3.9 GB |
| 4-bit | 1.9 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About Cosmos3-Edge
How much GPU memory does Cosmos3-Edge need?
About 9.3 GB at 16-bit and 2.3 GB at 4-bit: the weights (3.9B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run Cosmos3-Edge on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
What license is Cosmos3-Edge released under?
other, as its publisher declares it. Read the license text before commercial use.
What is Cosmos3-Edge's context length?
131,072 tokens, from the maximum position embeddings in its published configuration.