SAVRN
Search Contact SAVRN

Open-weight model

osim-4b-post-multi-utterance

by Catherine Guo bluewater327/osim-4b-post-multi-utterance

Parameters4B
Context32,768
Weights8.0 GB
License
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve osim-4b-post-multi-utterance (4B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 8.0 GB 9.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.0 GB 4.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.0 GB 2.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

The publisher has not written a card for this model.

Configuration

Architecture
Qwen3ForCausalLM
Context length (tokens)
32,768
Layers
36
Hidden size
2,560
Feed-forward size
9,728
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
151,936
RoPE base
1,000,000
Model type
qwen3

Identity and Version

Repository
bluewater327/osim-4b-post-multi-utterance
Publisher
Catherine Guo
Task
Not stated by the source
Modality
Other
Library
Not stated by the source
Parameters
4B parameters
Languages
Not stated by the source
Revision
667dcb853e3a0cc80b552b3f226f08669fd6979e
First published
2026-09-18
Last updated
2026-09-18

Files and Weights

13 files, 8.1 GB in total. The weights are 2 files totalling 8.0 GB in safetensors.

Weights2 files · 8.0 GB
Configuration5 files · 35.8 KB
Tokenizer4 files · 15.9 MB
Other1 file · 5.3 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights5.0 GB 2f177b1f7987
model-00002-of-00002.safetensorsWeights3.1 GB 7b56a4938349
added_tokens.jsonConfiguration707 B
config.jsonConfiguration1.5 KB
generation_config.jsonConfiguration117 B
model.safetensors.index.jsonConfiguration32.9 KB
special_tokens_map.jsonConfiguration613 B
chat_template.jinjaOther5.3 KB
.gitattributesRepository1.6 KB
merges.txtTokenizer1.7 MB
tokenizer.jsonTokenizer11.4 MB aeb13307a71a
tokenizer_config.jsonTokenizer5.4 KB
vocab.jsonTokenizer2.8 MB

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
8.0 GB
Download from Catherine Guo

Released by Catherine Guo through its official repository on Hugging Face.

Memory Requirements

PrecisionWeights in memory
As published8.0 GB
16-bit8.0 GB
8-bit4.0 GB
4-bit2.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About osim-4b-post-multi-utterance

How much GPU memory does osim-4b-post-multi-utterance need?

About 9.7 GB at 16-bit and 2.4 GB at 4-bit: the weights (4B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run osim-4b-post-multi-utterance on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What is osim-4b-post-multi-utterance's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.