SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Omni-Edu-4B-FP8

by Hao Liang lhpku20010120/Omni-Edu-4B-FP8

Omni-Edu-4B-FP8 is an open-weight model for image and text to text from Hao Liang, released under other. It has 4.5B parameters and a 262,144-token context. At 16-bit it needs about 10.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Open Foundation Models for Learning and Teaching OmniEdu-4B-FP8 is the FP8-quantized release of OmniEdu-4B.

Parameters4.5B
Context262,144
Weights5.5 GB
Licenseother
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Omni-Edu-4B-FP8 (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.1 GB 10.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.5 GB 5.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

Omni-Edu-4B-FP8 on every accelerator the SAVRN Index prices, at every precision

Model Card

Open Foundation Models for Learning and Teaching OmniEdu-4B-FP8 is the FP8-quantized release of OmniEdu-4B. The architecture, tokenizer, chat template and instruction-tuning data are unchanged: this checkpoint is the same model with its linear layers stored in FP8, which shrinks the weights from 8.46 GiB to 5.13 GiB and is intended for deployment on smaller GPUs and for higher-throughput serving. This checkpoint accompanies OmniEdu: Open Foundation Models for Learning and University · University of the Chinese Academy of Sciences · Zhongguancun Academy Quantized with LLM Compressor into the compressed-tensors FP8BLOCK scheme (library version 0.19.0), without calibration data: The vision…

Excerpt from the card by Hao Liang, licensed other.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
2,560
Feed-forward size
9,216
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5
Quantization
compressed-tensors

Identity and Version

Repository
lhpku20010120/Omni-Edu-4B-FP8
Publisher
Hao Liang
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
4.5B parameters
Languages
Not stated by the source
Revision
3e21accd0f93afee9371ad108de5dcbf22b763a7
First published
2026-10-05
Last updated
2026-10-06

Files and Weights

10 files, 5.5 GB in total. The weights are 1 file totalling 5.5 GB in safetensors.

Weights1 file · 5.5 GB
Configuration4 files · 19.7 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 11.5 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights5.5 GB 80361f663b83
config.jsonConfiguration11.5 KB —
generation_config.jsonConfiguration164 B —
processor_config.jsonConfiguration1.3 KB —
recipe.yamlConfiguration6.7 KB —
README.mdDocumentation11.5 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB 87a7830d63fc
tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
other
Access
Open weights, no gate
Download size
5.5 GB
Download from Hao Liang

Released by Hao Liang through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published5.5 GB
16-bit9.1 GB
8-bit4.5 GB
4-bit2.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Omni-Edu-4B-FP8

How much GPU memory does Omni-Edu-4B-FP8 need?

About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Omni-Edu-4B-FP8 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is Omni-Edu-4B-FP8 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is Omni-Edu-4B-FP8's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

h2o-lightning-4b

H2O.ai

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option: - yes/no (noul): gives the probability that a statement is true; It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2olightningshim.py). Each decision is one forward pass and one output token, so the cost is input tokens only. On JevBench v1.6.1, H2O-Lightning-4B v1.1 is #1 on the official Composite Score of open-weight systems…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Vinci-Piccolo-1.0

SimpleDirect

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Omni-Edu-4B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-160updates

XinLi

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

ezjev-4b-s2

Everett

A typed-decision model (Jev-style /v1/systemone: choice, noul and score questions answered with a probability per option from one forward pass). LoRA fine-tune of Qwen/Qwen3.5-4B, merged into full weights, in two stages: 1. Stage 1 (ezjev-4b): LoRA r=16 on all language-model linear layers (attention, Gated-DeltaNet, MLP), LR 1e-4, one epoch over ~85k rows / ~106k questions. Loss: cross-entropy + Brier over the option labels, llm2jev chat prompt (thinking off). 2. Stage 2: a second LoRA at LR 5e-5 on ~30k rows targeting weak task families (RAG hallucination, product relevance, code-output selection, select-all-that-apply, clinical NLI, phishing emails, claim verification) with 40% replay.…

Open weights other 4.5B parameters 262,144 tokens transformers