SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

ezjev-4b-s2

by Everett everettjf/ezjev-4b-s2

ezjev-4b-s2 is an open-weight model for image and text to text from Everett, released under other. It has 4.5B parameters and a 262,144-token context. At 16-bit it needs about 10.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

A typed-decision model (Jev-style /v1/systemone: choice, noul and score questions answered with a probability per option from one forward pass). LoRA fine-tune of Qwen/Qwen3.5-4B, merged into full weights, in two stages: 1.

Parameters4.5B
Context262,144
Weights9.1 GB
Licenseother
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve ezjev-4b-s2 (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.1 GB 10.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.5 GB 5.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

ezjev-4b-s2 on every accelerator the SAVRN Index prices, at every precision

Model Card

A typed-decision model (Jev-style /v1/systemone: choice, noul and score questions answered with a probability per option from one forward pass). LoRA fine-tune of Qwen/Qwen3.5-4B, merged into full weights, in two stages: 1. Stage 1 (ezjev-4b): LoRA r=16 on all language-model linear layers (attention, Gated-DeltaNet, MLP), LR 1e-4, one epoch over ~85k rows / ~106k questions. Loss: cross-entropy + Brier over the option labels, llm2jev chat prompt (thinking off). 2. Stage 2: a second LoRA at LR 5e-5 on ~30k rows targeting weak task families (RAG hallucination, product relevance, code-output selection, select-all-that-apply, clinical NLI, phishing emails, claim verification) with 40% replay.…

Excerpt from the card by Everett, licensed other.

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
2,560
Feed-forward size
9,216
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
everettjf/ezjev-4b-s2
Publisher
Everett
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
4.5B parameters
Languages
jev
Revision
3d77eb1565de5b04e77fa3b599b273656fd22a62
First published
2026-10-03
Last updated
2026-10-04

Files and Weights

10 files, 9.1 GB in total. The weights are 1 file totalling 9.1 GB in safetensors.

Weights1 file · 9.1 GB
Configuration4 files · 6.8 KB
Tokenizer2 files · 20.0 MB
Documentation1 file · 3.1 KB
Other1 file · 7.8 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model.safetensorsWeights9.1 GB e06fdb9bae4b
config.jsonConfiguration2.9 KB —
ezjev.jsonConfiguration2.5 KB —
generation_config.jsonConfiguration116 B —
processor_config.jsonConfiguration1.2 KB —
README.mdDocumentation3.1 KB —
chat_template.jinjaOther7.8 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer20.0 MB a5cd9732badc
tokenizer_config.jsonTokenizer1.5 KB —

License and Download

License
other
Access
Open weights, no gate
Download size
9.1 GB
Download from Everett

Released by Everett through its official repository on Hugging Face.

Built From

Memory Requirements

PrecisionWeights in memory
As published9.1 GB
16-bit9.1 GB
8-bit4.5 GB
4-bit2.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About ezjev-4b-s2

How much GPU memory does ezjev-4b-s2 need?

About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run ezjev-4b-s2 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What license is ezjev-4b-s2 released under?

other, as its publisher declares it. Read the license text before commercial use.

What is ezjev-4b-s2's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

h2o-lightning-4b

H2O.ai

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option: - yes/no (noul): gives the probability that a statement is true; It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2olightningshim.py). Each decision is one forward pass and one output token, so the cost is input tokens only. On JevBench v1.6.1, H2O-Lightning-4B v1.1 is #1 on the official Composite Score of open-weight systems…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Vinci-Piccolo-1.0

SimpleDirect

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Omni-Edu-4B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-160updates

XinLi

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Omni-Edu-4B-FP8

Hao Liang

Open Foundation Models for Learning and Teaching OmniEdu-4B-FP8 is the FP8-quantized release of OmniEdu-4B. The architecture, tokenizer, chat template and instruction-tuning data are unchanged: this checkpoint is the same model with its linear layers stored in FP8, which shrinks the weights from 8.46 GiB to 5.13 GiB and is intended for deployment on smaller GPUs and for higher-throughput serving. This checkpoint accompanies OmniEdu: Open Foundation Models for Learning and University · University of the Chinese Academy of Sciences · Zhongguancun Academy Quantized with LLM Compressor into the compressed-tensors FP8BLOCK scheme (library version 0.19.0), without calibration data: The vision…

Open weights other 4.5B parameters 262,144 tokens transformers