SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

h2o-lightning-4b

by H2O.ai h2oai/h2o-lightning-4b

h2o-lightning-4b is an open-weight model for image and text to text from H2O.ai, released under Apache License 2.0. It has 4.5B parameters and a 262,144-token context. At 16-bit it needs about 10.9 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 135 downloads a month.

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B.

Parameters4.5B
Context262,144
Weights9.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads135

Runs On

What it takes to serve h2o-lightning-4b (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.1 GB 10.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.5 GB 5.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.3 GB 2.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 8, 2026.

h2o-lightning-4b on every accelerator the SAVRN Index prices, at every precision

Model Card

By H2O.ai, published under apache-2.0, revision d5c279cf3fa7.

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option:

  • choice: picks one of a set of named options;
  • yes/no (noul): gives the probability that a statement is true;
  • score: gives an ordinal level on a stated scale, with its probability distribution.

It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2o_lightning_shim.py). Each decision is one forward pass and one output token, so the cost is input tokens only.

#1 on JevBench (October 7, 2026)

Read the full model card (2,619 words)

Configuration

Architecture
Qwen3_5ForConditionalGeneration
Context length (tokens)
262,144
Layers
32
Hidden size
2,560
Feed-forward size
9,216
Attention heads
16
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5

Identity and Version

Repository
h2oai/h2o-lightning-4b
Publisher
H2O.ai
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
4.5B parameters
Languages
Not stated by the source
Revision
d5c279cf3fa717b6094f5db4d241204d412a4182
First published
2026-10-02
Last updated
2026-10-08

Files and Weights

23 files, 9.6 GB in total. The weights are 2 files totalling 9.6 GB in safetensors.

Weights2 files · 9.6 GB
Configuration8 files · 80.3 KB
Tokenizer4 files · 22.9 MB
Documentation2 files · 31.9 KB
Other6 files · 386.0 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
adapter/adapter_model.safetensorsWeights519.5 MB 98a4b55083e9
model.safetensorsWeights9.1 GB 2d51561f221e
adapter/adapter_config.jsonConfiguration440 B —
config.jsonConfiguration2.9 KB —
generation_config.jsonConfiguration116 B —
h2o_lightning_shim.pyConfiguration47.4 KB —
preprocessor_config.jsonConfiguration390 B —
serve_config.jsonConfiguration2.2 KB —
test_shim.pyConfiguration26.6 KB —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.5 KB —
README.mdDocumentation20.4 KB —
assets/jevbench_capability_vs_cost_2026-10-07.pngOther97.0 KB —
assets/jevbench_composite_2026-10-07.pngOther190.3 KB 8333a5a4f510
assets/jevbench_four_axes_2026-10-07.pngOther57.9 KB —
chat_template.jinjaOther7.8 KB —
example_receipt.pngOther31.1 KB —
serve.shOther2.0 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
9.6 GB
Download from H2O.ai

Released by H2O.ai through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published9.6 GB
16-bit9.1 GB
8-bit4.5 GB
4-bit2.3 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About h2o-lightning-4b

How much GPU memory does h2o-lightning-4b need?

About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run h2o-lightning-4b on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use h2o-lightning-4b commercially?

Yes. h2o-lightning-4b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is h2o-lightning-4b's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Vinci-Piccolo-1.0

SimpleDirect

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Omni-Edu-4B

Hao Liang

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-baseline-label-160updates

XinLi

The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

DN-MOPD-Qwen3.5-4B-160updates

XinLi

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

ezjev-4b-s2

Everett

A typed-decision model (Jev-style /v1/systemone: choice, noul and score questions answered with a probability per option from one forward pass). LoRA fine-tune of Qwen/Qwen3.5-4B, merged into full weights, in two stages: 1. Stage 1 (ezjev-4b): LoRA r=16 on all language-model linear layers (attention, Gated-DeltaNet, MLP), LR 1e-4, one epoch over ~85k rows / ~106k questions. Loss: cross-entropy + Brier over the option labels, llm2jev chat prompt (thinking off). 2. Stage 2: a second LoRA at LR 5e-5 on ~30k rows targeting weak task families (RAG hallucination, product relevance, code-output selection, select-all-that-apply, clinical NLI, phishing emails, claim verification) with 40% replay.…

Open weights other 4.5B parameters 262,144 tokens transformers

Model · Image and text to text

Omni-Edu-4B-FP8

Hao Liang

Open Foundation Models for Learning and Teaching OmniEdu-4B-FP8 is the FP8-quantized release of OmniEdu-4B. The architecture, tokenizer, chat template and instruction-tuning data are unchanged: this checkpoint is the same model with its linear layers stored in FP8, which shrinks the weights from 8.46 GiB to 5.13 GiB and is intended for deployment on smaller GPUs and for higher-throughput serving. This checkpoint accompanies OmniEdu: Open Foundation Models for Learning and University · University of the Chinese Academy of Sciences · Zhongguancun Academy Quantized with LLM Compressor into the compressed-tensors FP8BLOCK scheme (library version 0.19.0), without calibration data: The vision…

Open weights other 4.5B parameters 262,144 tokens transformers