# h2o-lightning-4b by H2O.ai: Open-Weight Model
Source: https://savrn.com/models/h2o-lightning-4b
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve h2o-lightning-4b (4.5B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 9.1 GB | 10.9 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 4.5 GB | 5.4 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 2.3 GB | 2.7 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 8, 2026.

[h2o-lightning-4b on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/h2o-lightning-4b/gpus)

## Model Card

By H2O.ai, published under apache-2.0, revision d5c279cf3fa7.

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on [Qwen/Qwen3.5-4B](https://savrn.com/models/qwen3-5-4b). It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option:

- choice: picks one of a set of named options;
- yes/no (noul): gives the probability that a statement is true;
- score: gives an ordinal level on a stated scale, with its probability distribution.

It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2o_lightning_shim.py). Each decision is one forward pass and one output token, so the cost is input tokens only.

### #1 on JevBench (October 7, 2026)

[Read the full model card (2,619 words)](https://savrn.com/models/h2o-lightning-4b/card)

## Configuration

Architecture

Qwen3_5ForConditionalGeneration

Context length (tokens)

262,144

Layers

32

Hidden size

2,560

Feed-forward size

9,216

Attention heads

16

Key/value heads

4

Head dimension

256

Vocabulary size

248,320

Model type

qwen3_5

## Identity and Version

Repository

h2oai/h2o-lightning-4b

Publisher

H2O.ai

Task

Image and text to text

Modality

Image and text

Library

transformers

Parameters

4.5B parameters

Languages

Not stated by the source

Revision

d5c279cf3fa717b6094f5db4d241204d412a4182

First published

2026-10-02

Last updated

2026-10-08

## Files and Weights

23 files, 9.6 GB in total. The weights are 2 files totalling 9.6 GB in safetensors.

Weights2 files · 9.6 GB

Configuration8 files · 80.3 KB

Tokenizer4 files · 22.9 MB

Documentation2 files · 31.9 KB

Other6 files · 386.0 KB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| adapter/adapter_model.safetensors | Weights | 519.5 MB | 98a4b55083e9 |
| model.safetensors | Weights | 9.1 GB | 2d51561f221e |
| adapter/adapter_config.json | Configuration | 440 B | — |
| config.json | Configuration | 2.9 KB | — |
| generation_config.json | Configuration | 116 B | — |
| h2o_lightning_shim.py | Configuration | 47.4 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| serve_config.json | Configuration | 2.2 KB | — |
| test_shim.py | Configuration | 26.6 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| LICENSE | Documentation | 11.5 KB | — |
| README.md | Documentation | 20.4 KB | — |
| assets/jevbench_capability_vs_cost_2026-10-07.png | Other | 97.0 KB | — |
| assets/jevbench_composite_2026-10-07.png | Other | 190.3 KB | 8333a5a4f510 |
| assets/jevbench_four_axes_2026-10-07.png | Other | 57.9 KB | — |
| chat_template.jinja | Other | 7.8 KB | — |
| example_receipt.png | Other | 31.1 KB | — |
| serve.sh | Other | 2.0 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 5f9e4d4901a9 |
| tokenizer_config.json | Tokenizer | 16.7 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

9.6 GB

[Download from H2O.ai](https://huggingface.co/h2oai/h2o-lightning-4b)

Released by H2O.ai through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Built From

- Derived from [Qwen/Qwen3.5-4B](https://savrn.com/models/qwen3-5-4b)

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 9.6 GB |
| 16-bit | 9.1 GB |
| 8-bit | 4.5 GB |
| 4-bit | 2.3 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About h2o-lightning-4b

### How much GPU memory does h2o-lightning-4b need?

About 10.9 GB at 16-bit and 2.7 GB at 4-bit: the weights (4.5B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run h2o-lightning-4b on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use h2o-lightning-4b commercially?

Yes. h2o-lightning-4b is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

### What is h2o-lightning-4b's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Image and text to text

### [Vinci-Piccolo-1.0](https://savrn.com/models/vinci-piccolo-1-0)

[SimpleDirect](https://savrn.com/model-publishers/simpledirect)

Vinci Piccolo is a small, open-weight chat model fine-tuned for character and honesty — the first model in the Vinci family from SimpleDirect. The character you'd want in an AI, open and small enough to run yourself. Try it: chat app — free · ollama run hf.co/simpledirect/Vinci-Piccolo-1.0-GGUF Weight-file size is not a runtime-memory requirement — model loading, KV cache, context length and batching all need memory beyond the weights. See Hardware requirements below for the figures we do give. Most fine-tuning optimizes for capability. Vinci Piccolo is fine-tuned for something else: a consistent character and an honest disposition. It is trained against a written, public Constitution that…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/vinci-piccolo-1-0)

Model · Image and text to text

### [Omni-Edu-4B](https://savrn.com/models/omni-edu-4b)

[Hao Liang](https://savrn.com/model-publishers/lhpku20010120)

This model is a fine-tuned version of Qwen/Qwen3.5-4B-Base on the Omni-Edu-70K dataset. The following hyperparameters were used during training: - learningrate: 5e-06 - trainbatchsize: 1 - evalbatchsize: 8 - distributedtype: multi-GPU - numdevices: 8 - gradientaccumulationsteps: 8 - totaltrainbatchsize: 64 - totalevalbatchsize: 64 - lrschedulertype: cosine - lrschedulerwarmupsteps: 0.1 - numepochs: 3.0 - Transformers 5.2.0 - Pytorch 2.10.0 - Datasets 4.0.0 - Tokenizers 0.22.2

Open weights other 4.5B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/omni-edu-4b)

Model · Image and text to text

### [DN-MOPD-Qwen3.5-4B-baseline-label-160updates](https://savrn.com/models/dn-mopd-qwen3-5-4b-baseline-label-160updates)

[XinLi](https://savrn.com/model-publishers/xinli1997)

The Label baseline of the DN-MOPD paper at Qwen3.5-4B continued to 160 updates (paper Table 5): multi-teacher on-policy distillation with label routing (each prompt is scored by the expert of its domain, every domain multiplier is 1). Released for comparison with DN-MOPD-Qwen3.5-4B; it is not the proposed method. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in recipes/qwen3.5/ and docs/recipe.md. This model was trained and evaluated with the non-thinking chat format. Pass enablethinking=False to the…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/dn-mopd-qwen3-5-4b-baseline-label-160updates)

Model · Image and text to text

### [DN-MOPD-Qwen3.5-4B-160updates](https://savrn.com/models/dn-mopd-qwen3-5-4b-160updates)

[XinLi](https://savrn.com/model-publishers/xinli1997)

A Qwen3.5-4B student trained with DN-MOPD (Domain-Normalized Multi-Teacher On-Policy Distillation) continued to 160 updates (paper Table 5). Three same-size RL experts (math, code, instruction following) teach one student on its own responses; each prompt is scored by the expert of its domain, and DN-MOPD rescales each domain's token-level feedback by its measured spread, wd = clip(σall / σd, 0.25, 4), so that no domain dominates the shared update. Paper: Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation (arXiv:2609.35347, project page) · Code: github.com/LiXin97/DN-MOPD The full recipe, with the launch scripts for every row of the paper's tables, is in…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/dn-mopd-qwen3-5-4b-160updates)

Model · Image and text to text

### [ezjev-4b-s2](https://savrn.com/models/ezjev-4b-s2)

[Everett](https://savrn.com/model-publishers/everettjf)

A typed-decision model (Jev-style /v1/systemone: choice, noul and score questions answered with a probability per option from one forward pass). LoRA fine-tune of Qwen/Qwen3.5-4B, merged into full weights, in two stages: 1. Stage 1 (ezjev-4b): LoRA r=16 on all language-model linear layers (attention, Gated-DeltaNet, MLP), LR 1e-4, one epoch over ~85k rows / ~106k questions. Loss: cross-entropy + Brier over the option labels, llm2jev chat prompt (thinking off). 2. Stage 2: a second LoRA at LR 5e-5 on ~30k rows targeting weak task families (RAG hallucination, product relevance, code-output selection, select-all-that-apply, clinical NLI, phishing emails, claim verification) with 40% replay.…

Open weights other 4.5B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/ezjev-4b-s2)

Model · Image and text to text

### [Omni-Edu-4B-FP8](https://savrn.com/models/omni-edu-4b-fp8)

[Hao Liang](https://savrn.com/model-publishers/lhpku20010120)

Open Foundation Models for Learning and Teaching OmniEdu-4B-FP8 is the FP8-quantized release of OmniEdu-4B. The architecture, tokenizer, chat template and instruction-tuning data are unchanged: this checkpoint is the same model with its linear layers stored in FP8, which shrinks the weights from 8.46 GiB to 5.13 GiB and is intended for deployment on smaller GPUs and for higher-throughput serving. This checkpoint accompanies OmniEdu: Open Foundation Models for Learning and University · University of the Chinese Academy of Sciences · Zhongguancun Academy Quantized with LLM Compressor into the compressed-tensors FP8BLOCK scheme (library version 0.19.0), without calibration data: The vision…

Open weights other 4.5B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/omni-edu-4b-fp8)

## H2O.ai

[All models and datasets](https://savrn.com/model-publishers/h2oai)

## Versions

- [d5c279cf3fa7](https://savrn.com/models/h2o-lightning-4b/versions/d5c279cf3fa7) · current 2026-10-08

## Explore More

- [All image and text to text models](https://savrn.com/models/tasks/image-and-text-to-text)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-08.
- [Hugging Face record](https://huggingface.co/h2oai/h2o-lightning-4b)
- [How the hub is built](https://savrn.com/model-hub/methodology)
