h2o-lightning-4b · Model Card
h2o-lightning-4b: Model Card
Written by H2O.ai, published under apache-2.0, revision d5c279cf3fa7, read 2026-10-08. Shown as written; SAVRN's own facts about this model are on its page.
H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option:
- choice: picks one of a set of named options;
- yes/no (
noul): gives the probability that a statement is true; - score: gives an ordinal level on a stated scale, with its probability distribution.
It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2o_lightning_shim.py). Each
decision is one forward pass and one output token, so the cost is input tokens only.
#1 on JevBench (October 7, 2026)
On JevBench v1.6.1, H2O-Lightning-4B v1.1 is #1 on the official Composite Score of open-weight systems, and it outscores Jev itself: 72.5 to 71.5. It is also ahead of every open 12B, 26B and 31B model on the board. (leaderboard, read 2026-10-07)
| Composite | Intelligence | Calibration | Speed | Cost | $ per 1,000 decisions | |
|---|---|---|---|---|---|---|
| H2O-Lightning-4B v1.1 (4B) | 72.5 | 60.0 | 90.0 | 92.6 | 60.3 | $0.021 (est.) |
| Jev 1.13.0 (reference, hosted) | 71.5 | 63.6 | 90.6 | 91.5 | 54.7 | $0.032 |
| Quyet-1.0-Large (31B) | 71.4 | 73.4 | 90.0 | 86.9 | 50.5 | $0.045 (est.) |
The composite is the equal-weight harmonic mean of the four axes. The 31B models lead on raw intelligence, while this 4B matches their calibration at a fraction of the cost and latency.
| Four axes, against a 31B | Capability against cost |
|---|---|
How we got there: - A held-out lockbox: independently written evaluation items that are never trained on, used to compare candidate models before a release. - Rules written down first: ship / no-ship criteria are fixed before results are seen. - No benchmark test items in training: all training data is deduplicated against every evaluation set, exact and near-duplicate. - Gap hunting: comparing against other systems by topic and question type, running adversarial sets, and filling the gaps with new, verified data. - Data quality first: real, licence-clean data, with labels from human annotation or agreement among several independent strong models. - Calibration as a first-class goal: one forward pass, one output token, and probabilities that mean what they say.
The board measured v1.1, the text model. v1.2.1 adds image input; a request without an image takes the same text path.
Measured
Measured on one NVIDIA RTX PRO 4500 Blackwell (32 GB) with serve.sh as shipped (vLLM 0.30.0; text temperature
0.8, image temperature 0.65, images up to 1,048,576 pixels).
| JevBench | public items | result |
|---|---|---|
| JevBench (text) | the 231 public items, JevBench's own CLI | 205 / 231 (88.7 %): easy 48/48 · standard 71/72 · hard 86/111 |
| ImageJevBench | the 8 published examples (the only public items released with images) | 8 / 8 |
| image validation sets | 1,877 questions over 16 public licence-clean sets (table below) | 87.4 % |
Details follow.
Text: JevBench's public set
From JevBench's own CLI, with --adapter typesafe, on the 231 public items (datasets/public/easy.jsonl,
original.jsonl and hard.jsonl; dataset hash dc3995d8…) at JevBench commit bb05a33, with the commands below.
| RTX PRO 4500 Blackwell 32 GB | |
|---|---|
| correct | 205 / 231 (88.7 %) |
| by tier: easy / standard / hard | 48/48 · 71/72 · 86/111 |
| answered and valid | 231 / 231 |
| yes/no answers with P(yes) strictly between 0.2 and 0.8 | 0 of 74 |
| mean input tokens per decision | 654 |
| latency p50 / p95, standard tier, serial | 0.033 s / 0.035 s |
| latency p50 / p95, all 231 items, serial | 0.034 s / 0.221 s |
Latency depends on the GPU.
Identity check for an evaluator: a correct setup reproduces 205/231 with the tier split above. Five hard items
sit on near-ties (top two within 0.03) that GPU numerics can flip. Four are wrong here:
hard-opus-a-temporal_numeric-07, hard-opus-a-temporal_numeric-12, hard-opus-c-temporal_numeric-04 and
hard-sol-b-temporal_numeric-03. One is right: hard-sol-c-judge_hard-13.
Images
(a) The published ImageJevBench examples: 8 / 8. ImageJevBench's own harness is not public, and only 8 of its items are published with their images (on the benchmark's page: two everyday photos, four screenshots with five labelled click markers, one geometry diagram, one financial table). We read all 8 through this API: 8 correct. These 8 items are not a benchmark score.
(b) Public validation sets with human or source ground truth. Multiple-choice questions we built from each
dataset's own annotations (human labels, boxes, receipt fields, table cells, a chart's own data); none come from a
model. None of these images or questions was used for training. Read through this API, one image per request, as a
data URI in state.
| set (licence) | what is asked | n | correct | ECE |
|---|---|---|---|---|
| Open Images V7 validation (CC BY 4.0 annotations) | is there a ‹class› in the photo (human-verified labels) | 150 | 88.0 % | 0.034 |
| Open Images V7 validation | how many ‹class› in the photo (human boxes, 2-8) | 150 | 63.3 % | 0.121 |
| VizWiz validation (CC BY 4.0) | can the question be answered from this photo | 120 | 84.2 % | 0.057 |
| VizWiz validation | the answer (6+ of 10 annotators) | 120 | 100.0 % | 0.015 |
| TextVQA validation (CC BY 4.0) | text in the photo (6+ of 10 annotators) | 120 | 99.2 % | 0.024 |
| CORD v2 test + validation (CC BY 4.0) | receipt total | 95 | 99.0 % | 0.014 |
| CORD v2 test + validation | one item's price | 29 | 96.5 % | 0.043 |
| CORD v2 test + validation | number of item lines | 113 | 85.0 % | 0.050 |
| TAT-QA dev (CC BY 4.0) | financial-report table questions (the table shown as an image) | 120 | 95.8 % | 0.037 |
| DocLayNet test (CDLA-Permissive-1.0) | which section heading is on the page | 120 | 100.0 % | 0.011 |
| DocLayNet test | how many tables are on the page (0-3) | 120 | 75.8 % | 0.057 |
| DocLayNet test | does the page contain a picture | 120 | 90.0 % | 0.039 |
| Our World in Data charts (CC BY 4.0) | which country is highest or lowest (the chart's own data) | 100 | 98.0 % | 0.058 |
| Our World in Data charts | did a value rise between two years | 100 | 95.0 % | 0.069 |
| Our World in Data charts | roughly what is a value | 100 | 98.0 % | 0.066 |
| Rico app screens (CC BY 4.0) | which of 5 marked elements reaches a goal (human widget captions) | 200 | 65.0 % | 0.094 |
| all | 1,877 | 87.4 % | 0.017 |
ECE: 10-bin expected calibration error of the top answer's probability.
Run it
Hardware measured: one RTX PRO 4500 Blackwell (32 GB); the weights are 9.1 GB. Install vLLM 0.30.0 in a fresh virtual environment
(for example python3 -m venv .venv && . .venv/bin/activate):
pip install "vllm==0.30.0+cu129" "torchcodec==0.16.0+cu129" --extra-index-url https://wheels.vllm.ai/0.30.0/cu129 --extra-index-url https://download.pytorch.org/whl/cu129
The shim needs Python 3.8 or later and nothing else.
hf download h2oai/h2o-lightning-4b --revision v1.2.2 --local-dir h2o-lightning-4b
1. Serve the model. Keep config.json as shipped: its "head_dtype": "float32" makes vLLM compute the logits in
fp32. Inputs over the context limit get HTTP 422. The last two options set the image limits (up to 4 images
per request, each scaled to at most 1.6 megapixels inside vLLM).
vllm serve ./h2o-lightning-4b --served-model-name h2oai/h2o-lightning-4b --host 127.0.0.1 --port 8000 \
--max-model-len 40960 --gpu-memory-utilization 0.90 \
--limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'
If vLLM stops at startup because FlashInfer cannot compile its sampling kernels (an older system CUDA toolkit), start it
with VLLM_USE_FLASHINFER_SAMPLER=0 in the environment: the decisions do not use the sampler.
2. Start the shim on the same machine:
python3 h2o-lightning-4b/h2o_lightning_shim.py --config h2o-lightning-4b/serve_config.json --vllm http://127.0.0.1:8000 --port 8741
bash h2o-lightning-4b/serve.sh runs steps 1 and 2 together. Before answering, the
shim checks that vLLM serves h2oai/h2o-lightning-4b and that every label is one token at the answer slot; if not, it
answers 503, so a misconfigured server stops a run instead of scoring it. Before the first image request it also checks
that vLLM's chat template renders the same prompt as the text path.
3. Warm up until this returns HTTP 200:
curl -s http://127.0.0.1:8741/v1/systemone -H 'Content-Type: application/json' -d '{"state": "warm-up",
"questions": {"decision": {"type": "noul", "instructions": "Is this a warm-up?",
"criteria": {"true": "yes", "false": "no"}}}}'
4. Run JevBench's harness:
cd <JEVBENCH> && python3 -m jevbench.cli run \
--tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
--adapter typesafe --endpoint http://127.0.0.1:8741 --key-env '' --model h2oai/h2o-lightning-4b \
--cost-basis self_hosted_gpu --reserve-usd 0 --results <OUT>/results.jsonl --raw-dir <OUT>/raw \
--ledger <OUT>/ledger.jsonl --manifest <OUT>/manifest.json --run-label h2o-lightning-4b --delay-s 0
The API. Send requests to POST /v1/systemone. One request can carry several questions about the same state:
{"state": "Ticket 4471: since 09:10 the checkout page returns HTTP 500 for every customer; no orders have gone through in 40 minutes. Severity guide: P1 = a revenue path fully down; P2 = degraded but working; P3 = cosmetic.",
"questions": {
"priority": {"type": "choice", "instructions": "Which priority does the severity guide assign?",
"criteria": {"p1": "Revenue path fully down", "p2": "Degraded but working", "p3": "Cosmetic"}},
"all_users": {"type": "noul", "instructions": "The problem affects every customer.",
"criteria": {"true": "Affects everyone", "false": "Affects only some"}},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Low", "Medium", "High"]}}}
The response (rounded here):
{"model": "h2oai/h2o-lightning-4b",
"answers": {"priority": {"type": "choice", "choice": "p1",
"probabilities": {"p1": 0.944, "p2": 0.040, "p3": 0.017}, "confidence": 0.916},
"all_users": {"type": "noul", "noul": 0.929},
"urgency": {"type": "score", "score": 1.789, "legend": {"0": "Low", "1": "Medium", "2": "High"},
"probabilities": {"0": 0.033, "1": 0.146, "2": 0.821}, "confidence": 0.732}},
"probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
"usage": {"input_tokens": 231, "output_tokens": 0, "latency_ms": 70.3}}
usage.input_tokens counts the prompt head that the questions share once.
Images. Put each image in the request as a data:image/...;base64,... URI. It can go anywhere inside state (as a
string, an object value or a list item), or in a top-level images list next to state and questions. The images
list also takes bare base64 strings and http(s) URLs (vLLM fetches a URL itself, so the server needs network access
for those; a URL it cannot fetch gets HTTP 422). Chat-style parts ({"type": "image_url", "image_url": {"url": ...}}) and multipart/form-data (the JSON in a
request field, each image as a file part) work too.
- Each image taken out of state leaves [image N] in its place, so the record can refer to it. Images in the
top-level list are numbered first.
- Up to 4 images per request and 20 MB per image.
- Each image is scaled to at most 1,048,576 pixels, and image answers use their own temperature, 0.65 (serve_config.json,
"image"); text answers keep 0.8.
- A request without an image takes exactly the text path above.
This example uses example_receipt.png from this repository:
import base64, json, urllib.request
img = base64.b64encode(open("h2o-lightning-4b/example_receipt.png", "rb").read()).decode()
req = {"state": {"receipt": "data:image/png;base64," + img,
"claim": {"employee": "J. Ortiz", "category": "office supplies", "amount": "444.26"}},
"questions": {
"matches": {"type": "noul", "instructions": "The receipt total equals claim.amount.",
"criteria": {"true": "Same amount", "false": "Different amount"}},
"items": {"type": "choice", "instructions": "How many item lines are on the receipt?",
"criteria": {"three": "3", "four": "4", "five": "5", "six": "6"}},
"legible": {"type": "score", "instructions": "How legible is the receipt?",
"criteria": ["Unreadable", "Partly readable", "Fully readable"]}}}
r = urllib.request.urlopen(urllib.request.Request("http://127.0.0.1:8741/v1/systemone", json.dumps(req).encode(),
{"Content-Type": "application/json"}))
print(json.dumps(json.load(r), indent=1))
The response (rounded here):
{"model": "h2oai/h2o-lightning-4b",
"answers": {"matches": {"type": "noul", "noul": 0.964},
"items": {"type": "choice", "choice": "six",
"probabilities": {"three": 0.006, "four": 0.012, "five": 0.024, "six": 0.958}, "confidence": 0.943},
"legible": {"type": "score", "score": 1.926,
"legend": {"0": "Unreadable", "1": "Partly readable", "2": "Fully readable"},
"probabilities": {"0": 0.011, "1": 0.052, "2": 0.937}, "confidence": 0.906}},
"probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
"usage": {"input_tokens": 585, "output_tokens": 0, "latency_ms": 141.0, "images": 1}}
usage.input_tokens includes the image tokens; usage.images counts the images read.
Status codes. An input the shim cannot take gets HTTP 422, which the runner scores as one wrong answer. That covers: - over the context limit; - more than 255 options; - an unknown question type or a malformed question; - more than 4 images, or an image that is not valid base64 image data.
An empty request gets 400. vLLM's 401, 403 and 429 pass through. Any other vLLM failure is a 502, which counts toward the runner's three-consecutive-failures stop.
Tests (no GPU): python3 h2o-lightning-4b/test_shim.py.
Hybrid mode: decisions and text from one server
adapter/ holds H2O-Lightning-4B's decision adapter as a LoRA for the base model
Qwen/Qwen3.5-4B. Serve the base model with the adapter on vLLM, and one
server answers both kinds of request:
- decisions with model: "h2oai/h2o-lightning-4b" (the adapter), through the same shim and readout as the
merged model;
- free text with model: "Qwen/Qwen3.5-4B" (the base model, no adapter).
On JevBench's 231 public items, the adapter scores the same as the merged model (205/231) and gives the same top answer on 230 of 231 items. Measured on one RTX A6000: 94 ms per decision (p50, versus 60 ms for the merged model); about 1.3 s for a 90-token reply from the base model.
0. Get the adapter into the same folder as the download above (shim and serve_config.json included):
hf download h2oai/h2o-lightning-4b --include "adapter/*" --local-dir h2o-lightning-4b && cd h2o-lightning-4b
1. Serve the base model with the adapter (vLLM 0.30.0, from the folder that contains adapter/):
VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve Qwen/Qwen3.5-4B --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --host 127.0.0.1 --port 8000 \
--max-model-len 40960 --gpu-memory-utilization 0.90 --max-num-seqs 64 \
--enable-lora --max-lora-rank 64 --max-loras 1 --lora-modules h2oai/h2o-lightning-4b=./adapter \
--limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'
--max-num-seqs: Qwen3.5's linear-attention layers need one state block per running sequence, so keep this at or below what fits. vLLM stops at startup with a clear message if it's too high.VLLM_USE_FLASHINFER_SAMPLER=0is needed only where FlashInfer can't compile its sampler (an older CUDA toolkit). The decisions don't use the sampler.
2. Put the shim in front for decisions. Use the same shim and config as the merged model. The LoRA's name is the served model name the shim expects:
python3 h2o_lightning_shim.py --config serve_config.json --vllm http://127.0.0.1:8000 --port 8741
Decisions then go to POST http://127.0.0.1:8741/v1/systemone exactly as described above.
3. Text from the base model on the same server:
curl -s http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model": "Qwen/Qwen3.5-4B",
"messages": [{"role": "user", "content": "Write a two-sentence status update for customers about a checkout outage that is now fixed."}],
"max_tokens": 200, "chat_template_kwargs": {"enable_thinking": false}}'
Check that the adapter is active. The two lines must differ; if they are identical, vLLM is serving the base model under the adapter's name:
for m in Qwen/Qwen3.5-4B h2oai/h2o-lightning-4b; do curl -s http://127.0.0.1:8000/v1/completions -H 'Content-Type: application/json' -d "{\"model\": \"$m\", \"prompt\": \"Ticket: checkout is down for every customer. Priority (P1, P2 or P3):\", \"max_tokens\": 1, \"temperature\": 0, \"logprobs\": 5}" | python3 -c "import json,sys; print(sys.argv[1], json.load(sys.stdin)['choices'][0]['logprobs']['top_logprobs'][0])" "$m"; done
Must-rename note for converted or re-exported adapters. vLLM loads Qwen3.5-4B as
Qwen3_5ForConditionalGeneration, so the adapter's tensor names must bebase_model.model.model.language_model.layers.N.…. The file shipped here already uses those names. An adapter saved from a text-only load of the model (base_model.model.model.layers.N.…) is accepted by vLLM without any warning and changes nothing: the output equals the base model. Rename the keys (insertlanguage_model.aftermodel.model.) and run the check above.
The adapter's own scaling is the standard lora_alpha / r. Use it as is, with no extra scale.
Limitations
- Evaluated mostly in English. Decides best with up to about 16 options; up to 255 are accepted.
- The probabilities are calibrated for decision questions of this kind, not for open-ended text.
- Not a chat model: it answers one decision per question and does not generate explanations.
- Images: no video input; counting many small objects and reading values off line charts are the weakest image skills we measured.
Disclosures
- No JevBench or ImageJevBench data, public or otherwise, was used for training.
- No rules keyed to any benchmark item's wording, id or answer; no network calls at serving time (except fetching an image URL a request supplies).
- Price basis for an evaluator: Qwen3.5-4B at bf16, one output token per decision; 654 mean input tokens on the
public text set. Image tokens are counted in
usage.input_tokens.
License
Apache-2.0. Built on Qwen/Qwen3.5-4B by the Qwen team (Alibaba Cloud), Apache-2.0; served with unmodified vLLM (Apache-2.0).