SAVRN
Search Contact SAVRN

h2o-lightning-4b · Model Card

h2o-lightning-4b: Model Card

Written by H2O.ai, published under apache-2.0, revision d5c279cf3fa7, read 2026-10-08. Shown as written; SAVRN's own facts about this model are on its page.

H2O-Lightning-4B is a 4B-parameter decision model from H2O.ai, built on Qwen/Qwen3.5-4B. It answers typed decision questions about a record (a document, a ticket, a policy, a conversation) and, from v1.2, about images that come with the record (photos, screenshots, scanned documents, charts). It returns a probability for every option:

  • choice: picks one of a set of named options;
  • yes/no (noul): gives the probability that a statement is true;
  • score: gives an ordinal level on a stated scale, with its probability distribution.

It runs on unmodified vLLM 0.30.0 with a small standard-library shim in front (h2o_lightning_shim.py). Each decision is one forward pass and one output token, so the cost is input tokens only.

#1 on JevBench (October 7, 2026)

On JevBench v1.6.1, H2O-Lightning-4B v1.1 is #1 on the official Composite Score of open-weight systems, and it outscores Jev itself: 72.5 to 71.5. It is also ahead of every open 12B, 26B and 31B model on the board. (leaderboard, read 2026-10-07)

Composite Intelligence Calibration Speed Cost $ per 1,000 decisions
H2O-Lightning-4B v1.1 (4B) 72.5 60.0 90.0 92.6 60.3 $0.021 (est.)
Jev 1.13.0 (reference, hosted) 71.5 63.6 90.6 91.5 54.7 $0.032
Quyet-1.0-Large (31B) 71.4 73.4 90.0 86.9 50.5 $0.045 (est.)

The composite is the equal-weight harmonic mean of the four axes. The 31B models lead on raw intelligence, while this 4B matches their calibration at a fraction of the cost and latency.

Four axes, against a 31B Capability against cost

How we got there: - A held-out lockbox: independently written evaluation items that are never trained on, used to compare candidate models before a release. - Rules written down first: ship / no-ship criteria are fixed before results are seen. - No benchmark test items in training: all training data is deduplicated against every evaluation set, exact and near-duplicate. - Gap hunting: comparing against other systems by topic and question type, running adversarial sets, and filling the gaps with new, verified data. - Data quality first: real, licence-clean data, with labels from human annotation or agreement among several independent strong models. - Calibration as a first-class goal: one forward pass, one output token, and probabilities that mean what they say.

The board measured v1.1, the text model. v1.2.1 adds image input; a request without an image takes the same text path.

Measured

Measured on one NVIDIA RTX PRO 4500 Blackwell (32 GB) with serve.sh as shipped (vLLM 0.30.0; text temperature 0.8, image temperature 0.65, images up to 1,048,576 pixels).

JevBench public items result
JevBench (text) the 231 public items, JevBench's own CLI 205 / 231 (88.7 %): easy 48/48 · standard 71/72 · hard 86/111
ImageJevBench the 8 published examples (the only public items released with images) 8 / 8
image validation sets 1,877 questions over 16 public licence-clean sets (table below) 87.4 %

Details follow.

Text: JevBench's public set

From JevBench's own CLI, with --adapter typesafe, on the 231 public items (datasets/public/easy.jsonl, original.jsonl and hard.jsonl; dataset hash dc3995d8…) at JevBench commit bb05a33, with the commands below.

RTX PRO 4500 Blackwell 32 GB
correct 205 / 231 (88.7 %)
by tier: easy / standard / hard 48/48 · 71/72 · 86/111
answered and valid 231 / 231
yes/no answers with P(yes) strictly between 0.2 and 0.8 0 of 74
mean input tokens per decision 654
latency p50 / p95, standard tier, serial 0.033 s / 0.035 s
latency p50 / p95, all 231 items, serial 0.034 s / 0.221 s

Latency depends on the GPU.

Identity check for an evaluator: a correct setup reproduces 205/231 with the tier split above. Five hard items sit on near-ties (top two within 0.03) that GPU numerics can flip. Four are wrong here: hard-opus-a-temporal_numeric-07, hard-opus-a-temporal_numeric-12, hard-opus-c-temporal_numeric-04 and hard-sol-b-temporal_numeric-03. One is right: hard-sol-c-judge_hard-13.

Images

(a) The published ImageJevBench examples: 8 / 8. ImageJevBench's own harness is not public, and only 8 of its items are published with their images (on the benchmark's page: two everyday photos, four screenshots with five labelled click markers, one geometry diagram, one financial table). We read all 8 through this API: 8 correct. These 8 items are not a benchmark score.

(b) Public validation sets with human or source ground truth. Multiple-choice questions we built from each dataset's own annotations (human labels, boxes, receipt fields, table cells, a chart's own data); none come from a model. None of these images or questions was used for training. Read through this API, one image per request, as a data URI in state.

set (licence) what is asked n correct ECE
Open Images V7 validation (CC BY 4.0 annotations) is there a ‹class› in the photo (human-verified labels) 150 88.0 % 0.034
Open Images V7 validation how many ‹class› in the photo (human boxes, 2-8) 150 63.3 % 0.121
VizWiz validation (CC BY 4.0) can the question be answered from this photo 120 84.2 % 0.057
VizWiz validation the answer (6+ of 10 annotators) 120 100.0 % 0.015
TextVQA validation (CC BY 4.0) text in the photo (6+ of 10 annotators) 120 99.2 % 0.024
CORD v2 test + validation (CC BY 4.0) receipt total 95 99.0 % 0.014
CORD v2 test + validation one item's price 29 96.5 % 0.043
CORD v2 test + validation number of item lines 113 85.0 % 0.050
TAT-QA dev (CC BY 4.0) financial-report table questions (the table shown as an image) 120 95.8 % 0.037
DocLayNet test (CDLA-Permissive-1.0) which section heading is on the page 120 100.0 % 0.011
DocLayNet test how many tables are on the page (0-3) 120 75.8 % 0.057
DocLayNet test does the page contain a picture 120 90.0 % 0.039
Our World in Data charts (CC BY 4.0) which country is highest or lowest (the chart's own data) 100 98.0 % 0.058
Our World in Data charts did a value rise between two years 100 95.0 % 0.069
Our World in Data charts roughly what is a value 100 98.0 % 0.066
Rico app screens (CC BY 4.0) which of 5 marked elements reaches a goal (human widget captions) 200 65.0 % 0.094
all 1,877 87.4 % 0.017

ECE: 10-bin expected calibration error of the top answer's probability.

Run it

Hardware measured: one RTX PRO 4500 Blackwell (32 GB); the weights are 9.1 GB. Install vLLM 0.30.0 in a fresh virtual environment (for example python3 -m venv .venv && . .venv/bin/activate):

pip install "vllm==0.30.0+cu129" "torchcodec==0.16.0+cu129" --extra-index-url https://wheels.vllm.ai/0.30.0/cu129 --extra-index-url https://download.pytorch.org/whl/cu129

The shim needs Python 3.8 or later and nothing else.

hf download h2oai/h2o-lightning-4b --revision v1.2.2 --local-dir h2o-lightning-4b

1. Serve the model. Keep config.json as shipped: its "head_dtype": "float32" makes vLLM compute the logits in fp32. Inputs over the context limit get HTTP 422. The last two options set the image limits (up to 4 images per request, each scaled to at most 1.6 megapixels inside vLLM).

vllm serve ./h2o-lightning-4b --served-model-name h2oai/h2o-lightning-4b --host 127.0.0.1 --port 8000 \
  --max-model-len 40960 --gpu-memory-utilization 0.90 \
  --limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'

If vLLM stops at startup because FlashInfer cannot compile its sampling kernels (an older system CUDA toolkit), start it with VLLM_USE_FLASHINFER_SAMPLER=0 in the environment: the decisions do not use the sampler.

2. Start the shim on the same machine:

python3 h2o-lightning-4b/h2o_lightning_shim.py --config h2o-lightning-4b/serve_config.json --vllm http://127.0.0.1:8000 --port 8741

bash h2o-lightning-4b/serve.sh runs steps 1 and 2 together. Before answering, the shim checks that vLLM serves h2oai/h2o-lightning-4b and that every label is one token at the answer slot; if not, it answers 503, so a misconfigured server stops a run instead of scoring it. Before the first image request it also checks that vLLM's chat template renders the same prompt as the text path.

3. Warm up until this returns HTTP 200:

curl -s http://127.0.0.1:8741/v1/systemone -H 'Content-Type: application/json' -d '{"state": "warm-up",
  "questions": {"decision": {"type": "noul", "instructions": "Is this a warm-up?",
  "criteria": {"true": "yes", "false": "no"}}}}'

4. Run JevBench's harness:

cd <JEVBENCH> && python3 -m jevbench.cli run \
  --tasks datasets/public/easy.jsonl,datasets/public/original.jsonl,datasets/public/hard.jsonl \
  --adapter typesafe --endpoint http://127.0.0.1:8741 --key-env '' --model h2oai/h2o-lightning-4b \
  --cost-basis self_hosted_gpu --reserve-usd 0 --results <OUT>/results.jsonl --raw-dir <OUT>/raw \
  --ledger <OUT>/ledger.jsonl --manifest <OUT>/manifest.json --run-label h2o-lightning-4b --delay-s 0

The API. Send requests to POST /v1/systemone. One request can carry several questions about the same state:

{"state": "Ticket 4471: since 09:10 the checkout page returns HTTP 500 for every customer; no orders have gone through in 40 minutes. Severity guide: P1 = a revenue path fully down; P2 = degraded but working; P3 = cosmetic.",
 "questions": {
   "priority":  {"type": "choice", "instructions": "Which priority does the severity guide assign?",
                 "criteria": {"p1": "Revenue path fully down", "p2": "Degraded but working", "p3": "Cosmetic"}},
   "all_users": {"type": "noul", "instructions": "The problem affects every customer.",
                 "criteria": {"true": "Affects everyone", "false": "Affects only some"}},
   "urgency":   {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Low", "Medium", "High"]}}}

The response (rounded here):

{"model": "h2oai/h2o-lightning-4b",
 "answers": {"priority": {"type": "choice", "choice": "p1",
                          "probabilities": {"p1": 0.944, "p2": 0.040, "p3": 0.017}, "confidence": 0.916},
             "all_users": {"type": "noul", "noul": 0.929},
             "urgency": {"type": "score", "score": 1.789, "legend": {"0": "Low", "1": "Medium", "2": "High"},
                         "probabilities": {"0": 0.033, "1": 0.146, "2": 0.821}, "confidence": 0.732}},
 "probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
 "usage": {"input_tokens": 231, "output_tokens": 0, "latency_ms": 70.3}}

usage.input_tokens counts the prompt head that the questions share once.

Images. Put each image in the request as a data:image/...;base64,... URI. It can go anywhere inside state (as a string, an object value or a list item), or in a top-level images list next to state and questions. The images list also takes bare base64 strings and http(s) URLs (vLLM fetches a URL itself, so the server needs network access for those; a URL it cannot fetch gets HTTP 422). Chat-style parts ({"type": "image_url", "image_url": {"url": ...}}) and multipart/form-data (the JSON in a request field, each image as a file part) work too. - Each image taken out of state leaves [image N] in its place, so the record can refer to it. Images in the top-level list are numbered first. - Up to 4 images per request and 20 MB per image. - Each image is scaled to at most 1,048,576 pixels, and image answers use their own temperature, 0.65 (serve_config.json, "image"); text answers keep 0.8. - A request without an image takes exactly the text path above.

This example uses example_receipt.png from this repository:

import base64, json, urllib.request
img = base64.b64encode(open("h2o-lightning-4b/example_receipt.png", "rb").read()).decode()
req = {"state": {"receipt": "data:image/png;base64," + img,
                 "claim": {"employee": "J. Ortiz", "category": "office supplies", "amount": "444.26"}},
       "questions": {
         "matches":  {"type": "noul", "instructions": "The receipt total equals claim.amount.",
                      "criteria": {"true": "Same amount", "false": "Different amount"}},
         "items":    {"type": "choice", "instructions": "How many item lines are on the receipt?",
                      "criteria": {"three": "3", "four": "4", "five": "5", "six": "6"}},
         "legible":  {"type": "score", "instructions": "How legible is the receipt?",
                      "criteria": ["Unreadable", "Partly readable", "Fully readable"]}}}
r = urllib.request.urlopen(urllib.request.Request("http://127.0.0.1:8741/v1/systemone", json.dumps(req).encode(),
                                                  {"Content-Type": "application/json"}))
print(json.dumps(json.load(r), indent=1))

The response (rounded here):

{"model": "h2oai/h2o-lightning-4b",
 "answers": {"matches": {"type": "noul", "noul": 0.964},
             "items": {"type": "choice", "choice": "six",
                       "probabilities": {"three": 0.006, "four": 0.012, "five": 0.024, "six": 0.958}, "confidence": 0.943},
             "legible": {"type": "score", "score": 1.926,
                         "legend": {"0": "Unreadable", "1": "Partly readable", "2": "Fully readable"},
                         "probabilities": {"0": 0.011, "1": 0.052, "2": 0.937}, "confidence": 0.906}},
 "probability_source": "native", "effort_used": {"tier": "fast", "passes_per_decision": 1},
 "usage": {"input_tokens": 585, "output_tokens": 0, "latency_ms": 141.0, "images": 1}}

usage.input_tokens includes the image tokens; usage.images counts the images read.

Status codes. An input the shim cannot take gets HTTP 422, which the runner scores as one wrong answer. That covers: - over the context limit; - more than 255 options; - an unknown question type or a malformed question; - more than 4 images, or an image that is not valid base64 image data.

An empty request gets 400. vLLM's 401, 403 and 429 pass through. Any other vLLM failure is a 502, which counts toward the runner's three-consecutive-failures stop.

Tests (no GPU): python3 h2o-lightning-4b/test_shim.py.

Hybrid mode: decisions and text from one server

adapter/ holds H2O-Lightning-4B's decision adapter as a LoRA for the base model Qwen/Qwen3.5-4B. Serve the base model with the adapter on vLLM, and one server answers both kinds of request: - decisions with model: "h2oai/h2o-lightning-4b" (the adapter), through the same shim and readout as the merged model; - free text with model: "Qwen/Qwen3.5-4B" (the base model, no adapter).

On JevBench's 231 public items, the adapter scores the same as the merged model (205/231) and gives the same top answer on 230 of 231 items. Measured on one RTX A6000: 94 ms per decision (p50, versus 60 ms for the merged model); about 1.3 s for a 90-token reply from the base model.

0. Get the adapter into the same folder as the download above (shim and serve_config.json included):

hf download h2oai/h2o-lightning-4b --include "adapter/*" --local-dir h2o-lightning-4b && cd h2o-lightning-4b

1. Serve the base model with the adapter (vLLM 0.30.0, from the folder that contains adapter/):

VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve Qwen/Qwen3.5-4B --revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a --host 127.0.0.1 --port 8000 \
  --max-model-len 40960 --gpu-memory-utilization 0.90 --max-num-seqs 64 \
  --enable-lora --max-lora-rank 64 --max-loras 1 --lora-modules h2oai/h2o-lightning-4b=./adapter \
  --limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'
  • --max-num-seqs: Qwen3.5's linear-attention layers need one state block per running sequence, so keep this at or below what fits. vLLM stops at startup with a clear message if it's too high.
  • VLLM_USE_FLASHINFER_SAMPLER=0 is needed only where FlashInfer can't compile its sampler (an older CUDA toolkit). The decisions don't use the sampler.

2. Put the shim in front for decisions. Use the same shim and config as the merged model. The LoRA's name is the served model name the shim expects:

python3 h2o_lightning_shim.py --config serve_config.json --vllm http://127.0.0.1:8000 --port 8741

Decisions then go to POST http://127.0.0.1:8741/v1/systemone exactly as described above.

3. Text from the base model on the same server:

curl -s http://127.0.0.1:8000/v1/chat/completions -H 'Content-Type: application/json' -d '{"model": "Qwen/Qwen3.5-4B",
  "messages": [{"role": "user", "content": "Write a two-sentence status update for customers about a checkout outage that is now fixed."}],
  "max_tokens": 200, "chat_template_kwargs": {"enable_thinking": false}}'

Check that the adapter is active. The two lines must differ; if they are identical, vLLM is serving the base model under the adapter's name:

for m in Qwen/Qwen3.5-4B h2oai/h2o-lightning-4b; do curl -s http://127.0.0.1:8000/v1/completions -H 'Content-Type: application/json' -d "{\"model\": \"$m\", \"prompt\": \"Ticket: checkout is down for every customer. Priority (P1, P2 or P3):\", \"max_tokens\": 1, \"temperature\": 0, \"logprobs\": 5}" | python3 -c "import json,sys; print(sys.argv[1], json.load(sys.stdin)['choices'][0]['logprobs']['top_logprobs'][0])" "$m"; done

Must-rename note for converted or re-exported adapters. vLLM loads Qwen3.5-4B as Qwen3_5ForConditionalGeneration, so the adapter's tensor names must be base_model.model.model.language_model.layers.N.…. The file shipped here already uses those names. An adapter saved from a text-only load of the model (base_model.model.model.layers.N.…) is accepted by vLLM without any warning and changes nothing: the output equals the base model. Rename the keys (insert language_model. after model.model.) and run the check above.

The adapter's own scaling is the standard lora_alpha / r. Use it as is, with no extra scale.

Limitations

  • Evaluated mostly in English. Decides best with up to about 16 options; up to 255 are accepted.
  • The probabilities are calibrated for decision questions of this kind, not for open-ended text.
  • Not a chat model: it answers one decision per question and does not generate explanations.
  • Images: no video input; counting many small objects and reading values off line charts are the weakest image skills we measured.

Disclosures

  • No JevBench or ImageJevBench data, public or otherwise, was used for training.
  • No rules keyed to any benchmark item's wording, id or answer; no network calls at serving time (except fetching an image URL a request supplies).
  • Price basis for an evaluator: Qwen3.5-4B at bf16, one output token per decision; 654 mean input tokens on the public text set. Image tokens are counted in usage.input_tokens.

License

Apache-2.0. Built on Qwen/Qwen3.5-4B by the Qwen team (Alibaba Cloud), Apache-2.0; served with unmodified vLLM (Apache-2.0).