laya-browser-mind2web is an open-weight model from Hitesh S, released under Apache License 2.0. It has 421M parameters. At 16-bit it needs about 1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
A Laya (421M, ModernBERT-large encoder + typed decision heads) checkpoint fine-tuned to pick the next browser action — an operation (CLICK / TYPETEXT / SELECT) and a target element — from a page's DOM and a natural-language goal.
Runs On
What it takes to serve laya-browser-mind2web (421M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.8 GB | 1.0 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.4 GB | 0.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.
laya-browser-mind2web on every accelerator the SAVRN Index prices, at every precision
Model Card
By Hitesh S, published under apache-2.0, revision de7369e52665.
A Laya (421M, ModernBERT-large encoder + typed decision heads) checkpoint fine-tuned to pick the next browser action — an operation (CLICK / TYPETEXT / SELECT) and a target element — from a page's DOM and a natural-language goal. Trained on Mind2Web (Deng et al., NeurIPS 2023 Datasets and Benchmarks Track, CC-BY-4.0). Built as the policy for laya-system-use, a fork of browser-use/jev-ultrafast that replaces the hosted TypeSafe Jev API with this open-weight model. This is a research/hobby-project checkpoint, not a production model. Read the Limitations section before using it. Given a page's interactive elements (label, role, current value) and a goal, it answers two typed questions in one…
Read Hitesh S's full model card
Laya Browser (Mind2Web fine-tune)
A Laya (421M, ModernBERT-large encoder + typed decision heads)
checkpoint fine-tuned to pick the next browser action — an operation (CLICK / TYPE_TEXT / SELECT) and a
target element — from a page's DOM and a natural-language goal. Trained on Mind2Web
(Deng et al., NeurIPS 2023 Datasets and Benchmarks Track, CC-BY-4.0). Built as the policy for
laya-system-use, a fork of
browser-use/jev-ultrafast that replaces the hosted TypeSafe Jev
API with this open-weight model.
This is a research/hobby-project checkpoint, not a production model. Read the Limitations section before using it.
What it does
Given a page's interactive elements (label, role, current value) and a goal, it answers two typed questions in
one forward pass: which operation to perform, and which element to act on. It does not generate text — a
separate small local LLM fills in typed values (TYPE_TEXT), and the element ranking is done by an untrained
lexical shortlister before this model ever sees the candidates (see Methodology).
Methodology
Data pipeline
- Parsing. Mind2Web's raw HTML/candidate format is converted into a neutral
Candidate(id, label, role, value, ops)record. Labels come from the DOM (aria-label, text content,alt,title,placeholder,value,name, in that order), falling back to the element's role when none exist — matching what the live browser snapshot exports, so training and serving never see different inputs. - Shortlisting. Mind2Web pages average ~140-580 candidate elements; the fine-tuned model's context budget only fits ~20. An untrained lexical ranker (rare-word-weighted overlap between the goal/history and each candidate's label, plus a small role prior and page-position prior) keeps the top 20 per operation. This ranker was tuned on Mind2Web's training data only: it raised recall@20 from 0.686 to 0.755 (train) / 0.662 to 0.790 (dev) before any model training began.
- Reclaiming non-interactive gold elements. ~15% of steps have a labeled target that isn't itself
clickable/typeable (an icon inside a button, a
<label>around an input). These are mapped to the nearest interactive descendant or ancestor in the DOM tree, recovering steps that would otherwise be unusable. - Case building. Each usable step becomes 1-2 typed questions (
operation, and<op>_targetwhen more than one candidate survives the shortlist). A step with exactly one candidate needs no question at all — the live agent skips the model entirely in that case, and training treats it as free.
Training
- Base:
convaiinnovations/laya(ModernBERT-large, 421M params), unmodified. - Data: Mind2Web train split, 8,613 usable typed-decision items (from 949 tasks, 69 websites), tokenized at
max_len=768/head_max_len=448(zero items overflowed this budget; longest was 503 tokens). - Method: RLCD — proper-scoring-rule policy gradient (GRPO-style noised-logit sampling) plus soft cross-entropy guidance, following Laya's own fine-tuning recipe.
- Hardware: Kaggle, 2× NVIDIA T4, DDP (
torch.distributed, NCCL). - Schedule: 4 epochs, micro-batch 8, grad-accum 4 (effective batch 64), encoder LR 2.5e-5, head LR 1e-4, cosine schedule, exploration noise σ 0.4→0.1 across epochs. Wall-clock: ~33 minutes.
- Calibration: post-training temperature fitting, per Laya's option-count bucket (
choice:2,choice:3-5,choice:6-10,choice:11+), on held-out dev items — fitted values:[3.86, 2.96, 2.36, 2.16]respectively. (An earlier bug where the base checkpoint's own bucket temperatures silently overrode the fitted ones — making the dominant 11+-option bucket run at T=0.10, i.e. near one-hot overconfidence — was found and fixed before this checkpoint was trained; every bucket here is freshly fit.) - Checkpoint safety: weights and a placeholder-calibration config are saved before calibration runs, so a calibration failure can never discard a finished training run.
What was tried and rejected
Two follow-up experiments were run and did not improve on this checkpoint — included here for honesty, not as a footnote: - More epochs (8 instead of 4). Looked like a clear win on the held-out dev split (element accuracy 0.487 → 0.524 macro), but did not generalize: on the actual test splits it tied or lost to the 4-epoch checkpoint on 2 of 3 splits. Classic dev-set overfitting; not adopted. - LoRA (rank 8, frozen base encoder, same schedule). Training loss stayed far higher throughout (~1.4-1.5 vs ~0.3-1.0 for full fine-tuning) and element accuracy on dev dropped to roughly zero-shot level (0.29, vs 0.67 for full fine-tuning). 0.9% trainable parameters was not enough capacity for this loss within 4 epochs at this rank/LR. Not adopted.
Evaluation
Evaluated on Mind2Web's three official, held-out test splits — cross-task, cross-website, and cross-domain
generalization — with task counts matching the paper exactly (252 / 177 / 912 tasks). No sampling: every step of
every task in each split was scored. Metrics follow the Mind2Web paper's protocol exactly: macro-averaged per
task (a task's score is the mean correctness across all of its steps, then averaged across tasks), with
success_rate requiring every step in a task to succeed. A step this model was never asked about (its gold
element didn't survive the shortlist, or Mind2Web itself gave no usable candidate for it) counts as a hard
failure for these macro metrics — the same standard MindAct implicitly holds itself to.
| Split | Model | op_acc (macro) | element_acc (macro) | Step SR (macro) | Success Rate |
|---|---|---|---|---|---|
| Cross-Task | lexical ranker (no model) | 0.508 | 0.178 | 0.106 | 0.008 |
| Cross-Task | zero-shot Laya (no fine-tune) | 0.466 | 0.178 | 0.101 | 0.000 |
| Cross-Task | this checkpoint | 0.520 | 0.503 | 0.390 | 0.040 |
| Cross-Website | ranker | 0.471 | 0.185 | 0.102 | 0.006 |
| Cross-Website | zero-shot | 0.431 | 0.171 | 0.087 | 0.000 |
| Cross-Website | this checkpoint | 0.478 | 0.482 | 0.329 | 0.023 |
| Cross-Domain | ranker | 0.499 | 0.192 | 0.104 | 0.006 |
| Cross-Domain | zero-shot | 0.444 | 0.187 | 0.103 | 0.006 |
| Cross-Domain | this checkpoint | 0.498 | 0.484 | 0.354 | 0.026 |
Element accuracy roughly triples over both baselines on every split, confirming the gain is from fine-tuning, not from the base checkpoint alone (zero-shot stays near chance throughout).
Against the published MindAct benchmark
MindAct (Deng et al., NeurIPS 2023, Table 2), best backbone (Flan-T5-XL), same macro protocol, pulled directly from the paper:
| Cross-Task | Cross-Website | Cross-Domain | |
|---|---|---|---|
| MindAct (Flan-T5-XL): Ele.Acc / Op F1 / Step SR / SR | 55.1 / 75.7 / 52.0 / 5.2 | 42.0 / 65.2 / 38.9 / 5.1 | 42.1 / 66.5 / 39.6 / 2.9 |
| This checkpoint: Ele.Acc / op_acc / Step SR / SR | 50.3 / 52.0 / 39.0 / 4.0 | 48.2 / 47.8 / 32.9 / 2.3 | 48.4 / 49.8 / 35.4 / 2.6 |
Competitive on element accuracy, behind on step/task success — for two identifiable, honest reasons (not spin):
1. Candidate recall is the ceiling. MindAct's trained DeBERTa-base ranker gets 85-89% Recall@50. This
checkpoint's untrained lexical shortlister gets 73-79% Recall@20 — half the candidate budget, no training.
Element accuracy cannot exceed what survives the shortlist.
2. op_acc isn't Operation F1. MindAct's Op F1 checks the value typed/selected, not just the action type.
This model's text-filling component (a separate local LLM) isn't wired into this evaluation, so op_acc here
only checks CLICK/TYPE_TEXT/SELECT correctness, not typed-value correctness. This makes the comparison to
MindAct's Op F1 column directionally favorable to this checkpoint, not apples-to-apples.
Limitations
- Shortlist recall bounds everything downstream. ~21-27% of steps lose their correct element before this model ever sees it. Improving the shortlister (a trained ranker, a larger K) would likely help more than further fine-tuning.
- No typed-value evaluation. See above —
op_accis action-type accuracy only. - Small, single-domain data. 8,613 items from one dataset, one training run, no cross-validation.
- Live-page gap untested. Evaluated entirely on Mind2Web's dataset snapshots, not on live, rendered pages through the actual browser agent. Distribution shift (viewport-only visibility vs. Mind2Web's full-page candidate lists, real-world DOM noise) is expected but not yet measured.
- Calibrated confidences, not calibrated correctness. Temperature fitting makes the confidence scores meaningful; it does not change which answer is chosen.
Usage
import laya
agent = laya.load("Quantum08/laya-browser-mind2web")
result = agent.system_one(state, questions) # same typed-question API as upstream Laya / TypeSafe Jev
state and questions follow the format built by jev_ultrafast/formatter.py
and jev_ultrafast/shortlister.py
in the source repo — using this checkpoint directly with a different candidate format will not reproduce these
numbers, since the shortlister and label-rendering are load-bearing parts of the pipeline, not just preprocessing.
License
Apache-2.0, inherited from the base model. Mind2Web is CC-BY-4.0 (Deng, Xin, Su, et al. "Mind2Web: Towards a Generalist Agent for the Web." NeurIPS 2023 Datasets and Benchmarks Track).
Identity and Version
- Repository
- Quantum08/laya-browser-mind2web
- Publisher
- Hitesh S
- Task
- Not stated by the source
- Modality
- Other
- Library
- Not stated by the source
- Parameters
- 421M parameters
- Languages
- en
- Revision
- de7369e52665b53cf27bf17332d737090a0ae3da
- First published
- 2026-09-23
- Last updated
- 2026-09-23
Files and Weights
7 files, 846.2 MB in total. The weights are 1 file totalling 842.6 MB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model.safetensors | Weights | 842.6 MB | f0a393c89e89 |
| encoder/config.json | Configuration | 2.1 KB | — |
| rl_agent_config.json | Configuration | 743 B | — |
| README.md | Documentation | 10.1 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer/tokenizer.json | Tokenizer | 3.6 MB | — |
| tokenizer/tokenizer_config.json | Tokenizer | 308 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 842.6 MB
Released by Hitesh S through its official repository on Hugging Face. Read the license.
Built From
- Derived from convaiinnovations/laya
- Trained on (disclosed) osunlp/Mind2Web
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 842.6 MB |
| 16-bit | 0.8 GB |
| 8-bit | 0.4 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About laya-browser-mind2web
How much GPU memory does laya-browser-mind2web need?
About 1 GB at 16-bit and 0.3 GB at 4-bit: the weights (421M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run laya-browser-mind2web on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use laya-browser-mind2web commercially?
Yes. laya-browser-mind2web is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.