FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…
Open-weight model · Text classification
laya-LiteRT
by LiteRT Community (FKA TFLite) litert-community/laya-LiteRT
laya-LiteRT is an open-weight model for text classification from LiteRT Community (FKA TFLite), released under Apache License 2.0. Its published files total 7.5 GB. It draws 117 downloads a month.
A state in, calibrated answers to typed questions out, no text generated: choice (one of the criteria), score (an ordered scale), noul (does the statement hold, as a probability). One forward per question.
Model Card
By LiteRT Community (FKA TFLite), published under apache-2.0, revision ea4c7d415890.
A state in, calibrated answers to typed questions out, no text generated: choice (one of the criteria), score (an ordered scale), noul (does the statement hold, as a probability). One forward per question. These are the convaiinnovations/laya checkpoints (revision 1c5edc17) converted to classic LiteRT graphs with one static window per file: the English model is ModernBERT-large with laya's decision head (421M parameters), the multilingual model is mmBERT-base with the same head (322M; the publisher lists 100+ languages). The request and answer forms are laya's own. Tested on a Samsung Galaxy S26 (SM-S942Q, Android 16 BP4A.251205.006) with LiteRT 2.2.0 CompiledModel, GPU with explicit FP32…
Read LiteRT Community (FKA TFLite)'s full model card
laya decision encoders for LiteRT (Android GPU FP32)
A state in, calibrated answers to typed questions out, no text generated: choice (one of the criteria), score (an ordered scale), noul (does the statement hold, as a probability). One forward per question. These are the convaiinnovations/laya checkpoints (revision 1c5edc17) converted to classic LiteRT graphs with one static window per file: the English model is ModernBERT-large with laya's decision head (421M parameters), the multilingual model is mmBERT-base with the same head (322M; the publisher lists 100+ languages). The request and answer forms are laya's own.
Tested on a Samsung Galaxy S26 (SM-S942Q, Android 16 BP4A.251205.006) with LiteRT 2.2.0 CompiledModel, GPU with explicit FP32 precision, on 2026-09-21. On every captured row the phone's marker scores match the publisher's fp32 Agent.predict (argmax identical, probabilities within 0.0001 after the same calibration), and on the 144 rows of a public three-option decision fixture the phone's choices agree with the publisher's implementation 144/144. Other phones and GPU families have not been tested here.
Files
| File | Checkpoint | Window | Weights | Bytes | Tested on the phone |
|---|---|---|---|---|---|
laya_ml_s256_fp32.tflite |
multilingual | 256 | fp32 | 1,287,375,376 | GPU FP32: 54 ms per question (median, warm); CPU 4 threads: 222 ms |
laya_ml_s512_fp32.tflite |
multilingual | 512 | fp32 | 1,288,292,880 | GPU FP32: 138 ms per question (median, warm); CPU 4 threads: 649 ms (2026-09-23, 201-row parity, 0 flips) |
laya_ml_s256_wfp16.tflite |
multilingual | 256 | float16 | 644,077,088 | CPU 4 threads: 245 ms per question (2026-09-21, 40-row parity, 0 flips). CPU profile only (see below) |
laya_en_s256_fp32.tflite |
English | 256 | fp32 | 1,684,754,688 | GPU FP32: 127 ms per question |
laya_en_s512_fp32.tflite |
English | 512 | fp32 | 1,685,803,264 | GPU FP32: 442 ms per question |
laya_en_s512_wfp16.tflite |
English | 512 | float16 | 843,929,120 | CPU 4 threads: 1,326 ms per question. CPU profile only |
laya_ml_act_head_fp32.tflite, laya_act_head_fp32.tflite |
multilingual, English | any | fp32 | 795,816 / 1,057,960 | the act head of each checkpoint, run after the main graph |
Host files: en/ and multilingual/ hold each checkpoint's exact tokenizer.json, tokenizer_config.json and rl_agent_config.json from the pinned revision; multilingual/calibration.json is a temperature fit made for this conversion (below); laya_host.py is the reference host (laya 0.3.4's prompt builder and decoder, a CPU CompiledModel runner); HOST_CONTRACT.md specifies the sequence and the decoding for a port; hfmodels.json is what the Android SDK below reads; SHA256SUMS lists every file.
wfp16 files store the FULLY_CONNECTED and EMBEDDING_LOOKUP weights as float16 with a DEQUANTIZE to float32; activations stay float32 and inputs are int32 ids. On the tested phone LiteRT 2.2.0's GPU accelerator leaves those DEQUANTIZE nodes (and the EMBEDDING_LOOKUP) to the CPU partition and then fails to compile the model, so the descriptor gives the wfp16 variants a cpu profile only; the fp32 files compile on the GPU. Weight storage and GPU computation precision are separate settings: the tests used explicit FP32 (CompiledModel.GpuOptions(precision = FP32)); the default precision, INT8 and NPU execution were not tested.
Related package. litert-community/Laya-Multilingual-LiteRT holds the same multilingual checkpoint in a second form: the token-embedding lookup moved to the host (inputs_embeds input, a 393 MB float16 table), which lets the float16-weight graph (251 MB) compile on the same phone's GPU at 51 ms per question, plus a standalone Android sample. The graphs here keep the table inside and take token ids.
Graph contract
Main graph, signature serving_default: input_ids [1,N] int32, attention_mask [1,N] float32 (1 real, 0 pad), qtype_onehot [1,3] float32 (choice / score / noul) -> token_logits [1,N] float32, pooled_cls [1,D] float32 (D = 1024 English, 768 multilingual). Act head: pooled_cls [1,D] + feats [1,4] -> act_logits [1,2]. Right padding. The host builds [CLS] <type> question: <instructions> [SEP] [MASK] opt0 [MASK] opt1 … [SEP] <state> [SEP], gathers token_logits at the marker positions, divides by the temperature for the question type and option count, applies softmax, and rounds like the publisher's code. The token embedding table is inside the graph; the tokenizers are external files. laya_host.py reproduces all 611 captured rows (English 209, multilingual 201 per window).
Android
The hfmodels SDK (john-rocky/hfmodels-android, module hfmodels-litert) carries the Kotlin port of the tokenizers, the sequence builder and the decoder, downloads and verifies these files by their sha256, and runs the graphs on CompiledModel:
val model = HfModels(context).fromPretrained(ModelRef("litert-community/laya-LiteRT"), EncoderDecisions) // variant = "en_s256_fp32" for English
val d = model.decide(
mapOf("subject" to "Duplicate charge", "body" to "Please refund the duplicate charge."),
mapOf(
"department" to Question.Choice("Which department should handle this request?", linkedMapOf("billing" to "invoices, payments, refunds", "technical" to "bugs, outages", "other" to "everything else")),
"refund_requested" to Question.Noul("Does the user explicitly request a refund?"),
),
)
(d.answers["department"] as Answer.Choice).choice; (d.answers["refund_requested"] as Answer.Noul).noul; d.timing.questionMs
The SDK's device gate produced the numbers above (logs under litert/results/ there). Its samples/decide shows three uses with the milliseconds on screen: a voice gate, the clipboard before a paste, and a query against passages.
Python
laya_host.py runs the multilingual graph on the CPU with ai-edge-litert and the pinned tokenizer; neither torch nor the laya package is needed.
# pip install transformers==5.17.0 numpy==2.5.3 huggingface_hub==1.32.0 ai-edge-litert==2.1.6
from huggingface_hub import snapshot_download
from laya_host import LayaHost
root = snapshot_download("litert-community/laya-LiteRT", allow_patterns=["laya_ml_s256_wfp16.tflite", "laya_ml_act_head_fp32.tflite", "multilingual/*", "laya_host.py"])
with LayaHost(tokenizer_dir=f"{root}/multilingual", main_graph_path=f"{root}/laya_ml_s256_wfp16.tflite", act_graph_path=f"{root}/laya_ml_act_head_fp32.tflite",
window=256, head_max_len=256, temperatures_json=f"{root}/multilingual/calibration.json") as host:
print(host.predict({"message": "同じ注文の代金が二回引き落とされています。差額の返金をお願いします。"},
{"department": {"type": "choice", "instructions": "Which department should handle this request?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages, system errors", "sales": "pricing, new contracts", "other": "everything else"}}}))
Conversion and verification
Converted with litert-torch 0.9.3 from laya 0.3.4 (torch 2.12.1, transformers 5.17.0): the encoder plus the two decision-head layers as one graph per window, RoPE tables baked per window, exact-erf GELU kept, the marker gathering and the act features moved to the host. Parity against the publisher's fp32 Agent.predict on a Mac CPU (ai-edge-litert 2.1.6): English 209 rows at window 512 (140 at 256), multilingual 201 rows at each window; every fp32 file within 3e-5 in probability, every wfp16 file within 1.5e-3, argmax identical. The phone numbers are one run per row, the model loaded and warm, one question per call; the load (tokenizer, GPU compile) took 5 to 12 seconds. Later runs in the same session were slower (thermal status 1 at 40 C by the end), so the medians come from the first 60 rows of each run.
The multilingual checkpoint ships no temperature calibration (temperature: [1, 1, 1]). multilingual/calibration.json was fitted for this conversion on 4,415 labeled examples from public datasets (Civil Comments, MASSIVE en-US / ja-JP, JGLUE JNLI / JCommonsenseQA / JSTS, JMTEB livedoor; names, splits and terms are recorded inside the file), minimizing the negative log-likelihood over one temperature per question type and option-count bucket; it leaves every argmax unchanged. The choice:6-10 bucket keeps T = 1 (its fit did not meet the fixed calibration-error allowance). The English checkpoint keeps its upstream temperatures. Calibration on the reader's own data will fit better than this generic one.
Limits
Conversion agreement is not task accuracy: on the 144-row public fixture (three-option evidence questions the model was not trained for) the model's own accuracy was 0.59 multilingual and 0.61 English. The act head's probability saturated at 1.0 on the captured rows and is not a usable escalation signal. A state longer than the window is cut at its end. Japanese states go to the multilingual model with English question schemas, the way the publisher uses it; the English model is for English states.
License and attribution
Apache-2.0 for laya and the English ModernBERT-large encoder (Answer.AI and LightOn); MIT for the multilingual mmBERT-base encoder (JHU CLSP). The converted files are modified derivatives of the laya checkpoints: explicit graph rewrites, static windows, separate act heads, and optional float16 weight storage. LICENSE is laya's Apache-2.0 text; the tokenizer files are unchanged copies. Conversion, calibration fit and Android verification by mlboydaisuke.
Identity and Version
- Repository
- litert-community/laya-LiteRT
- Publisher
- LiteRT Community (FKA TFLite)
- Task
- Text classification
- Modality
- Text
- Library
- litert
- Parameters
- Not stated by the source
- Languages
- en, ja
- Revision
- ea4c7d41589027d51076e7ef0b700da103afa476
- First published
- 2026-09-21
- Last updated
- 2026-09-24
Files and Weights
22 files, 7.5 GB in total. The weights are 8 files totalling 7.4 GB in tflite.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| laya_act_head_fp32.tflite | Weights | 1.1 MB | c6f8de9b66e3 |
| laya_en_s256_fp32.tflite | Weights | 1.7 GB | 60cd44a92f70 |
| laya_en_s512_fp32.tflite | Weights | 1.7 GB | 2edc029fed80 |
| laya_en_s512_wfp16.tflite | Weights | 843.9 MB | 882643ae98cb |
| laya_ml_act_head_fp32.tflite | Weights | 795.8 KB | 30da532a8b75 |
| laya_ml_s256_fp32.tflite | Weights | 1.3 GB | 38d63fd6015a |
| laya_ml_s256_wfp16.tflite | Weights | 644.1 MB | 4e56e65788ae |
| laya_ml_s512_fp32.tflite | Weights | 1.3 GB | 1decd8d3e81a |
| en/rl_agent_config.json | Configuration | 745 B | — |
| hfmodels.json | Configuration | 20.7 KB | — |
| laya_host.py | Configuration | 14.0 KB | — |
| multilingual/calibration.json | Configuration | 9.2 KB | — |
| multilingual/rl_agent_config.json | Configuration | 472 B | — |
| HOST_CONTRACT.md | Documentation | 27.6 KB | — |
| LICENSE | Documentation | 10.2 KB | — |
| README.md | Documentation | 10.2 KB | — |
| SHA256SUMS | Other | 1.5 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| en/tokenizer.json | Tokenizer | 3.6 MB | — |
| en/tokenizer_config.json | Tokenizer | 308 B | — |
| multilingual/tokenizer.json | Tokenizer | 34.4 MB | 609d8f4c067c |
| multilingual/tokenizer_config.json | Tokenizer | 524 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 7.4 GB
Released by LiteRT Community (FKA TFLite) through its official repository on Hugging Face. Read the license.
Built From
- Derived from convaiinnovations/laya
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 7.4 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About laya-LiteRT
Can I use laya-LiteRT commercially?
Yes. laya-LiteRT is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.
FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…
https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).
This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.
With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…
