FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…
Open-weight model · Text classification
superfast-tiny-home-robotic
by Sr Aivante sraivante/superfast-tiny-home-robotic
superfast-tiny-home-robotic is an open-weight model for text classification from Sr Aivante, released under Apache License 2.0. Its published files total 179.0 MB.
An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent.
Model Card
By Sr Aivante, published under apache-2.0, revision 8623ed83f99f.
An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent. The package includes a trained 87,181-byte action model, a 1,002,414-device SQLite catalog, a Python API, an HTTP service, a browser playground, Docker files, and reproducible training and evaluation data. The runtime was previously packaged locally as edge-command-docker-v3-ui. The system resolves device names through explicit identifiers, aliases and SQLite FTS5 search. It prioritizes action patterns and capability checks, with a stored count-based unigram/bigram classifier for the fallback scoring path. The…
Read Sr Aivante's full model card
An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent. The package includes a trained 87,181-byte action model, a 1,002,414-device SQLite catalog, a Python API, an HTTP service, a browser playground, Docker files, and reproducible training and evaluation data.
Publisher: sraivante. Release: v1.0.0, 2026-09-26.
The runtime was previously packaged locally as edge-command-docker-v3-ui.
stop the robot arm 2 at factory 1 line 1 station 1
-> {"activity":"robotics","subject":"factory_1_line_1_station_1_robot_arm_2","action":"STOP"}
The system resolves device names through explicit identifiers, aliases and
SQLite FTS5 search. It prioritizes action patterns and capability checks, with
a stored count-based unigram/bigram classifier for the fallback scoring path.
The evaluation below measures the complete hybrid interpreter, not the
learned classifier in isolation. No external base model or neural checkpoint
is used. This is a custom Python package; load it with Interpreter.
Files and requirements
| Component | Size / scope |
|---|---|
action_model.json.gz |
87,181 bytes compressed; 180,309 training rows; 6,892 retained features; 12 action classes |
catalog.sqlite3 |
167,702,528 bytes; 1,001,410 base devices plus 1,004 legacy home devices |
legacy_subjects.json.gz |
3,863 bytes; canonical legacy identifiers |
| Runtime | Python 3 with SQLite FTS5; standard library only; no GPU or network required after download |
edge_command.py, build.py |
Inference, features and original model/catalog construction |
server.py, index.html, compose.yaml, Dockerfile |
HTTP API, browser UI and deployment |
evaluation.json, release_evaluation.json |
Original diagnostic and release-time rerun |
training_snapshot.json, release_manifest.json, SHA256SUMS |
Provenance, pinned dataset reference and file integrity |
edge-command-docker-v3-ui.tar.xz, ORIGINAL_README.md |
Original supplied package and its documentation |
The tiny figure describes the action model. The catalog is a separate 167.7 MB on-disk dependency. SuperFast is the release name, not a measured comparison with other models. No Raspberry Pi latency, memory or field benchmark has been recorded.
Download and run
pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotic', revision='v1.0.0', local_dir='superfast-tiny-home-robotic')"
cd superfast-tiny-home-robotic
python edge_command.py "stop the robot arm 2 at factory 1 line 1 station 1"
After downloading, inference works offline. From the repository directory:
from edge_command import Interpreter
engine = Interpreter()
try:
print(engine.parse("turn on kitchen light"))
# {"activity": "light", "subject": "kitchen_light", "action": "ON"}
finally:
engine.close()
Start the UI and HTTP API:
docker compose up --build -d
# Open http://127.0.0.1:8080/
curl -s http://127.0.0.1:8080/healthz
curl -s -X POST http://127.0.0.1:8080/parse \
-H 'Content-Type: application/json' \
-d '{"instruction":"turn on kitchen light"}'
POST /parse accepts {"instruction":"..."} and returns a parsed intent with
HTTP 200 or a JSON error with HTTP 422. GET /health and GET /healthz check
readiness. Compose binds the host port to loopback. Without Docker,
python server.py starts the same service on port 8080 and binds to all
interfaces; use it on a trusted local network. The service has no authentication.
Tasks, languages and output
Supported input is English and Romanized Hindi/Hinglish. Devanagari, speech recognition, conversation context and translation are not implemented.
{"activity":"robotics","subject":"factory_1_line_1_station_1_gripper","action":"START","requires_confirmation":true}
Successful responses contain activity, a registered subject, and a
catalog-supported action. The action set is ON, OFF, LOW, MEDIUM,
HIGH, OPEN, CLOSE, LOCK, UNLOCK, START, STOP, STATUS.
Sensitive activity categories, except STOP, and every UNLOCK request add
requires_confirmation: true. Failures return an error string and may
include candidate identifiers or a subject.
The interpreter produces intents only. It does not move hardware or obtain
live device state. STATUS requests a future status query; it is not a
measurement. A downstream controller must validate authorization, device
capabilities, state and interlocks. A parsed STOP is not an emergency stop.
Training data and provenance
The matching data is published at
SuperFast Tiny Home Robotic Data.
The immutable dataset commit is 5d997497d2cbd01f7104babd576498a9fe1f0975.
The original build.py thins the 2,141,541-row command source in file order
with random.Random(42) and selection probability 180000 / 2141541.
It selects 180,309 rows, counts unique word/number unigrams and adjacent
bigrams per record and action, and retains features occurring at least three
times subject to its length filter. Training is one streaming count-aggregation
pass, with no optimizer, gradient steps, epochs beyond that pass, or fine-tuning
stage. Runtime scoring uses count-plus-one smoothing. Source identifiers are
not explicitly removed from features. Original build hardware was not recorded.
For this release, replaying that procedure reproduced every stored class total, feature count, feature total, vocabulary size and training-row count exactly. All legacy identifiers and non-timing evaluation results also match. The source, selected rows, sampled diagnostic rows, catalog probes and rejection cases are included with hashes and source-row numbers.
No contemporaneous raw-source hash was stored with the original run. These
matches provide strong reproducibility evidence, while leaving the original
source-byte identity independently unproven. The missing original
devices-1M.jsonl.xz is represented by a lossless logical export of its stored
device fields from catalog.sqlite3; original formatting, discarded fields
and the catalog generator are unavailable. These distinctions are recorded in
training_snapshot.json.
Evaluation
| Diagnostic | Exact / total | Result |
|---|---|---|
| Sampled home-command triples | 928 / 1,000 | 92.8% |
| Registered catalog action/ID probes | 340 / 340 | 100% |
| Expected rejection cases | 6 / 6 | 100% |
The release-time rerun reproduced the original results and saved all predictions in the dataset repository. The template sample uses reservoir sampling with seed 198 from the same source used for training; 82 of 1,000 rows are also selected training rows. There is no independent held-out evaluation. Exact match checks the three target fields and permits an extra confirmation flag. Catalog probes use exact registered IDs and exercise capability coverage. The six rejection cases provide narrow regression coverage. Neither establishes generalization to real speech or safe hardware operation.
The template diagnostic includes 15 incorrect triples, 29 ambiguous-device rejections, 23 unknown/unspecified-device rejections and 5 unsupported-action rejections. Sample failures and the full predictions are published.
Reproduce
Download the matching dataset beside this repository:
python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotic-data', revision='5d997497d2cbd01f7104babd576498a9fe1f0975', local_dir='../superfast-tiny-home-robotic-data')"
python build.py ../superfast-tiny-home-robotic-data/source/devices-1M-recovered.jsonl.xz ../superfast-tiny-home-robotic-data/source/home-commands-v2-part1.jsonl.xz --out rebuilt
python evaluate.py ../superfast-tiny-home-robotic-data/source/home-commands-v2-part1.jsonl.xz --output evaluation-rerun.json
python edge_command.py --path rebuilt "turn on kitchen light"
evaluate.py evaluates the packaged model in its own directory. The last
command separately exercises the rebuilt directory. Compare decompressed
model JSON and logical catalog rows when checking a rebuild: gzip timestamps,
dictionary ordering and SQLite versions can change physical file hashes.
Limitations
- Synthetic templates dominate the command data. Novel wording, ASR errors, ambiguous aliases and conflicting instructions can fail or return wrong intents.
- Negation and multiple-command rejection use limited patterns. They are not comprehensive language understanding or a security boundary.
- Supports a single catalog-constrained command. Conditional rules, motion trajectories, destinations, numeric speed values and manipulation planning are not represented by this package.
- The catalog contains predefined identifiers and capabilities. New devices require an explicitly configured catalog; it is not live device discovery.
- No hardware operation, permissions system, telemetry integration or real-time control loop is provided. No comparative speed or field-reliability claim is made.
License and attribution
Apache License 2.0 applies to the user's original model, code, dataset material
and selection/arrangement. Copyright (c) 2026 sraivante. Third-party ownership,
licenses and runtime dependencies remain separate; the copyright statement
does not claim them. No third-party model weights are included. See LICENSE
and NOTICE. Dataset generator provenance and remaining historical uncertainty
are documented in the linked dataset card.
Identity and Version
- Repository
- sraivante/superfast-tiny-home-robotic
- Publisher
- Sr Aivante
- Task
- Text classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- Not stated by the source
- Languages
- en, hi
- Revision
- 8623ed83f99f53f6917ffbf8b430c0aeecd6b262
- First published
- 2026-09-26
- Last updated
- 2026-09-26
Files and Weights
22 files, 179.0 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| build.py | Configuration | 3.6 KB | — |
| compose.yaml | Configuration | 361 B | — |
| edge_command.py | Configuration | 9.8 KB | — |
| evaluate.py | Configuration | 3.4 KB | — |
| evaluation.json | Configuration | 7.2 KB | — |
| release_evaluation.json | Configuration | 7.7 KB | — |
| release_manifest.json | Configuration | 3.4 KB | — |
| server.py | Configuration | 2.9 KB | — |
| training_snapshot.json | Configuration | 9.9 KB | — |
| LICENSE | Documentation | 11.4 KB | — |
| NOTICE | Documentation | 780 B | — |
| ORIGINAL_README.md | Documentation | 5.2 KB | — |
| README.md | Documentation | 10.8 KB | — |
| Dockerfile | Other | 282 B | — |
| SHA256SUMS | Other | 1.7 KB | — |
| action_model.json.gz | Other | 87.2 KB | cd6ee9393112 |
| catalog.sqlite3 | Other | 167.7 MB | 4dafbdda19e2 |
| edge-command-docker-v3-ui.tar.xz | Other | 11.2 MB | 47c8881bddc3 |
| index.html | Other | 9.8 KB | — |
| legacy_subjects.json.gz | Other | 3.9 KB | 251e170adc38 |
| .dockerignore | Repository | 63 B | — |
| .gitattributes | Repository | 151 B | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
Released by Sr Aivante through its official repository on Hugging Face. Read the license.
Built From
- Trained on (disclosed) sraivante/superfast-tiny-home-robotic-data
Evaluations
Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.
| Benchmark | Conditions | Result | Reported by | Revision | Date |
|---|---|---|---|---|---|
| Template diagnostic (1,000 rows; overlaps training) | Configuration commandsTask Catalog-constrained instruction to device intentMetric Exact activity, subject and action match (hybrid…Comparison conditions not established | 92.8 | sraivante Publisher reported |
Evaluated revision not stated | — |
Questions About superfast-tiny-home-robotic
Can I use superfast-tiny-home-robotic commercially?
Yes. superfast-tiny-home-robotic is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
Similar Models
This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.
FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…
https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).
This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.
With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…
