SAVRN
Search Contact SAVRN

Open-weight model · Text classification

superfast-tiny-home-robotic

by Sr Aivante sraivante/superfast-tiny-home-robotic

superfast-tiny-home-robotic is an open-weight model for text classification from Sr Aivante, released under Apache License 2.0. Its published files total 179.0 MB.

An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent.

Parameters—
Context—
Weights179.0 MB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Model Card

By Sr Aivante, published under apache-2.0, revision 8623ed83f99f.

An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent. The package includes a trained 87,181-byte action model, a 1,002,414-device SQLite catalog, a Python API, an HTTP service, a browser playground, Docker files, and reproducible training and evaluation data. The runtime was previously packaged locally as edge-command-docker-v3-ui. The system resolves device names through explicit identifiers, aliases and SQLite FTS5 search. It prioritizes action patterns and capability checks, with a stored count-based unigram/bigram classifier for the fallback scoring path. The…

Read Sr Aivante's full model card

An offline English and Hinglish command interpreter for home devices and registered industrial/robotics devices. It converts one instruction into a catalog-validated JSON intent. The package includes a trained 87,181-byte action model, a 1,002,414-device SQLite catalog, a Python API, an HTTP service, a browser playground, Docker files, and reproducible training and evaluation data.

Publisher: sraivante. Release: v1.0.0, 2026-09-26. The runtime was previously packaged locally as edge-command-docker-v3-ui.

stop the robot arm 2 at factory 1 line 1 station 1
  -> {"activity":"robotics","subject":"factory_1_line_1_station_1_robot_arm_2","action":"STOP"}

The system resolves device names through explicit identifiers, aliases and SQLite FTS5 search. It prioritizes action patterns and capability checks, with a stored count-based unigram/bigram classifier for the fallback scoring path. The evaluation below measures the complete hybrid interpreter, not the learned classifier in isolation. No external base model or neural checkpoint is used. This is a custom Python package; load it with Interpreter.

Files and requirements

Component Size / scope
action_model.json.gz 87,181 bytes compressed; 180,309 training rows; 6,892 retained features; 12 action classes
catalog.sqlite3 167,702,528 bytes; 1,001,410 base devices plus 1,004 legacy home devices
legacy_subjects.json.gz 3,863 bytes; canonical legacy identifiers
Runtime Python 3 with SQLite FTS5; standard library only; no GPU or network required after download
edge_command.py, build.py Inference, features and original model/catalog construction
server.py, index.html, compose.yaml, Dockerfile HTTP API, browser UI and deployment
evaluation.json, release_evaluation.json Original diagnostic and release-time rerun
training_snapshot.json, release_manifest.json, SHA256SUMS Provenance, pinned dataset reference and file integrity
edge-command-docker-v3-ui.tar.xz, ORIGINAL_README.md Original supplied package and its documentation

The tiny figure describes the action model. The catalog is a separate 167.7 MB on-disk dependency. SuperFast is the release name, not a measured comparison with other models. No Raspberry Pi latency, memory or field benchmark has been recorded.

Download and run

pip install huggingface_hub
python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotic', revision='v1.0.0', local_dir='superfast-tiny-home-robotic')"
cd superfast-tiny-home-robotic
python edge_command.py "stop the robot arm 2 at factory 1 line 1 station 1"

After downloading, inference works offline. From the repository directory:

from edge_command import Interpreter

engine = Interpreter()
try:
    print(engine.parse("turn on kitchen light"))
    # {"activity": "light", "subject": "kitchen_light", "action": "ON"}
finally:
    engine.close()

Start the UI and HTTP API:

docker compose up --build -d
# Open http://127.0.0.1:8080/
curl -s http://127.0.0.1:8080/healthz
curl -s -X POST http://127.0.0.1:8080/parse \
  -H 'Content-Type: application/json' \
  -d '{"instruction":"turn on kitchen light"}'

POST /parse accepts {"instruction":"..."} and returns a parsed intent with HTTP 200 or a JSON error with HTTP 422. GET /health and GET /healthz check readiness. Compose binds the host port to loopback. Without Docker, python server.py starts the same service on port 8080 and binds to all interfaces; use it on a trusted local network. The service has no authentication.

Tasks, languages and output

Supported input is English and Romanized Hindi/Hinglish. Devanagari, speech recognition, conversation context and translation are not implemented.

{"activity":"robotics","subject":"factory_1_line_1_station_1_gripper","action":"START","requires_confirmation":true}

Successful responses contain activity, a registered subject, and a catalog-supported action. The action set is ON, OFF, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK, START, STOP, STATUS. Sensitive activity categories, except STOP, and every UNLOCK request add requires_confirmation: true. Failures return an error string and may include candidate identifiers or a subject.

The interpreter produces intents only. It does not move hardware or obtain live device state. STATUS requests a future status query; it is not a measurement. A downstream controller must validate authorization, device capabilities, state and interlocks. A parsed STOP is not an emergency stop.

Training data and provenance

The matching data is published at SuperFast Tiny Home Robotic Data. The immutable dataset commit is 5d997497d2cbd01f7104babd576498a9fe1f0975.

The original build.py thins the 2,141,541-row command source in file order with random.Random(42) and selection probability 180000 / 2141541. It selects 180,309 rows, counts unique word/number unigrams and adjacent bigrams per record and action, and retains features occurring at least three times subject to its length filter. Training is one streaming count-aggregation pass, with no optimizer, gradient steps, epochs beyond that pass, or fine-tuning stage. Runtime scoring uses count-plus-one smoothing. Source identifiers are not explicitly removed from features. Original build hardware was not recorded.

For this release, replaying that procedure reproduced every stored class total, feature count, feature total, vocabulary size and training-row count exactly. All legacy identifiers and non-timing evaluation results also match. The source, selected rows, sampled diagnostic rows, catalog probes and rejection cases are included with hashes and source-row numbers.

No contemporaneous raw-source hash was stored with the original run. These matches provide strong reproducibility evidence, while leaving the original source-byte identity independently unproven. The missing original devices-1M.jsonl.xz is represented by a lossless logical export of its stored device fields from catalog.sqlite3; original formatting, discarded fields and the catalog generator are unavailable. These distinctions are recorded in training_snapshot.json.

Evaluation

Diagnostic Exact / total Result
Sampled home-command triples 928 / 1,000 92.8%
Registered catalog action/ID probes 340 / 340 100%
Expected rejection cases 6 / 6 100%

The release-time rerun reproduced the original results and saved all predictions in the dataset repository. The template sample uses reservoir sampling with seed 198 from the same source used for training; 82 of 1,000 rows are also selected training rows. There is no independent held-out evaluation. Exact match checks the three target fields and permits an extra confirmation flag. Catalog probes use exact registered IDs and exercise capability coverage. The six rejection cases provide narrow regression coverage. Neither establishes generalization to real speech or safe hardware operation.

The template diagnostic includes 15 incorrect triples, 29 ambiguous-device rejections, 23 unknown/unspecified-device rejections and 5 unsupported-action rejections. Sample failures and the full predictions are published.

Reproduce

Download the matching dataset beside this repository:

python -c "from huggingface_hub import snapshot_download; snapshot_download('sraivante/superfast-tiny-home-robotic-data', revision='5d997497d2cbd01f7104babd576498a9fe1f0975', local_dir='../superfast-tiny-home-robotic-data')"
python build.py ../superfast-tiny-home-robotic-data/source/devices-1M-recovered.jsonl.xz ../superfast-tiny-home-robotic-data/source/home-commands-v2-part1.jsonl.xz --out rebuilt
python evaluate.py ../superfast-tiny-home-robotic-data/source/home-commands-v2-part1.jsonl.xz --output evaluation-rerun.json
python edge_command.py --path rebuilt "turn on kitchen light"

evaluate.py evaluates the packaged model in its own directory. The last command separately exercises the rebuilt directory. Compare decompressed model JSON and logical catalog rows when checking a rebuild: gzip timestamps, dictionary ordering and SQLite versions can change physical file hashes.

Limitations

  • Synthetic templates dominate the command data. Novel wording, ASR errors, ambiguous aliases and conflicting instructions can fail or return wrong intents.
  • Negation and multiple-command rejection use limited patterns. They are not comprehensive language understanding or a security boundary.
  • Supports a single catalog-constrained command. Conditional rules, motion trajectories, destinations, numeric speed values and manipulation planning are not represented by this package.
  • The catalog contains predefined identifiers and capabilities. New devices require an explicitly configured catalog; it is not live device discovery.
  • No hardware operation, permissions system, telemetry integration or real-time control loop is provided. No comparative speed or field-reliability claim is made.

License and attribution

Apache License 2.0 applies to the user's original model, code, dataset material and selection/arrangement. Copyright (c) 2026 sraivante. Third-party ownership, licenses and runtime dependencies remain separate; the copyright statement does not claim them. No third-party model weights are included. See LICENSE and NOTICE. Dataset generator provenance and remaining historical uncertainty are documented in the linked dataset card.

Identity and Version

Repository
sraivante/superfast-tiny-home-robotic
Publisher
Sr Aivante
Task
Text classification
Modality
Text
Library
Not stated by the source
Parameters
Not stated by the source
Languages
en, hi
Revision
8623ed83f99f53f6917ffbf8b430c0aeecd6b262
First published
2026-09-26
Last updated
2026-09-26

Files and Weights

22 files, 179.0 MB in total.

Configuration9 files · 48.1 KB
Documentation4 files · 28.2 KB
Other7 files · 179.0 MB
Repository2 files · 214 B
Every file
FileTypeSizeSHA-256
build.pyConfiguration3.6 KB —
compose.yamlConfiguration361 B —
edge_command.pyConfiguration9.8 KB —
evaluate.pyConfiguration3.4 KB —
evaluation.jsonConfiguration7.2 KB —
release_evaluation.jsonConfiguration7.7 KB —
release_manifest.jsonConfiguration3.4 KB —
server.pyConfiguration2.9 KB —
training_snapshot.jsonConfiguration9.9 KB —
LICENSEDocumentation11.4 KB —
NOTICEDocumentation780 B —
ORIGINAL_README.mdDocumentation5.2 KB —
README.mdDocumentation10.8 KB —
DockerfileOther282 B —
SHA256SUMSOther1.7 KB —
action_model.json.gzOther87.2 KB cd6ee9393112
catalog.sqlite3Other167.7 MB 4dafbdda19e2
edge-command-docker-v3-ui.tar.xzOther11.2 MB 47c8881bddc3
index.htmlOther9.8 KB —
legacy_subjects.json.gzOther3.9 KB 251e170adc38
.dockerignoreRepository63 B —
.gitattributesRepository151 B —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download from Sr Aivante

Released by Sr Aivante through its official repository on Hugging Face. Read the license.

Built From

Evaluations

Each result is shown as reported, with the conditions its reporter stated. None is a SAVRN measurement. A comparison lines two results up only when their configuration, unit and setup are all stated and identical.

BenchmarkConditionsResultReported byRevisionDate
Template diagnostic (1,000 rows; overlaps training) Configuration commandsTask Catalog-constrained instruction to device intentMetric Exact activity, subject and action match (hybrid…Comparison conditions not established 92.8 sraivante
Publisher reported
Evaluated revision not stated —

Questions About superfast-tiny-home-robotic

Can I use superfast-tiny-home-robotic commercially?

Yes. superfast-tiny-home-robotic is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text classification

finbert

Prosus AI

FinBERT is a pre-trained NLP model to analyze sentiment of financial text. It is built by further training the BERT language model in the finance domain, using a large financial corpus and thereby fine-tuning it for financial sentiment classification. Financial PhraseBank by Malo et al. (2014) is used for fine-tuning. For more details, please see the paper FinBERT: Financial Sentiment Analysis with Pre-trained Language Models and our related blog post on Medium. The model will give softmax outputs for three labels: positive, negative or neutral. About Prosus Prosus is a global consumer internet group and one of the largest technology investors in the world. Operating and investing globally…

Open weights 512 tokens transformers

Model · Text classification

twitter-roberta-base-sentiment-latest

Cardiff NLP

This is a RoBERTa-base model trained on ~124M tweets from January 2018 to December 2021, and finetuned for sentiment analysis with the TweetEval benchmark. The original Twitter-based RoBERTa model can be found here and the original reference paper is TweetEval. This model is suitable for English. 0 -> Negative; 1 -> Neutral; 2 -> Positive This sentiment analysis model has been integrated into TweetNLP. You can access the demo here.

Open weights cc-by-4.0 514 tokens transformers

Model · Text classification

finbert-tone

Yi

FinBERT is a BERT model pre-trained on financial communication text. The purpose is to enhance financial NLP research and practice. It is trained on the following three financial communication corpus. The total corpora size is 4.9B tokens. More technical details on FinBERT: Click Link This released finbert-tone model is the FinBERT model fine-tuned on 10,000 manually annotated (positive, negative, neutral) sentences from analyst reports. This model achieves superior performance on financial tone analysis task. If you are simply interested in using FinBERT for financial tone analysis, give it a try. If you use the model in your academic work, please cite the following paper: Huang, Allen H.…

Open weights 512 tokens transformers

Model · Text classification

ms-marco-MiniLM-L-6-v2

Joshua

https://huggingface.co/cross-encoder/ms-marco-MiniLM-L-6-v2 with ONNX weights to be compatible with Transformers.js. If you haven't already, you can install the Transformers.js JavaScript library from NPM using: Note: Having a separate repo for ONNX weights is intended to be a temporary solution until WebML gains more traction. If you would like to make your models web-ready, we recommend converting to ONNX using Optimum and structuring your repo like this one (with ONNX weights located in a subfolder named onnx).

Open weights 512 tokens transformers.js

Model · Text classification

twitter-xlm-roberta-base-sentiment

Cardiff NLP

This is a multilingual XLM-roBERTa-base model trained on ~198M tweets and finetuned for sentiment analysis. The sentiment fine-tuning was done on 8 languages (Ar, En, Fr, De, Hi, It, Sp, Pt) but it can be used for more languages (see paper for details). This model has been integrated into the TweetNLP library.

Open weights 514 tokens transformers

Model · Text classification

emotion-english-distilroberta-base

Hartmann

With this model, you can classify emotions in English text data. The model was trained on 6 diverse datasets (see Appendix below) and predicts Ekman's 6 basic emotions, plus a neutral class: 1) anger 2) disgust 3) fear 4) joy 5) neutral 6) sadness 7) surprise The model is a fine-tuned checkpoint of DistilRoBERTa-base. For a 'non-distilled' emotion model, please refer to the model card of the RoBERTa-large version. a) Run emotion model with 3 lines of code on single text example using Hugging Face's pipeline command on Google Colab: b) Run emotion model on multiple examples and full datasets (e.g.,.csv files) on Google Colab: Please reach out to [email protected] if you have any…

Open weights 514 tokens transformers