Dataset · Text classification
superfast-tiny-home-robotic-data
by Sr Aivante sraivante/superfast-tiny-home-robotic-data
Versioned training, device-catalog and diagnostic data for Publisher: sraivante. Release date: 2026-09-26. Dataset version: v1.0.0. Instructions in English and Romanized Hindi/Hinglish map to a single structured device intent.
Dataset Card
By Sr Aivante, published under apache-2.0, revision 5d997497d2cb.
Versioned training, device-catalog and diagnostic data for SuperFast Tiny Home Robotic v1.0.0. Publisher: sraivante. Release date: 2026-09-26. Dataset version: v1.0.0.
Instructions in English and Romanized Hindi/Hinglish map to a single structured device intent. The release preserves the raw command source, the recovered model training sample, exact diagnostic inputs, all predictions and the logical device catalog. There is no separate fine-tuning dataset for this count-based model.
Configurations and counts
| Configuration / split or source | Rows | Role |
|---|---|---|
commands / train |
180,309 | Reconstructed action-model training sample; all stored model statistics match |
commands / test |
1,000 | Template diagnostic from the same command source; 82 rows overlap train |
catalog / train |
1,002,414 | Device capability reference; not instruction-training examples |
catalog_probes / test |
340 | Exact-ID capability probes used by the original evaluation |
rejection_probes / test |
6 | Original negative/ambiguous/unsupported command probes |
source/home-commands-v2-part1.jsonl.xz |
2,141,541 | Complete command part consumed by the original build procedure |
source/devices-1M-recovered.jsonl.xz |
1,001,410 | Base catalog recovered from the stored SQLite rows |
The test split is a diagnostic, not a held-out generalization benchmark.
Training and diagnostic record-content overlap is also 82.
There is no separate held-out validation set and no independently collected
real-world test set. Configurations serve different purposes and should not be
concatenated. The full command source contains the selected training and test
rows. The complete catalog also contains 1,004 legacy home devices derived
from all command-source rows, including rows outside the sampled training set.
Load
from datasets import load_dataset
repo = "sraivante/superfast-tiny-home-robotic-data"
commands = load_dataset(repo, "commands", revision="v1.0.0")
print(commands["train"][0])
catalog = load_dataset(repo, "catalog", split="train", revision="v1.0.0", streaming=True)
print(next(iter(catalog)))
probes = load_dataset(repo, "catalog_probes", split="test", revision="v1.0.0")
rejections = load_dataset(repo, "rejection_probes", split="test", revision="v1.0.0")
The compressed JSONL files also work with Python's standard gzip, lzma
and json modules. The model's runtime does not require the datasets library.
Schema
Command records (commands train and test):
{"instruction":"could you stop master bedroom led strip asap please","output":{"activity":"light","subject":"master_bedroom_led_strip","action":"OFF"},"source_row":1}
This illustrates the schema using source row 1; the sampling procedure determines which rows appear in each split.
| Field | Type | Meaning |
|---|---|---|
instruction |
string | Input sentence, preserving case, spelling and punctuation |
output |
struct | Nested target with string activity, subject, action |
source_row |
integer | One-based line number in the included raw command source |
The raw source has instruction and output; source_row was added at release
time for reproducibility. Targets are objects, not JSON-encoded strings.
The 12 action classes are ON, OFF, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK,
START, STOP, STATUS.
Catalog records contain string fields subject, kind, activity, and an
actions list of strings. They preserve the stored catalog's capabilities and
labeling, including differing legacy activity names. There are 147 stored kind
labels and 31 activity labels; these are labels, not a harmonized ontology.
Catalog probes have instruction and nested output. Rejection probes have
instruction and expected_error. Predictions under evaluation/ additionally
store prediction and exact; this directory is not a training configuration.
Origin and exact training association
The local model was built from home-commands-v2-part1.jsonl.xz and a
million-device catalog. The command source was located by its recorded name,
preserved byte-for-byte, and replayed using the original build.py:
- Iterate the entire source in file order with Python
random.Random(42). - Keep each row when the next draw is at most
180000 / 2141541. - Count the same unique unigram and bigram features per action, using the original vocabulary filters.
- Compare every field in the resulting model with
action_model.json.gz.
Every stored model statistic matched: 180,309 selected rows, 6,892 retained features, class totals, feature counts and feature totals. All 1,004 legacy identifiers also matched. The 1,000 diagnostic records were reconstructed using the original reservoir sampler with seed 198; all original non-timing metrics and recorded failure examples were reproduced.
The original run did not include a contemporaneous raw-source hash or row-index manifest. The present reconstruction and hashes provide strong reproducibility evidence but do not independently prove the historical raw file bytes were identical. No training run was replaced or silently retrained for publication.
The original devices-1M.jsonl.xz file was not found. The included recovered
source exports all 1,001,410 non-legacy device rows in stored rowid order from
the supplied SQLite artifact. It preserves subject, kind, activity, and
actions. Original raw formatting, any fields not stored in SQLite, compressed
bytes and the catalog-generation script cannot be recovered and are not claimed.
Recreate the two command split files and verify the original model statistics:
python provenance/recover_training_sample.py --out reproduced --model ../superfast-tiny-home-robotic/action_model.json.gz
This release recipe was checked to reproduce the packaged training and template
diagnostic gzip files byte-for-byte. The original model repository includes
build.py for rebuilding the classifier and catalog from the source files.
Generation and provenance
The home-command data is synthetic and templated. The included
generators/generate_dataset_v2.py expands device/location/action templates
with English and Hinglish wording, capitalization and filler variations. It is
byte-identical to the generator in
sraivante/home-commands-json-v3 at da10aca352a34cb2834a049867f6f2e9fcc30e3e,
whose card identifies the command data as synthetic and Apache-2.0 licensed.
The supplied part's original shuffle/partition step is not separately recorded;
use the included source snapshot to reproduce this model exactly.
Catalog identifiers are structured device/location labels from the user's supplied artifact, not measurements or evidence of installed real hardware. The base catalog's original generation process and source revision were not recorded. No external base-model weights or third-party text corpus are included. The only files taken from outside the current model folder are the specifically associated command source, its matching generator, and the Apache license text.
provenance/training_snapshot.json records source and model SHA-256 hashes,
source revisions where available, selection settings, split counts, class
distributions, overlap, catalog recovery and remaining uncertainty.
release_manifest.json and SHA256SUMS cover the published artifacts.
Evaluation and limitations
The hybrid interpreter reproduces 928/1,000 exact command triples, 340/340 catalog probes and 6/6 rejection cases. These are template and capability diagnostics, not blind real-world performance. All inputs and predictions are included, along with the original and release-time reports.
- ON/OFF commands dominate. Class distributions are recorded in the manifest.
- English and Romanized Hindi are represented; Devanagari is not supported.
- Template similarity and explicit train/test overlap inflate apparent generalization. Device IDs and capability labels are not learned robotics skills.
- Command training contains no general rejection class, trajectories, motion destinations, numeric control values, conditional rules or dialogue context.
- Catalog capabilities do not establish safe actions in a real environment. Deployment needs independent device-specific permissions and interlocks.
- The six rejection examples are too few to assess adversarial or safety behavior.
License and attribution
Copyright (c) 2026 sraivante for original dataset material and the original
selection/arrangement contributed by the user. Those contributions are released
under Apache License 2.0. This claim does not extend to third-party material;
existing third-party notices, ownership and licenses remain in force.
The related generator is preserved from the user's Apache-2.0 dataset release.
See LICENSE and NOTICE. Missing original catalog-source bytes are explicitly
identified above rather than represented as an original raw snapshot.
Structure
commands 181,309 rows
| Split | Rows | Size |
|---|---|---|
| train | 180,309 | 19.0 MB |
| test | 1,000 | 107.0 KB |
catalog 1,002,414 rows
| Split | Rows | Size |
|---|---|---|
| train | 1,002,414 | 95.5 MB |
catalog_probes 340 rows
| Split | Rows | Size |
|---|---|---|
| test | 340 | 33.6 KB |
rejection_probes 6 rows
| Split | Rows | Size |
|---|---|---|
| test | 6 | 473 B |
Details
- Repository
- sraivante/superfast-tiny-home-robotic-data
- Publisher
- Sr Aivante
- Task category
- Text classification
- Tags
- robotics, home-automation, text-to-json
- Size category
- 1M<n<10M
- Languages
- en, hi
- Revision
- 5d997497d2cbd01f7104babd576498a9fe1f0975
- Last updated
- 2026-09-26
Files
20 files, 34.4 MB in total.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| data/action_train.jsonl.gz | Data | 4.5 MB | 3a3abe686cb6 |
| data/catalog.jsonl.gz | Data | 3.4 MB | 61e6388b1192 |
| data/catalog_probes.jsonl.gz | Data | 2.4 KB | 21a5f13137d9 |
| data/rejection_probes.jsonl.gz | Data | 264 B | 27d54f4e695e |
| data/template_diagnostic.jsonl.gz | Data | 28.6 KB | cd552c20b35b |
| evaluation/catalog_predictions.jsonl.gz | Data | 3.9 KB | 0b6ea05fd5e2 |
| evaluation/original_evaluation.json | Data | 7.2 KB | — |
| evaluation/release_evaluation.json | Data | 7.7 KB | — |
| evaluation/template_predictions.jsonl.gz | Data | 35.2 KB | e60bc5e383f3 |
| provenance/training_snapshot.json | Data | 9.8 KB | — |
| release_manifest.json | Data | 3.3 KB | — |
| LICENSE | Documentation | 11.4 KB | — |
| NOTICE | Documentation | 780 B | — |
| README.md | Documentation | 10.2 KB | — |
| SHA256SUMS | Other | 1.8 KB | — |
| generators/generate_dataset_v2.py | Other | 29.2 KB | — |
| provenance/recover_training_sample.py | Other | 3.7 KB | — |
| source/devices-1M-recovered.jsonl.xz | Other | 936.6 KB | 127a09bcf286 |
| source/home-commands-v2-part1.jsonl.xz | Other | 25.4 MB | 6f82a7f56708 |
| .gitattributes | Repository | 151 B | — |
License and Download
- License
- apache-2.0
- Access
- No access gate
Released by Sr Aivante through its official repository on Hugging Face. Read the license.
Models Trained on This Dataset
- Trained on (disclosed)superfast-tiny-home-robotic