SAVRN
Search Contact SAVRN

Dataset · Text classification

superfast-tiny-home-robotic-data

by Sr Aivante sraivante/superfast-tiny-home-robotic-data

Versioned training, device-catalog and diagnostic data for Publisher: sraivante. Release date: 2026-09-26. Dataset version: v1.0.0. Instructions in English and Romanized Hindi/Hinglish map to a single structured device intent.

Rows1,184,069
Configurations4
Size34.4 MB
Licenseapache-2.0
AccessPublicly accessible
Monthly Downloads—

Dataset Card

By Sr Aivante, published under apache-2.0, revision 5d997497d2cb.

Versioned training, device-catalog and diagnostic data for SuperFast Tiny Home Robotic v1.0.0. Publisher: sraivante. Release date: 2026-09-26. Dataset version: v1.0.0.

Instructions in English and Romanized Hindi/Hinglish map to a single structured device intent. The release preserves the raw command source, the recovered model training sample, exact diagnostic inputs, all predictions and the logical device catalog. There is no separate fine-tuning dataset for this count-based model.

Configurations and counts

Configuration / split or source Rows Role
commands / train 180,309 Reconstructed action-model training sample; all stored model statistics match
commands / test 1,000 Template diagnostic from the same command source; 82 rows overlap train
catalog / train 1,002,414 Device capability reference; not instruction-training examples
catalog_probes / test 340 Exact-ID capability probes used by the original evaluation
rejection_probes / test 6 Original negative/ambiguous/unsupported command probes
source/home-commands-v2-part1.jsonl.xz 2,141,541 Complete command part consumed by the original build procedure
source/devices-1M-recovered.jsonl.xz 1,001,410 Base catalog recovered from the stored SQLite rows

The test split is a diagnostic, not a held-out generalization benchmark. Training and diagnostic record-content overlap is also 82. There is no separate held-out validation set and no independently collected real-world test set. Configurations serve different purposes and should not be concatenated. The full command source contains the selected training and test rows. The complete catalog also contains 1,004 legacy home devices derived from all command-source rows, including rows outside the sampled training set.

Load

from datasets import load_dataset

repo = "sraivante/superfast-tiny-home-robotic-data"
commands = load_dataset(repo, "commands", revision="v1.0.0")
print(commands["train"][0])

catalog = load_dataset(repo, "catalog", split="train", revision="v1.0.0", streaming=True)
print(next(iter(catalog)))
probes = load_dataset(repo, "catalog_probes", split="test", revision="v1.0.0")
rejections = load_dataset(repo, "rejection_probes", split="test", revision="v1.0.0")

The compressed JSONL files also work with Python's standard gzip, lzma and json modules. The model's runtime does not require the datasets library.

Schema

Command records (commands train and test):

{"instruction":"could you stop master bedroom led strip asap please","output":{"activity":"light","subject":"master_bedroom_led_strip","action":"OFF"},"source_row":1}

This illustrates the schema using source row 1; the sampling procedure determines which rows appear in each split.

Field Type Meaning
instruction string Input sentence, preserving case, spelling and punctuation
output struct Nested target with string activity, subject, action
source_row integer One-based line number in the included raw command source

The raw source has instruction and output; source_row was added at release time for reproducibility. Targets are objects, not JSON-encoded strings. The 12 action classes are ON, OFF, LOW, MEDIUM, HIGH, OPEN, CLOSE, LOCK, UNLOCK, START, STOP, STATUS.

Catalog records contain string fields subject, kind, activity, and an actions list of strings. They preserve the stored catalog's capabilities and labeling, including differing legacy activity names. There are 147 stored kind labels and 31 activity labels; these are labels, not a harmonized ontology. Catalog probes have instruction and nested output. Rejection probes have instruction and expected_error. Predictions under evaluation/ additionally store prediction and exact; this directory is not a training configuration.

Origin and exact training association

The local model was built from home-commands-v2-part1.jsonl.xz and a million-device catalog. The command source was located by its recorded name, preserved byte-for-byte, and replayed using the original build.py:

  1. Iterate the entire source in file order with Python random.Random(42).
  2. Keep each row when the next draw is at most 180000 / 2141541.
  3. Count the same unique unigram and bigram features per action, using the original vocabulary filters.
  4. Compare every field in the resulting model with action_model.json.gz.

Every stored model statistic matched: 180,309 selected rows, 6,892 retained features, class totals, feature counts and feature totals. All 1,004 legacy identifiers also matched. The 1,000 diagnostic records were reconstructed using the original reservoir sampler with seed 198; all original non-timing metrics and recorded failure examples were reproduced.

The original run did not include a contemporaneous raw-source hash or row-index manifest. The present reconstruction and hashes provide strong reproducibility evidence but do not independently prove the historical raw file bytes were identical. No training run was replaced or silently retrained for publication.

The original devices-1M.jsonl.xz file was not found. The included recovered source exports all 1,001,410 non-legacy device rows in stored rowid order from the supplied SQLite artifact. It preserves subject, kind, activity, and actions. Original raw formatting, any fields not stored in SQLite, compressed bytes and the catalog-generation script cannot be recovered and are not claimed.

Recreate the two command split files and verify the original model statistics:

python provenance/recover_training_sample.py --out reproduced --model ../superfast-tiny-home-robotic/action_model.json.gz

This release recipe was checked to reproduce the packaged training and template diagnostic gzip files byte-for-byte. The original model repository includes build.py for rebuilding the classifier and catalog from the source files.

Generation and provenance

The home-command data is synthetic and templated. The included generators/generate_dataset_v2.py expands device/location/action templates with English and Hinglish wording, capitalization and filler variations. It is byte-identical to the generator in sraivante/home-commands-json-v3 at da10aca352a34cb2834a049867f6f2e9fcc30e3e, whose card identifies the command data as synthetic and Apache-2.0 licensed. The supplied part's original shuffle/partition step is not separately recorded; use the included source snapshot to reproduce this model exactly.

Catalog identifiers are structured device/location labels from the user's supplied artifact, not measurements or evidence of installed real hardware. The base catalog's original generation process and source revision were not recorded. No external base-model weights or third-party text corpus are included. The only files taken from outside the current model folder are the specifically associated command source, its matching generator, and the Apache license text.

provenance/training_snapshot.json records source and model SHA-256 hashes, source revisions where available, selection settings, split counts, class distributions, overlap, catalog recovery and remaining uncertainty. release_manifest.json and SHA256SUMS cover the published artifacts.

Evaluation and limitations

The hybrid interpreter reproduces 928/1,000 exact command triples, 340/340 catalog probes and 6/6 rejection cases. These are template and capability diagnostics, not blind real-world performance. All inputs and predictions are included, along with the original and release-time reports.

  • ON/OFF commands dominate. Class distributions are recorded in the manifest.
  • English and Romanized Hindi are represented; Devanagari is not supported.
  • Template similarity and explicit train/test overlap inflate apparent generalization. Device IDs and capability labels are not learned robotics skills.
  • Command training contains no general rejection class, trajectories, motion destinations, numeric control values, conditional rules or dialogue context.
  • Catalog capabilities do not establish safe actions in a real environment. Deployment needs independent device-specific permissions and interlocks.
  • The six rejection examples are too few to assess adversarial or safety behavior.

License and attribution

Copyright (c) 2026 sraivante for original dataset material and the original selection/arrangement contributed by the user. Those contributions are released under Apache License 2.0. This claim does not extend to third-party material; existing third-party notices, ownership and licenses remain in force. The related generator is preserved from the user's Apache-2.0 dataset release. See LICENSE and NOTICE. Missing original catalog-source bytes are explicitly identified above rather than represented as an original raw snapshot.

Structure

commands 181,309 rows

SplitRowsSize
train180,30919.0 MB
test1,000107.0 KB
instructionstringoutputvaluesource_rowint64

catalog 1,002,414 rows

SplitRowsSize
train1,002,41495.5 MB
subjectstringkindstringactivitystringactionslist

catalog_probes 340 rows

SplitRowsSize
test34033.6 KB
instructionstringoutputvalue

rejection_probes 6 rows

SplitRowsSize
test6473 B
instructionstringexpected_errorstring

Details

Repository
sraivante/superfast-tiny-home-robotic-data
Publisher
Sr Aivante
Task category
Text classification
Tags
robotics, home-automation, text-to-json
Size category
1M<n<10M
Languages
en, hi
Revision
5d997497d2cbd01f7104babd576498a9fe1f0975
Last updated
2026-09-26

Files

20 files, 34.4 MB in total.

Data11 files · 8.0 MB
Documentation3 files · 22.3 KB
Other5 files · 26.4 MB
Repository1 file · 151 B
Every file
FileTypeSizeSHA-256
data/action_train.jsonl.gzData4.5 MB3a3abe686cb6
data/catalog.jsonl.gzData3.4 MB61e6388b1192
data/catalog_probes.jsonl.gzData2.4 KB21a5f13137d9
data/rejection_probes.jsonl.gzData264 B27d54f4e695e
data/template_diagnostic.jsonl.gzData28.6 KBcd552c20b35b
evaluation/catalog_predictions.jsonl.gzData3.9 KB0b6ea05fd5e2
evaluation/original_evaluation.jsonData7.2 KB—
evaluation/release_evaluation.jsonData7.7 KB—
evaluation/template_predictions.jsonl.gzData35.2 KBe60bc5e383f3
provenance/training_snapshot.jsonData9.8 KB—
release_manifest.jsonData3.3 KB—
LICENSEDocumentation11.4 KB—
NOTICEDocumentation780 B—
README.mdDocumentation10.2 KB—
SHA256SUMSOther1.8 KB—
generators/generate_dataset_v2.pyOther29.2 KB—
provenance/recover_training_sample.pyOther3.7 KB—
source/devices-1M-recovered.jsonl.xzOther936.6 KB127a09bcf286
source/home-commands-v2-part1.jsonl.xzOther25.4 MB6f82a7f56708
.gitattributesRepository151 B—

License and Download

License
apache-2.0
Access
No access gate
Download from Sr Aivante

Released by Sr Aivante through its official repository on Hugging Face. Read the license.

Models Trained on This Dataset