Try MadStickArt online → — generate a storyboard in your browser before downloading. Adjust composition and camera controls, then export clean or annotated PNGs and scene JSON. The playground runs the released neural decoder locally and uses optional rule-based scene suggestions; Strands Decider is available in the downloadable Python pipeline.
MadStickArt is a small, trained 12,192-parameter image model for clean stick-figure storyboard scenes at 956 × 400 pixels, exactly 2.39:1. This first public release is danger-room/MadStickArt-v0.1.0, licensed under Apache 2.0. Pin revision v0.1.0 for reproducible downloads.
The model is a custom PyTorch layout-conditioned neural raster decoder. A reusable scene pipeline accepts story beats, composition/framing, depth/staging, shot size, camera angle/movement, exact actor placements, sequential numbering, and production metadata. Strands Decider 2B selects missing scene controls; a procedural layout supplies explicit geometry; the learned decoder generates sharp pixels from four coarse ink channels. Numbering and metadata are drawn with a deterministic text overlay.
Model details
| Property |
Value |
| Model family |
MadStickArt |
| Release |
v0.1.0 |
| Developer/publisher |
danger-room |
| License |
Apache-2.0 |
| Architecture |
Low-resolution convolutional residual decoder with x4 PixelShuffle |
| Learned parameters |
12,192 |
| Input |
Float32 ink masks [batch, 4, 100, 239]: actors, setting, props, motion |
| Output |
Ink coverage [batch, 1, 400, 956], converted to grayscale pixels |
| Default resolution |
956 × 400, exact 239:100 aspect ratio |
| Weight precision |
float32 |
| Primary weights |
model.safetensors |
| Compatible CLI checkpoint |
best.pt, loaded with weights_only=True |
| Decoder training |
From scratch, synthetic layout-to-image supervision |
The decoder learns a correction over a bilinear union of the input channels. Its learned convolutions run on the coarse input grid. It supports flexible input grids; this release's training and evaluation use 956 × 400. Strands Decider is an optional separate planner, not the raster model's ancestor or a bundled weight component. It answers typed choice questions rather than writing scene descriptions.
This custom package has its own loading API. Publication does not provision a hosted inference service or add Transformers/Diffusers pipeline() compatibility.
Download and run
Download the pinned model repository with the Hugging Face CLI:
hf download danger-room/MadStickArt-v0.1.0 \
--revision v0.1.0 --local-dir ./MadStickArt-v0.1.0
cd MadStickArt-v0.1.0
UV_CACHE_DIR="$PWD/.cache/uv" uv sync --locked --extra planner
uv run --no-sync madstick storyboard examples/scifi-storyboard.jsonl \
--checkpoint best.pt --device mps --planner strands --planner-device mlx \
--output generated-storyboard
Use --device cpu when Apple MPS is unavailable. The Strands MLX planner is intended for Apple silicon; --planner-device cpu uses its CPU backend, or explicitly choose --planner rules for an offline deterministic planner. On Apple silicon, first Strands use downloads its separate adapter and roughly 4.5 GB of Qwen base weights to the project's .cache/huggingface; no planner weights are bundled here.
Alternatively, install only the decoder/package dependencies with pip install . and load the Safetensors release:
from madstick.hub import load_from_pretrained
from madstick.model import predict_image
from madstick.render import conditioning
from madstick.schema import SceneSpec
model = load_from_pretrained(
"danger-room/MadStickArt-v0.1.0", revision="v0.1.0", device="auto"
)
scene = SceneSpec(
beat="Two scientists talk beside a spacecraft window.",
setting="space", shot="medium", staging="two_shot",
action="talk", details=["window"], seed=17,
)
image = predict_image(model, conditioning(scene), width=956, height=400)
image.save("scientists.png")
Keep one model and pipeline instance alive when generating many frames. Each pipeline frame saves a clean PNG, an annotated PNG, and a JSON sidecar. The clean image omits numbering/metadata text and can retain camera-motion arrows. The sidecar records all original metadata, choices, confidence, fallback status, and timings.
Scene controls and intended use
The intended use is fast, rough storyboarding, spatial blocking, shot planning, and simple sci-fi scene prototyping.
| Control |
Supported vocabulary |
| Shot |
establishing, wide, medium, close_up, over_shoulder |
| Focal point |
left, center, right |
| Staging |
single, two_shot, group, foreground_background |
| Camera angle |
eye_level, low, high |
| Camera movement |
static, pan_left, pan_right, tilt_up, tilt_down, dolly_in, dolly_out, tracking |
| Setting |
interior, street, forest, desert, space, waterfront |
| Action |
stand, walk, run, talk, point, reach, sit, confront, embrace |
| Emotion |
neutral, tense, joyful, sad, surprised |
Props support sparse silhouettes such as doors, windows/screens, tables, chairs, vehicles, briefcases/crates, lamps, boats, fences, rocks, plants, clocks, bicycles, ladders, and papers. An unsupported detail becomes a neutral box. Exact actor positions can be specified with normalized x, y (foot position), scale (height), pose, and facing controls.
Explicit supported controls are preserved. Missing fields can be selected in one Strands call. Low-confidence choices use a documented per-field rule fallback (default confidence threshold 0.45), with the model's probabilities and candidate retained in the sidecar. Dependency or inference failures raise errors; the pipeline only uses the rule planner when explicitly selected.
Coverage is bounded by the procedural layout and supported scene vocabulary. Character identity, unconstrained multi-frame continuity, novel object shapes, detailed recreation of a movie scene, photorealism, and general free-form text-to-image generation are outside the first release's trained scope. Camera angles are simple layout cues; lens values are metadata, not optical simulation.
Training data and procedure
The included catalog contains 66 original generalized sci-fi beats associated with 12 films and 10 television series. Titles include Arrival, Blade Runner, Dune, Alien, Interstellar, Star Wars, Star Trek, The Expanse, Battlestar Galactica, and Foundation. The titles describe inspiration for the beats; they are not sources of downloaded visual training material.
The scene pipeline configured the catalog with the actual Strands model and rendered 2,048 synthetic pairs through controlled layout and pose augmentation. Targets are newly rendered antialiased stick figures. No movie frames, copied scripts, or third-party storyboard images are included. The generated dataset itself is omitted; the code and catalog reproduce the generation workflow. Exact MPS retraining is not guaranteed bit-identical.
| Split |
Scenes |
Source titles |
| Train |
1,583 |
17 titles |
| Validation |
186 |
Firefly; The Terminator |
| Test |
279 |
Star Trek; Star Wars; The Fifth Element |
All variants of a source title remain in one split. Augmentation decisions are recorded separately from the initial planner choices and preserve explicit controls. Title holdouts measure generalization across synthetic layouts, not understanding of unseen real cinematography.
- Dataset manifest SHA256:
861c9852e0dc1b5db87e3bc17184e9c0df6efbbb31a1cd9efb3e6ad4a0c97538.
- Training seed: 42; selected epoch: 12 of 12.
- Batch size: 8; AdamW learning rate: 0.002; weight decay: 0.00001.
- Loss: foreground-weighted coverage, edge differences, and soft Dice.
- Checkpoint selection: validation loss; test evaluated after selection.
- Training device: Apple M5 Max Metal GPU via PyTorch MPS.
- Recorded training/evaluation run: 33.60 seconds, excluding data loading and initial baseline evaluation.
Reproduce into a new directory:
uv sync --locked --extra planner --extra dev
uv run --no-sync madstick dataset --catalog examples/beats.jsonl \
--count 2048 --seed 42 --planner strands --planner-device mlx \
--output artifacts/dataset
uv run --no-sync madstick train --dataset artifacts/dataset \
--output artifacts/model-retrained --epochs 12 --batch-size 8 --device mps
uv run --no-sync pytest -q
Evaluation
Results on 279 test scenes, selected independently of model checkpoint selection:
| Metric |
Trained decoder |
Bilinear union baseline |
| Foreground IoU |
0.8362 |
0.2802 |
| Foreground precision |
0.8639 |
0.5312 |
| Foreground recall |
0.9631 |
0.3722 |
| Pixel MAE |
0.00602 |
0.02806 |
Foreground metrics use an ink threshold of 0.35. Sparse blank backgrounds make unweighted MAE misleading; the reports include a blank-image baseline and foreground-specific metrics. These numbers compare the decoder to its synthetic teacher, not a human assessment of narrative quality.
Measured performance
Recorded on Apple M5 Max, 40 GPU cores, 128 GB unified memory, batch size 1, 956 × 400, PyTorch MPS for the image decoder and MLX for Strands. Both models were warmed up.
| Scope |
Images/second |
Median |
p95 |
| Fresh Strands decisions → clean + annotated PNGs + JSON |
8.69 |
114.6 ms |
120.6 ms |
| Configured scene → new layout → clean + annotated PNGs + JSON |
97.75 |
10.29 ms |
11.30 ms |
| Neural forward only, device-resident layout |
1,598.13 |
0.443 ms |
1.41 ms |
The fresh run performed 100 uncached model calls and 700 scene-choice questions. It includes layout construction, decoder inference, CPU transfer, annotations, and image/JSON writes. Downloads and cold model/framework startup are excluded. These are workload-specific measurements, not a guarantee for base M5 hardware, arbitrary prompts, or larger images.
See training report, benchmark report, and dataset summary. Report paths have been made repository-relative.
Versioning and provenance
The repository name preserves the model family MadStickArt with a version suffix -v0.1.0. A matching Hub Git tag v0.1.0 pins this release. Future incompatible changes increment the major version, compatible capabilities increment the minor version, and corrections increment the patch version. This is the project's release policy using Hub revisions and familiar semantic versioning.
The package includes the exact implementation, examples, locked dependencies, hashes, and the following separately downloaded planner pins:
- Strands code:
75c9fd32e664954cdc18481434018aa507eee8fb.
- Adapter:
StrandsAgents/strands-decider-2B-hobson-v19, revision bb282d786bc251fd4e3068de3ada9ddbb38127cd.
- Base:
Qwen/Qwen3.5-2B-Base, revision b1485b2fa6dfa1287294f269f5fb618e03d52d7c.
- The upstream adapter export describes the base revision as inferred from training time; the pipeline records that note and uses the pinned revision.
Dependencies retain their respective licenses. See LICENSE and NOTICE. Development skills from huggingface/skills are installed only in the source project's .agents/skills; they are not bundled into this model repository.