Nocturne v1.2d (Teacher) — the first Nocturne model to beat v1 on the frozen test
Audio Spectrogram Transformer, ~86M parameters, 2,196 output columns. This is the first candidate in
the Nocturne line to clear our pre-committed release gate against the v1 teacher. It got there by fixing
two supervision bugs, not by adding data — which is the most useful thing in this card.
Read the limitations section before deploying it. On an unseen recording site this model is
substantially worse than v1, and our much smaller distilled student beats it outright on the very test
it was released for. Both are stated with numbers below.
Results on the frozen harness
Identical 13,710-clip held-out test set, per-class thresholds calibrated on validation, scored over the
2,182 name-aligned core classes. Rows are identical across all models by construction, so these columns
are directly comparable.
| model |
params |
macro-F1 (calibrated) |
mAP |
| v1 teacher (baseline) |
86M |
0.148766 |
0.149849 |
| v1.2d (this model) |
86M |
0.151051 |
0.154110 |
| v1.2 (weak labels) |
86M |
0.126 |
0.126 |
| nocturne-v1-mini (student) |
9.4M |
0.154198 |
0.160848 |
Deltas against v1: +0.002286 macro-F1, +0.004262 mAP. Both exceed the noninferiority margins of
0.00169 and 0.00152, which were derived from a paired bootstrap over the test rows and written into the
run's ship condition before this model was scored. That ordering is the point: the margin was not
chosen to let the candidate through.
What actually produced the gain
Three earlier runs isolate it, and the answer is not more data.
| run |
change |
calibrated macro-F1 |
| v1.2 |
added AnuraSet with weak recording-level labels |
0.126 — regression |
| v1.2b |
strong labels, medium-quality subset |
0.1476 — parity at best |
| v1.2c |
strong labels + fixed annotation parser |
0.1474 — no gain from the parser alone |
| v1.2d |
parser fix + multi-hot targets + mAP checkpoint selection |
0.1511 |
Two bugs, both in supervision rather than architecture:
- The annotation parser dropped 56% of AnuraSet annotations. The suffix on a strong-label column is
call quality (
_L/_M/_H), not sex. Our parser assumed sex and silently discarded everything it
did not recognise.
- Targets were one-hot when the task is multi-label. About 37% of strong-label rows were the same
clip under a different species. One-hot training presented those clips repeatedly with contradictory
negatives. Collapsing to multi-hot took training rows from 137,068 to 86,703 — fewer rows, better model.
Checkpoint selection mattered as much. Best epoch was 18 of 40, chosen by validation mAP. Under the
previous [email protected] rule this run would have shipped a much later and worse checkpoint.
Limitations — please read these
It is worse than v1 at a recording site it has never heard. On a held-out AnuraSet site, top-1
accuracy is 0.138 against v1's 0.544. Adding an anuran corpus did not buy site generalisation; the
v1.2 family learned site signatures. If you are deploying to a new field location with no local
validation data, use nocturne-v1-teacher instead. This diagnostic covers 2,209 rows and only 5
scorable classes, so treat it as a warning signal rather than a precise measurement — but the direction
has been consistent across every model in this family.
Our 9.4M student beats it. nocturne-v1-mini scores higher on the same frozen test at roughly a
ninth of the parameters and about ten times the speed. If you want the best numbers on this benchmark,
or anything running at the edge, take the mini. This model is published because it is the first teacher
to clear the gate and because the result behind it is worth having on the record, not because it is the
best model we have.
The vocabulary contains near-duplicate entries. 2,196 columns cover about 1,896 unique species;
some appear both as Genus species and Genus_species from differing source conventions. Scores can
split across the pair. The 2,182-class core set used for evaluation is name-aligned to handle this.
Not for bird identification. Use BirdNET or Perch. Not for legal or conservation decisions
without field verification. Coverage is biased toward temperate zones and well-recorded taxa.
Usage
from model import load_nocturne, predict_file
model, vocab, thresholds = load_nocturne(".") # strict load, transformers-version aware
print(predict_file(model, "clip.wav", vocab, thresholds, top_k=5))
model.py remaps parameter names across transformers versions and then loads strictly. These weights
were saved under transformers ≥5.16, which renamed every AST attention parameter. On an older build the
names will not match, and loading with strict=False appears to succeed while leaving the entire backbone
at its AudioSet initialisation — the model then returns confident nonsense. Verified loading cleanly on
transformers 4.57.6 and 5.16.1. If it raises, install a matching transformers rather than relaxing the check.
Files
model.safetensors (verified bit-identical to the training checkpoint before upload) · vocab.json ·
thresholds.json (per-class, calibrated on validation) · eval_report.json (the full frozen-harness
report behind the numbers above) · training_config.yaml (including the ship condition) · config.json
Try the line without installing anything
https://nocturne.runstratus.com/ serves v1, mini and v1.2. This checkpoint is not yet a route
there; v1 remains the default because of the unseen-site behaviour described above.
Licence: weights CC-BY-4.0, code Apache-2.0.
vs BirdNET on non-bird taxa (release headline)
Same protocol as v1's card: 1,000 randomly sampled non-bird clips from the iNat Sounds 2024 test split (never trained on), BirdNET v2.4 via birdnetlib (min_conf floor as its author recommends) and this model given identical inputs; metric is top-1 species accuracy. Run 2026-09-23 on a DGX Spark (CPU inference), harness soundscape/birdnet_benchmark.py built with this checkpoint's 2,196-class vocab.
| Model |
Top-1 on non-bird clips (n=1,000) |
| BirdNET (v2.4, bird-focused) |
6.7% |
| Nocturne v1.2d teacher |
73.3% |
~11× lift. v1.1 measured 76.7% vs 6.7% on a 300-clip sample; the two samples differ in size and draw, so treat 73.3 and 76.7 as the same finding, not a regression. BirdNET's vocabulary is bird-only, so its non-bird accuracy is expected to be near zero — the point of this table is that Nocturne fills that gap, not that BirdNET is bad at birds.