Myosotis-1-base (100M)
Myosotis-1-base is the first flagship release from us, introducing a 100-million parameter recurrent language model built on the FWKV architecture.
Myosotis-1 is engineered to never truly forget—using a mathematically clamped exponential decay that guarantees an infinite effective context window while maintaining blazing-fast inference on consumer hardware.
Architecture at a Glance
| Component |
Specification |
| Type |
RWKV-style Gating |
| Total Parameters |
~100 Million |
Hidden Dimension (d_model) |
768 |
Embedding Bottleneck (d_emb) |
192 |
Layers (n_layers) |
13 |
| FFN Expansion Factor |
4× (GELU activation) |
| Context Length |
1024 tokens (packed training) |
| Vocabulary |
50,257 (GPT-2 tokenizer |
| Weight Tying |
Fully tied, factorized input/output head |
Core Technical Innovations
1. The FWKV Recurrent Core
Instead of pairwise attention, Myosotis uses a fixed-size state vector updated via a gated linear recurrence:
$$ S_t = S_{t-1} \odot W + k_t \odot v_t $$