# Myosotis-1-base by FWKV Project: Open-Weight Model
Source: https://savrn.com/models/myosotis-1-base
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Myosotis-1-base (102M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 0.2 GB | 0.2 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 8-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 0.1 GB | 0.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[Myosotis-1-base on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/myosotis-1-base/gpus)

## Model Card

By FWKV Project, published under apache-2.0, revision 392bfa209da1.

### Myosotis-1-base (100M)

Myosotis-1-base is the first flagship release from us, introducing a 100-million parameter recurrent language model built on the FWKV architecture.

Myosotis-1 is engineered to never truly forget—using a mathematically clamped exponential decay that guarantees an infinite effective context window while maintaining blazing-fast inference on consumer hardware.

### Architecture at a Glance

| Component | Specification |
| --- | --- |
| Type | RWKV-style Gating |
| Total Parameters | ~100 Million |
| Hidden Dimension (d_model) | 768 |
| Embedding Bottleneck (d_emb) | 192 |
| Layers (n_layers) | 13 |
| FFN Expansion Factor | 4× (GELU activation) |
| Context Length | 1024 tokens (packed training) |
| Vocabulary | 50,257 (GPT-2 tokenizer |
| Weight Tying | Fully tied, factorized input/output head |

### Core Technical Innovations

#### 1. The FWKV Recurrent Core

Instead of pairwise attention, Myosotis uses a fixed-size state vector updated via a gated linear recurrence:

$$ S_t = S_{t-1} \odot W + k_t \odot v_t $$

[Read the full model card (789 words)](https://savrn.com/models/myosotis-1-base/card)

## Configuration

Architecture

FWKVLanguageModel

Vocabulary size

50,257

Model type

fwkv

## Identity and Version

Repository

FWKV/Myosotis-1-base

Publisher

FWKV Project

Task

Text generation

Modality

Text

Library

transformers

Parameters

102M parameters

Languages

en

Revision

392bfa209da18fd749f257740359e8e008067d7b

First published

2026-09-01

Last updated

2026-09-24

## Files and Weights

10 files, 411.0 MB in total. The weights are 1 file totalling 407.5 MB in safetensors.

Weights1 file · 407.5 MB

Configuration4 files · 10.9 KB

Tokenizer2 files · 3.6 MB

Documentation2 files · 7.3 KB

Repository1 file · 1.5 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model.safetensors | Weights | 407.5 MB | 191eb88fb1f3 |
| config.json | Configuration | 1.1 KB | — |
| configuration_fwkv.py | Configuration | 826 B | — |
| generation_config.json | Configuration | 172 B | — |
| modeling_fwkv.py | Configuration | 8.9 KB | — |
| .ipynb_checkpoints/README-checkpoint.md | Documentation | 308 B | — |
| README.md | Documentation | 7.0 KB | — |
| .gitattributes | Repository | 1.5 KB | — |
| tokenizer.json | Tokenizer | 3.6 MB | — |
| tokenizer_config.json | Tokenizer | 326 B | — |

## License and Download

License

apache-2.0

Access

Open weights, no gate

Download size

407.5 MB

[Download from FWKV Project](https://huggingface.co/FWKV/Myosotis-1-base)

Released by FWKV Project through its official repository on Hugging Face. [Read the license](https://www.apache.org/licenses/LICENSE-2.0).

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 407.5 MB |
| 16-bit | 0.2 GB |
| 8-bit | 0.1 GB |
| 4-bit | 0.1 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Myosotis-1-base

### How much GPU memory does Myosotis-1-base need?

About 0.2 GB at 16-bit and 0.1 GB at 4-bit: the weights (102M parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Myosotis-1-base on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### Can I use Myosotis-1-base commercially?

Yes. Myosotis-1-base is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

## Similar Models

Model · Text generation

### [uuu_fine_tune_gpt2](https://savrn.com/models/uuu-fine-tune-gpt2)

[David Lanz](https://savrn.com/model-publishers/davidlanz)

Fine tuning pre-trained language models for text generation. Pretrained model on Chinese language using a GPT2 for Large Language Head Model objective. transferlearning from DavidLanz/uuufinetunetaipower and fine-tuning with medical dataset for the GPT-2 architecture. You can use this model directly with a pipeline for text generation. Since the generation relies on some randomness, we

Open weights gpl 102M parameters transformers

[View model](https://savrn.com/models/uuu-fine-tune-gpt2)

Model · Text generation

### [macbert4csc-base-chinese](https://savrn.com/models/macbert4csc-base-chinese)

[Ming Xu (徐明)](https://savrn.com/model-publishers/shibing624)

macbert4csc-base-chinese evaluate SIGHAN2015 test data： 由于训练使用的数据使用了SIGHAN2015的训练集（复现paper），在SIGHAN2015的测试集上达到SOTA水平。 模型结构，魔改于softmaskedbert： 本项目开源在中文文本纠错项目：pycorrector，可支持macbert4csc模型，通过如下命令调用： 当然，你也可使用transformers调用： SIGHAN+Wang271K中文纠错数据集，数据格式： 如果需要训练macbert4csc，请参考https://github.com/shibing624/pycorrector/tree/master/pycorrector/macbert MacBERT is an improved BERT with novel MLM as correction pre-training task, which mitigates the discrepancy of pre-training and fine-tuning. Here is an example of our pre-training task. Except for the new pre-training task, we also incorporate the following techniques. Note that our MacBERT can be directly replaced with the original BERT as there is no…

Open weights apache-2.0 102M parameters 512 tokens transformers

[View model](https://savrn.com/models/macbert4csc-base-chinese)

Model · Text generation

### [SharperSwarm](https://savrn.com/models/sharperswarm)

[Convergent Intelligence](https://savrn.com/model-publishers/reaperdoesntknow)

SAGI (Swarm AGI) is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. V3.2 introduces a revolutionary Self-Assessment Layer, allowing the system to predict its own performance, identify skill gaps, and autonomously design its own learning curriculum. 1. Pre-Assessment: Predict success, identify risks, recommend strategy. 2. Execution: Generate with selected strategy. 3. Real-Time Monitoring: Catch and correct errors during generation. 4. Post-Assessment: Update skill…

Open weights apache-2.0 103M parameters 1,024 tokens transformers

[View model](https://savrn.com/models/sharperswarm)

Model · Text generation

### [tyrian-75m](https://savrn.com/models/tyrian-75m)

[Phil McCanham](https://savrn.com/model-publishers/redptam)

A 75M parameter decoder-only language model built entirely from scratch in PyTorch — no HuggingFace model classes, no nanoGPT wrapping. Every component (tokenizer, architecture, data pipeline, training loop, SFT) was written from scratch with Claude (Anthropic's AI assistant). Pretraining - ~17.8B tokens of English web text - AdamW (β₁=0.9, β₂=0.95), weight decay 0.1 SFT Fine-tuning - 100K examples from OpenHermes-2.5 - ChatML format with loss masking on user/system tokens Evaluated with log-likelihood scoring (no few-shot): Comparable to GPT-2 (117M) at 0.64× the parameter count. This is a small research model, built to learn how language models work from the ground up. It is not suitable…

Open weights mit 99M parameters transformers

[View model](https://savrn.com/models/tyrian-75m)

Model · Text generation

### [lightning-105m](https://savrn.com/models/lightning-105m)

[AobanZ](https://savrn.com/model-publishers/aobangaming)

Lightning is a small, autoregressive transformer which utilizes FlashAttention and MHA. This model is trained on a variety of books from a dataset(300 MB). This is the expanded version of Lightning-60m. Lightning utilizes FlashAttention and AdamW for performance and capability. Lightning is designed to provide quick, coherent outputs, improved with a larger size and weight. Lightning is intended to be used for research, analysis and fine-tuning, stories and other. It is not intended to be used for professional advice, real writing or any kind of heavy work as generated outputs may be incorrect. Lightning can be used directly for text generation, experimentation, and conversational…

Open weights mit 105M parameters transformers

[View model](https://savrn.com/models/lightning-105m)

Model · Text generation

### [LMLM_97M_1](https://savrn.com/models/lmlm-97m-1)

[Sr Aivante](https://savrn.com/model-publishers/sraivante)

A 97.6M-parameter language model trained from scratch, then fine-tuned for short, polite, everyday English conversation. It is a small, open research and learning model: you can read every line of its training code, run it on a laptop CPU, and see exactly where a model of this size is good and where it fails. - HellaSwag 33.6% (accnorm). That is above GPT-2 small (~30%) and below SmolLM2-135M (43.1%), which saw about 250x more training text. - The custom PyTorch code is included. The model does not use transformers; see How to use. This is QuickTalk run 8. "LMLM97M1" is its published name. 1. Learning and teaching how LLMs work. It is a complete, small, from-scratch GPT, with its training…

Open weights apache-2.0 98M parameters pytorch

[View model](https://savrn.com/models/lmlm-97m-1)

## FWKV Project

[All models and datasets](https://savrn.com/model-publishers/fwkv)

## Versions

- [392bfa209da1](https://savrn.com/models/myosotis-1-base/versions/392bfa209da1) · current 2026-09-24

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [All models under apache-2.0](https://savrn.com/models/licenses/apache-2-0)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-09-24.
- [Hugging Face record](https://huggingface.co/FWKV/Myosotis-1-base)
- [How the hub is built](https://savrn.com/model-hub/methodology)
