Open-weight model · Text generation
AMOR-GatedDeltaNet-440M
by Flyingodzilla FlyinGodzilla/AMOR-GatedDeltaNet-440M
AMOR-GatedDeltaNet-440M is an open-weight model for text generation from Flyingodzilla, released under MIT License. It has 442M parameters. At 16-bit it needs about 1.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
Runs On
What it takes to serve AMOR-GatedDeltaNet-440M (442M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 0.9 GB | 1.1 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 0.4 GB | 0.5 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 0.2 GB | 0.3 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 21, 2026.
AMOR-GatedDeltaNet-440M on every accelerator the SAVRN Index prices, at every precision
Model Card
Identity and Version
- Repository
- FlyinGodzilla/AMOR-GatedDeltaNet-440M
- Publisher
- Flyingodzilla
- Task
- Text generation
- Modality
- Text
- Library
- pytorch
- Parameters
- 442M parameters
- Languages
- en
- Revision
- df3baf82d2ffd01386a5e5d168201f3a16e57850
- First published
- 2026-09-21
- Last updated
- 2026-09-21
License and Download
- License
- mit
- Access
- Open weights, no gate
Released by Flyingodzilla through its official repository on Hugging Face. Read the license.
Built From
- Described by arXiv:2602.13215
- Trained on (disclosed) HuggingFaceFW/fineweb-edu
Memory Requirements
| Precision | Weights in memory |
|---|---|
| 16-bit | 0.9 GB |
| 8-bit | 0.4 GB |
| 4-bit | 0.2 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About AMOR-GatedDeltaNet-440M
How much GPU memory does AMOR-GatedDeltaNet-440M need?
About 1.1 GB at 16-bit and 0.3 GB at 4-bit: the weights (442M parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run AMOR-GatedDeltaNet-440M on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use AMOR-GatedDeltaNet-440M commercially?
Yes. AMOR-GatedDeltaNet-440M is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.
Similar Models
The three-language base with a 2,048-token context: AnuLM-Base-400M continued for 10,000 steps at block 2,048 with YaRN, on 46M tokens of the same Hindi / English / Python proportions it was originally trained on. Five hours on one RTX 5070 Ti. Full log and the honest reading of what it bought: docs/RESULTS.md §28, with the zero-shot measurement it is compared against Hindi and English Wikipedia are CC BY-SA, C4 is ODC-BY, and the Python slice is codeparrot-clean, de-duplicated GitHub Python with mixed licences. Not affiliated with Sarvam AI, AI4Bharat, BharatGen or the Government of India. Every other checkpoint in this project is trained at 512 tokens. This one answers what happens if you…
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and…
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and…
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language model. Compared with the state-of-the-art opensource language models, including the previous released Qwen1.5, Qwen2 has generally surpassed most opensource models and demonstrated competitiveness against proprietary models across a series of benchmarks targeting for language understanding, language generation, multilingual capability, coding, mathematics, reasoning, etc. For more details, please refer to our blog…
Part of the Standalone Models by Convergent Intelligence LLC: Research Division DistilQwen Collection — Our only BF16 series. Proof-weighted distillation from Qwen3-30B-A3B → 1.7B and 0.6B on H100. Three teacher variants (Instruct, Thinking, Coder), nine models, 2,788 combined downloads. The rest of the portfolio proves structure beats scale on CPU. This collection shows what happens when you give the methodology real hardware. This model is part of the Convergent Intelligence LLC: Research Division portfolio. All models in this portfolio are developed under the Discrepancy Calculus (DISC) framework — a measure-theoretic approach to understanding and controlling the gap between what a model…