SAVRN
Search Contact SAVRN

Open-weight model · Text generation

AMOR-Mamba2-440M

by Flyingodzilla FlyinGodzilla/AMOR-Mamba2-440M

AMOR-Mamba2-440M is an open-weight model for text generation from Flyingodzilla, released under MIT License. It has 449M parameters. At 16-bit it needs about 1.1 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Parameters449M
Context
Weights
Licensemit
AccessOpen weights
Monthly Downloads

Runs On

What it takes to serve AMOR-Mamba2-440M (449M parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 0.9 GB 1.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 0.4 GB 0.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 0.2 GB 0.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 21, 2026.

AMOR-Mamba2-440M on every accelerator the SAVRN Index prices, at every precision

Model Card

Identity and Version

Repository
FlyinGodzilla/AMOR-Mamba2-440M
Publisher
Flyingodzilla
Task
Text generation
Modality
Text
Library
pytorch
Parameters
449M parameters
Languages
en
Revision
a366f8b9680e91e2825c5e2de3f835fead175382
First published
2026-09-21
Last updated
2026-09-21

License and Download

License
mit
Access
Open weights, no gate
Download from Flyingodzilla

Released by Flyingodzilla through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
16-bit0.9 GB
8-bit0.4 GB
4-bit0.2 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About AMOR-Mamba2-440M

How much GPU memory does AMOR-Mamba2-440M need?

About 1.1 GB at 16-bit and 0.3 GB at 4-bit: the weights (449M parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run AMOR-Mamba2-440M on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use AMOR-Mamba2-440M commercially?

Yes. AMOR-Mamba2-440M is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

Similar Models

Model · Text generation

Qwen2.5-0.5B-Instruct

Qwen

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and…

Open weights apache-2.0 494M parameters 32,768 tokens transformers

Model · Text generation

Qwen2.5-0.5B

Qwen

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized expert models in these domains. - Significant improvements in instruction following, generating long texts (over 8K tokens), understanding structured data (e.g, tables), and generating structured outputs especially JSON. More resilient to the diversity of system prompts, enhancing role-play implementation and…

Open weights apache-2.0 494M parameters 32,768 tokens transformers

Model · Text generation

Qwen2-0.5B

Qwen

Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language model. Compared with the state-of-the-art opensource language models, including the previous released Qwen1.5, Qwen2 has generally surpassed most opensource models and demonstrated competitiveness against proprietary models across a series of benchmarks targeting for language understanding, language generation, multilingual capability, coding, mathematics, reasoning, etc. For more details, please refer to our blog…

Open weights apache-2.0 494M parameters 131,072 tokens transformers

Model · Text generation

DeepReasoning_1R

Convergent Intelligence

Part of the Standalone Models by Convergent Intelligence LLC: Research Division DistilQwen Collection — Our only BF16 series. Proof-weighted distillation from Qwen3-30B-A3B → 1.7B and 0.6B on H100. Three teacher variants (Instruct, Thinking, Coder), nine models, 2,788 combined downloads. The rest of the portfolio proves structure beats scale on CPU. This collection shows what happens when you give the methodology real hardware. This model is part of the Convergent Intelligence LLC: Research Division portfolio. All models in this portfolio are developed under the Discrepancy Calculus (DISC) framework — a measure-theoretic approach to understanding and controlling the gap between what a model…

Open weights 494M parameters 32,768 tokens transformers

Model · Text generation

math-slm-qwen2.5-0.5b-v4

Mandavi Singh

A domain-specific small language model for step-by-step math problem solving, built by team03 (SLM Learners) for the Pramana SLM++ Bootcamp Round 2 submission. For an OpenAI-compatible endpoint, serve with servehf.py (stdlib + transformers only, no Ollama needed). Precision note: training ran in bf16 compute (QLoRA 4-bit NF4 base), but the merged checkpoint uploaded here is float16 (the merge step reloads the base in fp16). - Public Hugging Face datasets pulled via pulldata.py; licenses verified through the HF API on 2026-09-05 and recorded in datamanifest.md. - Held-out eval set built with buildheldouteval.py from raw ExamBench rows never used in training, with a final overlap check that…

Open weights apache-2.0 494M parameters 32,768 tokens