SAVRN
Search Contact SAVRN

SAVRN Model Hub · Models by License

Open-Weight Models Under osl-3.0

5 models in the SAVRN Model Hub released under osl-3.0, from publishers including Convergent Intelligence.

5 models.

Model · Text generation

Qemma-sft

Convergent Intelligence

Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (0.6B) at the weight level (no adapters). Design: Gemma MLP/body + Qwen attention/head, projected and aligned to Gemma’s hidden size. The model is then SFT-tuned for stepwise reasoning. Use: research, instruction following, code/help, analysis, further SFT/RLHF. Limits: may hallucinate; not for safety-critical, medical, legal, or financial decisions. Follow dataset/model licenses. ~512 warm-start steps (Alpaca-style data) 256 Additional pretraining steps on (O1-OPEN/OpenO1-SFT) 128 SFT steps with (Jackrong/gpt-oss-120b-reasoning-STEM-5K) 256 SFT steps with (O1-OPEN/OpenO1-SFT) This model is part of the Convergent…

Open weights osl-3.0 32,768 tokens transformers

Model · Text generation

Qemma-redux

Convergent Intelligence

Redux This Model underwent an additional merge between Qemma-sft and Qwen3-0.6B, in addition to adding Rope Scaling. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (0.6B) at the weight level (no adapters). Design: Gemma MLP/body + Qwen attention/head, projected and aligned to Gemma’s hidden size. The model is then SFT-tuned for stepwise reasoning. This variant uses Yarn based Rope Scaling with 1:1 Ratio from maxpositionembeddings Use: research, instruction following, code/help, analysis, further SFT/RLHF. Limits: may hallucinate; not for safety-critical, medical, legal, or financial decisions. Follow dataset/model licenses. ~512 warm-start steps (Alpaca-style…

Open weights osl-3.0 32,768 tokens transformers

Model · Text generation

Qemma-Q14B

Convergent Intelligence

My mathematical formulation to utilize space projections to "measure" the Jump between points of discontinuity found in Non-Differentialable Functions. This Model underwent an additional merge between Qemma-redux and Qwen3-14B, in addition to adding Rope Scaling. Fusion Logic was updated to aid per layer fusion and post fusion embedding alignment. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (14B) at the weight level (no adapters). This variant uses Yarn based Rope Scaling with 1: Ratio from maxpositionembeddings = 524288 Gemma-3 backbone (26 layers, hidden 1152, MLP 6912) Qwen-style attention regrouped to Gemma’s 4×256 heads. (headdim=128, hidden=5120…

Open weights osl-3.0 1B parameters 524,288 tokens transformers

Model · Text generation

Qemma-Q1.7B

Convergent Intelligence

My mathematical formulation to utilize space projections to "measure" the Jump between points of discontinuity found in Non-Differentialable Functions. This Model underwent an additional merge between Qemma-redux and Qwen3-1.7B, in addition to adding Rope Scaling. Fusion Logic was updated to aid per layer fusion and post fusion embedding alignment. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (1.7B) at the weight level (no adapters). This variant uses Yarn based Rope Scaling with 1: Ratio from maxpositionembeddings = 242144 Gemma-3 backbone (26 layers, hidden 1152, MLP 6912) Qwen-style attention regrouped to Gemma’s 4×256 heads. (headdim=128, hidden=2048…

Open weights osl-3.0 1B parameters 262,144 tokens transformers

Model · Text generation

Qemma-GEI

Convergent Intelligence

My mathematical formulation to utilize space projections to "measure" the Jump between points of discontinuity found in Non-Differentialable Functions. This Model underwent an additional merge between Qemma-redux and Qwen3-0.6B, in addition to adding Rope Scaling. Fusion Logic was updated to aid per layer fusion and post fusion embedding alignment. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (0.6B) at the weight level (no adapters). This variant uses Yarn based Rope Scaling with 1:1 Ratio from maxpositionembeddings Use: research, instruction following, code/help, analysis, further SFT/RLHF. Limits: may hallucinate; not for safety-critical, medical, legal…

Open weights osl-3.0 1B parameters 131,072 tokens transformers

Who Publishes These Models

Questions

Which osl-3.0 models are most downloaded?

By monthly downloads reported by the Hugging Face Hub: Qemma-sft (3.8k); Qemma-redux (3.8k); Qemma-Q14B (3.7k).

Other licenses

See all