Model · Text generation
Qwen
We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode, named Qwen3-4B-Instruct-2507-FP8, featuring the following key enhancements: - Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage. - Substantial gains in long-tail knowledge coverage across multiple languages. - Markedly better alignment with user preferences in subjective and open-ended tasks, enabling more helpful responses and higher-quality text generation. - Enhanced capabilities in 256K long-context understanding. This repo contains the FP8 version of Qwen3-4B-Instruct-2507, which has the following…
Open weights
apache-2.0
4.4B parameters
262,144 tokens
transformers
Model · Text generation
Qwen
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in code generation, code reasoning and code fixing. Base on the strong Qwen2.5, we scale up the training tokens into 5.5 trillion including source code, text-code grounding, Synthetic data, etc. Qwen2.5-Coder-32B has become the current state-of-the-art open-source codeLLM, with its coding abilities matching those of GPT-4o. - A more…
Open weights
apache-2.0
7.6B parameters
32,768 tokens
transformers
Pick your build → -0f6e56) Wikipedia-Korean perplexity, lower is better. Q4KM = 5.79 baseline. English builds are tuned on English; see each repo. We measure Bonsai on the same machine with the same stock llama.cpp, and we tell you where we lose. [measured] Generation speed — POCKET wins on both CPU and GPU: [measured on a MacBook M3 Pro, 18 GB] — and on a laptop, POCKET wins every axis, including prompt processing: On a laptop GPU the arithmetic headroom that let Bonsai win prefill on an H100 is gone, so MoE sparsity wins across the board. POCKET-35B-Q2K runs on the M3 Pro's CPU at 19.5 tok/s — on an 18 GB Mac, run Q2K on CPU (-ngl 0); its 13 GB exceeds the recommended Metal budget.…
Open weights
apache-2.0
llama.cpp
Model · Text generation
Qwen
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: - Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model, ensuring optimal performance across various scenarios. - Significantly enhancement in its reasoning capabilities, surpassing previous QwQ (in thinking mode) and…
Open weights
apache-2.0
8.2B parameters
40,960 tokens
transformers
Model · Text generation
Qwen
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: - Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model, ensuring optimal performance across various scenarios. - Significantly enhancement in its reasoning capabilities, surpassing previous QwQ (in thinking mode) and…
Open weights
apache-2.0
32.8B parameters
40,960 tokens
transformers
Model · Text generation
Empero
Developed by Empero GGUF quantizations of empero-ai/Qwen3.8-4B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-4B architecture — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes. This card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the main model card. Headline results for the source model (CoT protocols, lm-evaluation-harness, identical settings base vs. student): Sizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes). Practical weight-size-based guidance at modest context — the KV cache is the dominant cost at long context and may require…
Open weights
apache-2.0
gguf
Model · Text generation
Qwen
Today, we're announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we're excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct. featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks, achieving results comparable to Claude Sonnet. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format.…
Open weights
apache-2.0
480.2B parameters
262,144 tokens
transformers
Model · Text generation
Empero
Developed by Empero GGUF quantizations of empero-ai/Qwen3.8-2B — a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-2B architecture, the smallest member of the family — for llama.cpp, Ollama, LM Studio, Jan, KoboldCpp, and other stock GGUF runtimes. This card is about choosing a file and running it. The capability writeup, full benchmark results, and best practices live on the main model card. Headline results for the source model (CoT protocols, lm-evaluation-harness, identical settings base vs. student): Sizes are exact decimal GB from the uploaded files (1 GB = 1,000,000,000 bytes). Practical weight-size-based guidance at modest context — the KV cache is the dominant…
Open weights
apache-2.0
gguf
A fast and efficient 14B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Media, Vision, Sound, Tool call, and Robotics tags. Built on a DeepSeek R1-14B architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer…
Open weights
mit
14.8B parameters
131,072 tokens
Text encoder weights from Google's T5 model
Open weights
apache-2.0
4.8B parameters
transformers
This is a modified version of google/translategemma-12b-it optimized for deployment with vLLM. No retraining was performed. Only configuration files and the chat template were modified. Model weights are identical to the original. As of 2025-01-29, vLLM does not natively support TranslateGemma's custom structured input format. See vllm-project/vllm#32446 for the upstream tracking issue. Until that is merged, this repo provides a workaround by modifying configuration files to make TranslateGemma compatible with vLLM's standard chat API. This conversion is based entirely on the work done by Infomaniak-AI/vllm-translategemma-4b-it. The same conversion approach was applied to the 12B model.…
Open weights
gemma
13.2B parameters
131,072 tokens
transformers
How do I pronounce the model's name? Watch a Youtube tutorial IDEFICS (Image-aware Decoder Enhanced à la Flamingo with Interleaved Cross-attentionS) is an open-access reproduction of Flamingo, a closed-source visual language model developed by Deepmind. Like GPT-4, the multimodal model accepts arbitrary sequences of image and text inputs and produces text outputs. IDEFICS is built solely on publicly available data and models. The model can answer questions about images, describe visual contents, create stories grounded on multiple images, or simply behave as a pure language model without visual inputs. IDEFICS is on par with the original closed-source model on various image-text benchmarks…
Open weights
other
8.9B parameters
2,048 tokens
transformers
Model · Text generation
Vxtzq
CrowdGPT's first community-distributed language model architecture. Crowd-v1 is the first official model architecture released for CrowdGPT, a community-driven distributed AI project. Unlike a conventional pretrained model release, Crowd-v1 is distributed with randomly initialized weights. The purpose of this release is to provide a common model definition and weight format that CrowdGPT clients can download and collectively train. The model is designed to be consumed by the CrowdGPT distributed training infrastructure, where individual participants contribute compute toward training a shared model. Crowd-v1 contains approximately 1 billion parameters. Grouped-Query Attention (GQA) Crowd-v1…
Open weights
mit
A fast and efficient 32B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Media, Vision, Sound, Tool call, and Robotics tags. Built on a DeepSeek R1-32B architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer…
Open weights
mit
32.8B parameters
131,072 tokens
A fast and efficient 7B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Media, Vision, Sound,Tool call, and Robotics tags. Built on a DeepSeek R1 -7B architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer…
Open weights
mit
7.6B parameters
131,072 tokens
Model · Text generation
AI Box
The MVP model was proposed in MVP: Multi-task Supervised Pre-training for Natural Language Generation by Tianyi Tang, Junyi Li, Wayne Xin Zhao and Ji-Rong Wen. The detailed information and instructions can be found https://github.com/RUCAIBox/MVP. MVP is supervised pre-trained using a mixture of labeled datasets. It follows a standard Transformer encoder-decoder architecture. MVP is specially designed for natural language generation and can be adapted to a wide range of generation tasks, including but not limited to summarization, data-to-text generation, open-ended dialogue system, story generation, question answering, question generation, task-oriented dialogue system, commonsense…
Open weights
apache-2.0
1,024 tokens
transformers
Status: training in progress. No weights are published yet — this card describes the recipe and the pilot results that motivate it. A ~1B masked-diffusion language model decoded with confidence-targeted steps, then spend a few extra passes rewriting only the tokens the model is least sure about. The point is inference cost. An autoregressive model needs one sequential forward pass per token. This one needs ~20 passes for a whole sequence, regardless of its length. Cost is K + R forward passes. One refill pass fixes any number of positions at once, because the model processes the whole sequence in parallel — that is what makes targeted repair cheaper than more denoising. Draft and refill are…
Open weights
apache-2.0
2,048 tokens
Fastino-Nemotron-3.5-Lightning-Finance is a 30B-parameter, 3B-active mixture-of-experts model specialized for financial reasoning, extraction, and research fine-tuned on LoRA with the Fastino Fine-Tuning Agent. The model targets financial document reasoning, numerical question answering over filings and tables, numeric span extraction, financial entity recognition, conversational analysis, and source-grounded financial research. The evaluation suite includes FinQA, TAT-QA, SEC-Num, FinEntity, BizFinBench, BigFinanceBench, ConvFinQA, and FiQA. The published weights are BF16 and require about 66 GB before runtime overhead. An 80 GB or larger GPU, or tensor parallelism across multiple GPUs, is…
Open weights
apache-2.0
31.6B parameters
262,144 tokens
transformers
A fast and efficient ~1.5B model optimized for CPU inference. The model was refactored with BitNet features and an updated tokenizer that includes new Routing, Tool call, and Robotics tags. Built on a redesigned DeepSeek R1 architecture with native ternary (BitNet-style) support and ready-to-run GGUF quantizations. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert model in RAG deployments, with the ONNX JiRack Java server as an alternative. - Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via CMSManhattan/JiRackPrecisionTokenizer - We are working to…
Open weights
mit
1.8B parameters
131,072 tokens
A compact Qwen3.5 0.8B repository with a practical GGUF quantization ladder for local inference. This card describes what is present in the repository. The public files do not document the fine-tuning dataset or provide evaluation results, so the Cyber label should be read as the repository variant name—not as a verified capability claim. With a recent llama.cpp build: The repository includes a BF16 mmproj file and its configuration includes vision components. That establishes that a projector artifact is present; it does not establish that the end-to-end multimodal path was validated for this release. Verify image input locally before depending on it. - Local experimentation with a small…
Open weights
apache-2.0
262,144 tokens
gguf
GGUF exports of reaperdoesntknow/Qwen3.5-2B-CyberSec for local inference with llama.cpp-compatible runtimes. The source model is associated with the Trendyol Cybersecurity Instruction Tuning Dataset. No benchmark or safety-evaluation results are published with this GGUF release. A BF16 projector file is present, and the source-model configuration includes vision components. This release does not include a documented multimodal smoke-test receipt. Verify the projector, prompt format, runtime version, and image path before claiming multimodal support. - Local qualitative evaluation of the source checkpoint. - CPU or consumer-GPU experimentation. - Comparison of BF16, Q80, and Q4KM output…
Open weights
apache-2.0
262,144 tokens
gguf
Suggest a title and description for any text. On-device titles and descriptions: a short factual title and a one- to two-sentence description for any passage of text. Swift (requirements) Then add the Title product to your target. The MLX trait is required: without it the module compiles as a stub. Get a title and a one or two sentence description for any passage of text, on device. Fine-tuned on transcript clips, but it works on any prose. The register is deliberately plain, with no emoji, no hashtags and no clickbait, and a description is meant to identify this passage rather than its topic. An MLX model directory. Load the folder, not a single file. The chat template is not incidental. A…
Open weights
other
352M parameters
32,768 tokens
mlx
Topology-Aware Knowledge Distillation from Qwen3-30B-A3B → 1.7B TopologicalQwen is a 1.7B parameter model distilled from Qwen3-30B-A3B using Topological Knowledge Distillation (TKD) — a methodology that treats the teacher's output distribution over a concatenated token stream as a bounded variation (BV) function and decomposes knowledge transfer into three channels via the Mesh Fundamental Identity: 1. Smooth distillation (AC component) — Standard KL divergence over regions where the teacher's distribution varies continuously. This is what every other KD method does and stops at. 2. Jump corrections (D^j f) — Explicit correction terms at conceptual boundaries where the teacher's…
Open weights
apache-2.0
2B parameters
40,960 tokens
transformers
With Blackhole Rope Dynamics This model builds on the original 421M TAMELM-AFMoER by introducing the Blackhole Rope (BHR) mechanism—a dynamic field-based routing system designed to stabilize, amplify, and concentrate information flow across multiple temporal scales. While the original AFMoER established efficiency in routing-based intelligence, the BHR variant explores how structured gravitational-like attractors can further enhance reasoning depth without exponential increases in computation or parameters. The Blackhole Rope is a symplectic, multiscale vortex mechanism inside AFMoER that: If AFMoER routes are like neuronal pathways, the Blackhole Rope is the myelinated tether that keeps…
Open weights
apache-2.0
transformers
Purpose: Full-scale cognitive reasoning model with self-organizing memory and generative symbolic evolution SymbioticLM-14B is a 17.8-billion-parameter symbolic–transformer hybrid that couples high-capacity neural representation with structured symbolic cognition. It supports persistent memory, entropic recall, multi-stage symbolic routing, and self-organizing knowledge structures. This is an experimental research checkpoint — the capability claims below describe architectural intent, not benchmarked results (see Limitations). This model is ideal for advanced reasoning agents, research assistants, and symbolic math/code generation systems. - Long-form symbolic theorem generation and proof…
Open weights
afl-3.0
14.8B parameters
40,960 tokens
transformers
Purpose: Long-memory symbolic reasoning + high-fidelity language generation SymbioticLM-8B is a state-of-the-art hybrid transformer model with built-in symbolic cognition. It combines an 8B Qwen-based transformer with modular symbolic processors and a persistent memory buffer. The model supports both general conversation and deep symbolic tasks such as theorem generation, logical chaining, and structured reasoning with retained memory across turns. - General symbolic reasoning and logical conversation - Code + math proof modeling - Not instruction-tuned (e.g., chat-style inputs may require prompt engineering) - Larger memory buffer may increase CPU load slightly - Symbolic inference is…
Open weights
afl-3.0
8.2B parameters
40,960 tokens
transformers
SymbioticLM is a hybrid symbolic–neural language model that integrates a frozen transformer backbone (Qwen2ForCausalLM) with a suite of symbolic cognitive modules for adaptive, interpretable reasoning. The architecture fuses neural token-level generation with symbolic introspection and reasoning: - Dynamic Thought Evolution with Helical Encoding and DNA-Inspired Memory (DTE-HDM) Enables structured long-term memory and spiral-context encoding across tokens. - Multi-Agent Symbiotic Response Mechanisms (M.A.S.R.M) Coordinates symbolic-neural agents via gated attention and adaptive response layers. - QwenExoCortex Projects contextual hidden states from the Qwen model into a symbolic fusion…
Open weights
afl-3.0
3.6B parameters
transformers
The first defense AI reasoning model on Hugging Face. Shepherd-Alpha is a tactical reasoning model fine-tuned on dual-perspective military scenario analysis using BiCell Depth Dispersal — a novel training methodology that partitions transformer layers by abstraction depth and trains them asymmetrically to separate representation encoding from task-specific reasoning. Developed by Convergent Intelligence LLC: Research Division Given a tactical scenario, Shepherd-Alpha produces structured dual-perspective analysis: - Attack reasoning — how an adversary would exploit the situation - Defense reasoning — how to counter, mitigate, and survive The model is trained to think like both attacker and…
Open weights
apache-2.0
1.7B parameters
40,960 tokens
transformers
Purpose: Lightweight, memory-augmented reasoning model for CPU and embedded inference SymbioticLM-1B is the compact version of the SymbioticAI architecture. It fuses Qwen’s rotary transformer design with a symbolic processing pipeline and a persistent episodic memory. Though smaller in parameter count, it retains the full cognitive engine: symbolic memory, dynamic thought evolution, and entropy-gated control. This model is ideal for symbolic reasoning in constrained environments — like research agents, lightweight assistants, and memory-efficient logical processing. - Procedural planning, math modeling, small-code generation - Less fluent in free-form language than larger variants…
Open weights
afl-3.0
596M parameters
40,960 tokens
transformers
SmolLM2Prover is a specialized, fine-tuned version of prithivMLmods/SmolLM2-CoT-360M. While retaining the strong conversational abilities of its base model, this version has been specifically enhanced to excel at deep thinking, logical reasoning, and higher-level mathematics, with a focus on generating step-by-step proofs and explanations (Chain-of-Thought). The model was fine-tuned using multiple rounds of Supervised Fine-Tuning (SFT) with the TRL library on a curated dataset, enhancing its ability to follow complex instructions and reason through problems. This model is intended to be used for text generation tasks that require logical reasoning or advanced conversation. The easiest way…
Open weights
apache-2.0
362M parameters
8,192 tokens
transformers
SAGI (Swarm AGI) is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. V3.2 introduces a revolutionary Self-Assessment Layer, allowing the system to predict its own performance, identify skill gaps, and autonomously design its own learning curriculum. 1. Pre-Assessment: Predict success, identify risks, recommend strategy. 2. Execution: Generate with selected strategy. 3. Real-Time Monitoring: Catch and correct errors during generation. 4. Post-Assessment: Update skill…
Open weights
apache-2.0
103M parameters
1,024 tokens
transformers
SAGI is a novel causal language model that integrates swarm intelligence dynamics with transformer architecture. The model treats cognition as a dynamic, adaptive system where multiple internal "agents" collaborate through differentiable routing, trust mechanisms, and shared memory. - Episodic + Semantic Memory: Dual memory system with trainable retrieval utility The swarm processes observations derived from token embeddings, updating its internal state S. This state conditions the transformer's attention patterns and feed-forward activations via learned projections, creating bidirectional information flow between symbolic (tokens) and subsymbolic (swarm dynamics) processing. - Educational…
Open weights
apache-2.0
53M parameters
2,048 tokens
transformers
A 2B-parameter Qwen3.5 fine-tune, part of the Opus-Distil line in the reaperdoesntknow open-weight portfolio. Text-generation / reasoning model, trained with Unsloth + Hugging Face TRL.
Open weights
apache-2.0
2.3B parameters
262,144 tokens
transformers
Benefits high quality CPU inference TQ2 on Llama.cpp and Ollama via QAT - Robotcs, Routing, Coding, Multimedia, Advanced tool calling via JiRackDeltaNetTokenizer - JiRack DeltaNet understand video and images that best for Robotics also A fast and efficient 27B model optimized for CPU inference. Built on a Qwen3.8-style DeltaNet architecture (hybrid attention + SSM), with an updated tokenizer that includes Routing, Media, Vision, Sound, Tool call, and Robotics tags. Ready-to-run GGUF quantizations, and native Ollama support with reasoning disabled by default for fast, direct responses. - JiRack is a cloud-ready model that helps save money on cloud infrastructure. It can be used as an expert…
Open weights
mit
27.3B parameters
262,144 tokens
Extended Reasoning Distillation from Qwen3-30B-A3B-Thinking → 1.7B The most downloaded model in the Convergent Intelligence portfolio. Qwen3-1.7B-Thinking-Distil captures extended deliberation patterns from the Qwen3-30B-A3B Thinking teacher — the variant that generates long-form reasoning chains before committing to an answer — and compresses them into a 1.7B student via supervised fine-tuning on the longwriter-6k dataset. The Thinking teacher produces the richest signal of the three teacher variants in the DistilQwen family (Instruct, Thinking, Coder). Where Instruct distillation captures clean instruction-following and Coder captures hierarchical decomposition, Thinking distillation…
Open weights
apache-2.0
2B parameters
40,960 tokens
transformers
An English Qwen3.5 2B checkpoint associated with the Trendyol Cybersecurity Instruction Tuning Dataset and exported in Transformers / Safetensors format. This release is intended for research and local experimentation. The repository does not currently publish benchmark or safety-evaluation results, so the model should not be treated as a validated cybersecurity authority. The configuration identifies a Qwen3.5 conditional-generation architecture with text and vision components. Use a recent Transformers release that supports this architecture. Dependency and device behavior can vary across Transformers versions. Pin a tested environment for reproducible use. - Research on small-model…
Open weights
apache-2.0
2.3B parameters
262,144 tokens
transformers
A 1.7B-parameter causal language model distilled from Qwen3-30B-A3B on 6,122 STEM chain-of-thought samples using discrepancy-informed knowledge distillation. The training objective emphasizes proof structure, detects reasoning pivot tokens through token-level divergence dynamics, smooths high-entropy student singularities before distillation, and monitors structural drift through discrepancy energy. Standard knowledge distillation treats all tokens uniformly. Even proof-weighted approaches typically apply a static multiplier over the entire derivation span. That helps, but it still misses the internal structure of reasoning: some regions are smooth procedural continuation, while others are…
Open weights
apache-2.0
2B parameters
40,960 tokens
transformers
A 1.7B model built in two stages: knowledge distillation from a 30B Coder teacher to establish a structured reasoning backbone, then supervised fine-tuning on ~54,600 logical inference problems. The Coder teacher's decomposition patterns meet formal propositional logic. The hypothesis: a model that learned STEM derivation from a Coder teacher (Stage 1) already has latent structure for sequential logic, state tracking, and compositional reasoning. Logical inference SFT (Stage 2) activates that structure explicitly — the model doesn't learn logic from scratch, it surfaces what the Coder teacher already gave it. Qwen3-1.7B distilled from Qwen3-Coder-30B-A3B-Instruct — the coding-specialized…
Open weights
apache-2.0
2B parameters
40,960 tokens
transformers
A 0.6B parameter model built in two stages: knowledge distillation from a 30B Thinking teacher to establish a structured reasoning backbone, then supervised fine-tuning on legal instruction data. 50x compression. Under 500MB quantized. Runs on a phone. The training order is the thesis: teach the model how to reason first (distillation from Thinking teacher), then teach it what to reason about (legal SFT). The Thinking teacher's extended deliberation traces transfer deeper reasoning structure than an Instruct teacher — critical when the student has only 0.6B parameters to work with. Qwen3-0.6B distilled from Qwen3-30B-A3B-Thinking-2507 — a Mixture-of-Experts model with 30B total parameters…
Open weights
apache-2.0
752M parameters
40,960 tokens
transformers
Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (0.6B) at the weight level (no adapters). Design: Gemma MLP/body + Qwen attention/head, projected and aligned to Gemma’s hidden size. The model is then SFT-tuned for stepwise reasoning. Use: research, instruction following, code/help, analysis, further SFT/RLHF. Limits: may hallucinate; not for safety-critical, medical, legal, or financial decisions. Follow dataset/model licenses. ~512 warm-start steps (Alpaca-style data) 256 Additional pretraining steps on (O1-OPEN/OpenO1-SFT) 128 SFT steps with (Jackrong/gpt-oss-120b-reasoning-STEM-5K) 256 SFT steps with (O1-OPEN/OpenO1-SFT) This model is part of the Convergent…
Open weights
osl-3.0
32,768 tokens
transformers
A 0.6B parameter model distilled from Qwen3-30B-A3B-Thinking on 6,122 STEM chain-of-thought samples. 50x parameter compression. The Thinking variant teacher produces richer extended reasoning traces than the Instruct variant, transferring deeper deliberation structure into the smallest possible student. The result: a model under 500MB quantized that produces structured STEM derivations because a 30B thinking model showed it how to reason. Two key differences from standard small-model distillation: 1. Thinking teacher, not Instruct teacher. The Qwen3-30B-A3B-Thinking variant generates extended internal reasoning before committing to an answer. Its softmax distributions are higher-entropy…
Open weights
apache-2.0
752M parameters
40,960 tokens
transformers
Redux This Model underwent an additional merge between Qemma-sft and Qwen3-0.6B, in addition to adding Rope Scaling. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (0.6B) at the weight level (no adapters). Design: Gemma MLP/body + Qwen attention/head, projected and aligned to Gemma’s hidden size. The model is then SFT-tuned for stepwise reasoning. This variant uses Yarn based Rope Scaling with 1:1 Ratio from maxpositionembeddings Use: research, instruction following, code/help, analysis, further SFT/RLHF. Limits: may hallucinate; not for safety-critical, medical, legal, or financial decisions. Follow dataset/model licenses. ~512 warm-start steps (Alpaca-style…
Open weights
osl-3.0
32,768 tokens
transformers
My mathematical formulation to utilize space projections to "measure" the Jump between points of discontinuity found in Non-Differentialable Functions. This Model underwent an additional merge between Qemma-redux and Qwen3-14B, in addition to adding Rope Scaling. Fusion Logic was updated to aid per layer fusion and post fusion embedding alignment. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (14B) at the weight level (no adapters). This variant uses Yarn based Rope Scaling with 1: Ratio from maxpositionembeddings = 524288 Gemma-3 backbone (26 layers, hidden 1152, MLP 6912) Qwen-style attention regrouped to Gemma’s 4×256 heads. (headdim=128, hidden=5120…
Open weights
osl-3.0
1B parameters
524,288 tokens
transformers
My mathematical formulation to utilize space projections to "measure" the Jump between points of discontinuity found in Non-Differentialable Functions. This Model underwent an additional merge between Qemma-redux and Qwen3-1.7B, in addition to adding Rope Scaling. Fusion Logic was updated to aid per layer fusion and post fusion embedding alignment. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (1.7B) at the weight level (no adapters). This variant uses Yarn based Rope Scaling with 1: Ratio from maxpositionembeddings = 242144 Gemma-3 backbone (26 layers, hidden 1152, MLP 6912) Qwen-style attention regrouped to Gemma’s 4×256 heads. (headdim=128, hidden=2048…
Open weights
osl-3.0
1B parameters
262,144 tokens
transformers
My mathematical formulation to utilize space projections to "measure" the Jump between points of discontinuity found in Non-Differentialable Functions. This Model underwent an additional merge between Qemma-redux and Qwen3-0.6B, in addition to adding Rope Scaling. Fusion Logic was updated to aid per layer fusion and post fusion embedding alignment. Qemma is a HuggingFace-native hybrid model that merges Gemma-3 (1B) and Qwen-3 (0.6B) at the weight level (no adapters). This variant uses Yarn based Rope Scaling with 1:1 Ratio from maxpositionembeddings Use: research, instruction following, code/help, analysis, further SFT/RLHF. Limits: may hallucinate; not for safety-critical, medical, legal…
Open weights
osl-3.0
1B parameters
131,072 tokens
transformers
A compact-but-capable ≈400M parameter causal LM that replaces dot-product attention with metric-native attention and augments sequence geometry with BlackHoleRoPE (a learnable, stable RoPE variant). Designed to train and run on modest hardware (CPU-first friendly) while staying fully compatible with Transformers. Datasets: yzhuang/Agentic-Long-Context-Understanding-QA, HuggingFaceH4/MATH-500 • Distance scores, not dot products. Heads score with L2, cosine, or diag-Mahalanobis distances. This gives direct control over geometry, often stabilizes training, and can be more sample-efficient. • BlackHoleRoPE positional encoding. • Q/K: pure unit-modulus rotation (unitary → numerically stable). •…
Open weights
2,048 tokens
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). This model is part of the Convergent Intelligence LLC: Research Division portfolio. All models in this portfolio are developed under the Discrepancy Calculus (DISC) framework — a measure-theoretic approach to understanding and controlling the gap between what a model should produce and what it actually produces. DISC treats training…
Open weights
apache-2.0
1,024 tokens
transformers
This model is a fine-tuned version of LiquidAI/LFM2.5-8B-A1B, adapted on the angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k dataset for English text-generation and reasoning-style responses. The fine-tuning run used a custom Convergent Intelligence optimizer stack, CIxOpt, designed for heterogeneous routing across parameter types. The goal of this checkpoint is to test whether a Liquid Foundation Model backbone can be adapted efficiently through targeted sparse participation rather than broad full-model modification. This is an experimental research checkpoint intended for continued evaluation, domain adaptation, and architecture/optimizer testing.…
Open weights
apache-2.0
8.5B parameters
128,000 tokens
transformers
A compact-but-capable ≈150M parameter causal LM that replaces dot-product attention with metric-native attention and augments sequence geometry with BlackHoleRoPE (a learnable, stable RoPE variant). Designed to train and run on modest hardware (CPU-first friendly) while staying fully compatible with • Distance scores, not dot products. Heads score with L2, cosine, or diag-Mahalanobis distances. This gives direct control over geometry, often stabilizes training, and can be more sample-efficient. • BlackHoleRoPE positional encoding. • Q/K: pure unit-modulus rotation (unitary → numerically stable). • V: bounded-energy gating (Penrose-inspired), optionally modulated by a discrepancy signal. •…
Open weights
apache-2.0
2,048 tokens
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). This model is part of the Convergent Intelligence LLC: Research Division portfolio. All models in this portfolio are developed under the Discrepancy Calculus (DISC) framework — a measure-theoretic approach to understanding and controlling the gap between what a model should produce and what it actually produces. DISC treats training…
Open weights
664M parameters
131,072 tokens
transformers
A geometry‑aware Transformer that mixes several attention mechanisms and routes them with a metric‑based router. MoA replaces the classic dot‑product attention with metric‑based attention and blends four distinct heads per Transformer block: A token‑wise router decides, for each token, which head(s) to use and applies feature‑gates (FiLM‑style) and router‑bias gates for up/down‑scaling. The FFN is a HyperFFN – three parallel branches (SwiGLU MLP, separable‑conv, low‑rank) combined by a branch router. LayerScale and optional DropPath keep training stable. Triangle‑inequality (TI) penalty on sampled triples to encourage true‑metric behaviour. Ball pruning – each head learns an origin \(oh\)…
Open weights
apache-2.0
1,024 tokens
transformers
This model is a fine-tuned derivative of google/gemma-3-270m, adapted using the Convergent Intelligence sparse fine-tuning setup originally tested on Liquid Foundation Models. The checkpoint was trained on reasoning-style English examples from angrygiraffe/claude-opus-4.6-4.7-reasoning-8.7k using a targeted adaptation strategy and the custom CIxOpt optimizer framework. The goal of this model is to test whether a compact Gemma 3 270M backbone can be shaped toward reasoning-style text generation through selective parameter participation rather than broad full-model modification. This is an experimental research checkpoint intended for evaluation, local testing, optimizer research, and…
Open weights
gemma
268M parameters
262,144 tokens
transformers
This is the model card of a transformers model that has been pushed on the Hub. Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. Use the code below to get started with the model. Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019). Part of the DistilQwen3 Series by Convergent Intelligence LLC: Research Division This model is part of a distillation chain built on Discrepancy Calculus — a measure-theoretic framework where the teacher's output distribution is decomposed via the Mesh Fundamental Identity into smooth (AC), jump, and Cantor components. The discrepancy operator…
Open weights
2B parameters
40,960 tokens
transformers
DualMind TKD Agentic 1.7B is a two-stage derivative of Qwen/Qwen3-1.7B. It combines topology-guided mathematical knowledge distillation with assistant-masked agentic and function-calling specialization. teacher distillation topology, gap-energy diagnostics, and phase-weighted Explore/Examine/Response supervision Stage 1 was designed to transfer mathematical reasoning behavior while placing additional learning pressure on derivation, verification, and high-discrepancy reasoning transitions. - Tool schemas, user messages, and tool-result messages were visible as context but excluded from direct loss - Mathematical replay was mixed into Stage 2 to reduce catastrophic forgetting The files in…
Open weights
other
1.7B parameters
40,960 tokens
transformers
Claude Opus 4.6 Reasoning Traces → 1.7B via DualMind SFT A 1.7B model trained on 2.5M+ tokens of Claude Opus 4.6 reasoning traces using the DualMind SFT methodology. The training data comes from Opus-4.6-Reasoning-3000x-filtered — a curated dataset of extended reasoning chains from Anthropic's most capable model, with refusals removed. This is the Opus variant of the DualMind family. Where the base DualMind model was trained on LogicInference data, this model absorbs the reasoning patterns of Claude Opus 4.6 — longer chains, more nuanced self-correction, and richer deliberative structure. The Opus teacher produces qualitatively different reasoning than synthetic logic datasets: it…
Open weights
apache-2.0
2B parameters
40,960 tokens
transformers
A 1.7B parameter dual-cognition model trained on Opus 4.6 reasoning traces. The model implements a three-phase cognitive loop — explore, examine, respond — where it reasons freely, critiques its own reasoning, then synthesizes a clean answer. This is the multi-model collision array collapsed into a single architecture. The dialectical structure that produces novel insights from architectural diversity is recreated through role-conditioned generation on shared weights. No extra parameters, no routing — same weights, different cognitive modes. DualMinded-Qwen3-1.7B is the product of a four-stage pipeline: Stage 1 — Multi-Teacher Distillation: Qwen3-30B-A3B in three variants (Instruct…
Open weights
apache-2.0
2B parameters
40,960 tokens
A 1.2B hybrid model (SSM + attention) built in two stages: knowledge distillation from a 24B MoE hybrid teacher on STEM chain-of-thought data, then supervised fine-tuning on logical inference. The first proof-weighted distillation + SFT pipeline on a non-transformer architecture. Liquid Foundation Models run at 239 tok/s on AMD CPU and fit under 1GB of RAM. This model adds structured STEM reasoning and formal logical inference to that efficiency substrate. LFM2.5-1.2B distilled from LFM2-24B-A2B — a 24B MoE hybrid (SSM + attention) with only 2B active parameters per token. Teacher and student share the LFM hybrid architecture, so the KL divergence transfers reasoning patterns between…
Open weights
apache-2.0
1.3B parameters
128,000 tokens
transformers
This model is a custom-code derivative of AxiomicLabs/GPT-X2-125M, adapted for experimental long-context causal language modeling and architecture research. The repository includes a Hugging Face Transformers-compatible GPT-X2 implementation with optional Symplectic Metric-RoPE Governor support and training utilities built around CIxOpt, a heterogeneous optimizer developed for efficient parameter routing across large projection matrices, sensitive normalization parameters, and optional governor modules. The model is intended as a research checkpoint for compact long-context generation, positional encoding experiments, optimizer testing, and continued fine-tuning. This implementation uses a…
Open weights
apache-2.0
126M parameters
32,768 tokens
transformers
Single Architecture, Dual Cognition — The Multi-Model Collision Array on Shared Weights DualMind is a 1.7B parameter model that implements dual-mental-modality reasoning — a single model with two internal voices sharing the same weights, differentiated only by role tokens: - — Unconstrained reasoning. Derivation, speculation, working through the problem freely. - — Adversarial self-response. The model reads its own explore output and critiques it. Error detection, verification, refinement. - — Clean synthesis. The final answer distilled from the internal dialogue. This is the multi-model collision array collapsed into a single architecture. The dialectical structure that produces novel…
Open weights
apache-2.0
2B parameters
40,960 tokens
transformers
This model is a fine-tuned version of reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored. It has been trained using TRL. This model was trained with SFT. This model is the DISC-refined node in the DistilQwen distillation chain. Discrepancy Calculus is a measure-theoretic framework that quantifies mismatch between integration and differentiation via the discrepancy operator: $$Df(x) = \lim{\varepsilon \downarrow 0} \frac{1}{\varepsilon} \intx^{x+\varepsilon} \frac{|f(t) - f(x)|}{|t - x|}\, dt$$ DISC refinement applies the Mesh Fundamental Identity decomposition ($f = \text{AC} + \text{jumps} + \text{Cantor}$) to the model's weight space, identifying and preserving structural boundaries that…
Open weights
2B parameters
40,960 tokens
transformers