"Experimental Mistral Nemo Mix" using... "Model Stock merges with Mergekit". This is a mix of some known "RP Uncensored" models. I'm not responsible of the use of this. After noticing some "errors" on my previous versions including the "RemiX" one. I decided to go "simplier". "Model Stock" cleanly mixes and filters out the noise from each model. natong19/Mistral-Nemo-Instruct-2407-abliterated # Base! natong19/Mistral-Nemo-Instruct-2407-abliterated # Base! natong19/Mistral-Nemo-Instruct-2407-abliterated # Base! natong19/Mistral-Nemo-Instruct-2407-abliterated # Base! natong19/Mistral-Nemo-Instruct-2407-abliterated # Base!.\mergekitconfig1 # Intelligence.\mergekitconfig2 # RP &…
Open weights
12.2B parameters
1,024,000 tokens
transformers
Merged checkpoint produced by the family-aware Delta-P2S experiment package.
Open weights
13B parameters
4,096 tokens
transformers
This is a modified version of google/translategemma-12b-it optimized for deployment with vLLM. No retraining was performed. Only configuration files and the chat template were modified. Model weights are identical to the original. As of 2025-01-29, vLLM does not natively support TranslateGemma's custom structured input format. See vllm-project/vllm#32446 for the upstream tracking issue. Until that is merged, this repo provides a workaround by modifying configuration files to make TranslateGemma compatible with vLLM's standard chat API. This conversion is based entirely on the work done by Infomaniak-AI/vllm-translategemma-4b-it. The same conversion approach was applied to the 12B model.…
Open weights
gemma
13.2B parameters
131,072 tokens
transformers
Model · Text generation
Qwen
Qwen1.5-MoE is a transformer-based MoE decoder-only language model pretrained on a large amount of data. For more details, please refer to our blog post and GitHub repo. Qwen1.5-MoE employs Mixture of Experts (MoE) architecture, where the models are upcycled from dense language models. For instance, Qwen1.5-MoE-A2.7B is upcycled from Qwen-1.8B. It has 14.3B parameters in total and 2.7B activated parameters during runtime, while achieving comparable performance to Qwen1.5-7B, it only requires 25% of the training resources. We also observed that the inference speed is 1.74 times that of Qwen1.5-7B. The code of Qwen1.5-MoE has been in the latest Hugging face transformers and we advise you to…
Open weights
other
14.3B parameters
8,192 tokens
transformers
Model · Text generation
NVIDIA
Gemma 4 26B IT is an open multimodal model built by Google DeepMind that handles text and image inputs, can process video as sequences of frames, and generates text output. It is designed to deliver frontier-level performance for reasoning, agentic workflows, coding, and multimodal understanding on consumer GPUs and workstations, with a 256K-token context window and support for over 140 languages. The model uses a hybrid attention mechanism that interleaves local sliding-window and full global attention, with unified Keys and Values in global layers and Proportional RoPE (p-RoPE) to support long-context performance. The NVIDIA Gemma 4 26B IT NVFP4 model is quantized with NVIDIA Model…
Open weights
apache-2.0
14.4B parameters
262,144 tokens
Model Optimizer
Model · Text generation
Qwen
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features: - Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model, ensuring optimal performance across various scenarios. - Significantly enhancement in its reasoning capabilities, surpassing previous QwQ (in thinking mode) and…
Open weights
apache-2.0
14.8B parameters
40,960 tokens
transformers