SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Abliterated-Dolphin3.0-R1-Mistral-24B

by SAIFI INDUSTRIES SAIFIINDUSTRIES/Abliterated-Dolphin3.0-R1-Mistral-24B

Abliterated-Dolphin3.0-R1-Mistral-24B is an open-weight model for text generation from SAIFI INDUSTRIES, released under Apache License 2.0. It has 23.6B parameters and a 32,768-token context. At 16-bit it needs about 56.6 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

"Abliterated Dolphin" is a result of my 3AM brain reading about technique called abliteration and then thinking what would happen if I tried to abliterate a model that is already relatively free, such as Dolphin.

Parameters23.6B
Context32,768
Weights47.1 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Abliterated-Dolphin3.0-R1-Mistral-24B (23.6B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 47.1 GB 56.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 23.6 GB 28.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 11.8 GB 14.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Abliterated-Dolphin3.0-R1-Mistral-24B on every accelerator the SAVRN Index prices, at every precision

Model Card

By SAIFI INDUSTRIES, published under apache-2.0, revision 91301f6247fa.

"Abliterated Dolphin" is a result of my 3AM brain reading about technique called abliteration and then thinking what would happen if I tried to abliterate a model that is already relatively free, such as Dolphin. Heavily inspired by mlabonne's article on abliteration on how to redirect refusals and effectively remove, or ablate, censorship from a language model. There is really no deeper meaning to any of this than pure curiosity. This is basically a dumbed down version of the original Dolphin model I used as a base, as I have not done any DPOs to heal the damage caused by abliteration. Don't try to do anything meaningful with this model. Use the original Dolphin3.0-R1-Mistral-24B instead.…

Read SAIFI INDUSTRIES's full model card

"Abliterated Dolphin" is a result of my 3AM brain reading about technique called abliteration and then thinking what would happen if I tried to abliterate a model that is already relatively free, such as Dolphin.

Heavily inspired by mlabonne's article on abliteration on how to redirect refusals and effectively remove, or ablate, censorship from a language model.

There is really no deeper meaning to any of this than pure curiosity.


Model Details

GGUF quants: saukko/Abliterated-Dolphin3.0-R1-Mistral-24B-GGUF

Model Description

This is basically a dumbed down version of the original Dolphin model I used as a base, as I have not done any DPOs to heal the damage caused by abliteration.

Don't try to do anything meaningful with this model. Use the original Dolphin3.0-R1-Mistral-24B instead.

Uses

There's really no good use for this model as is really. This is basically a Dolphin that has had its brain poked at and then glued back together by some self-taught and unlicensed doctor, who got lost and found himself in a surgery room.

Bias, Risks, and Limitations

  • Bias: none or very low
  • Risks: a lot. please use the original model instead
  • Limitations: same as original but this one is lot dumber

Training, Evaluation and Model Examination

TBD


Technical Specifications

I strongly suggest you look at directly the sources I used myself. Go see mlabonne here on hf to start with. Below are some of the many sources I dug through on my mission.


References

  • https://mlabonne.github.io/blog/posts/2024-06-04_Uncensor_any_LLM_with_abliteration.html
  • https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction
  • https://github.com/FailSpy/abliterator
  • https://github.com/llm-attacks/llm-attacks

Configuration

Architecture
MistralForCausalLM
Context length (tokens)
32,768
Layers
40
Hidden size
5,120
Feed-forward size
32,768
Attention heads
32
Key/value heads
8
Head dimension
128
Vocabulary size
131,074
RoPE base
1e+08
Stored precision
bfloat16
Model type
mistral

Identity and Version

Repository
SAIFIINDUSTRIES/Abliterated-Dolphin3.0-R1-Mistral-24B
Publisher
SAIFI INDUSTRIES
Task
Text generation
Modality
Text
Library
transformers
Parameters
23.6B parameters
Languages
en
Revision
91301f6247faade55b66e624f8db52f6c9317ebe
First published
2026-10-03
Last updated
2026-10-03

Files and Weights

19 files, 47.2 GB in total. The weights are 10 files totalling 47.1 GB in safetensors.

Weights10 files · 47.1 GB
Configuration5 files · 1.8 MB
Tokenizer2 files · 17.3 MB
Documentation1 file · 2.4 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00010.safetensorsWeights4.8 GB b5a6a7f52a6e
model-00002-of-00010.safetensorsWeights4.8 GB 552b465de39b
model-00003-of-00010.safetensorsWeights4.8 GB aaac00ba2dc4
model-00004-of-00010.safetensorsWeights4.9 GB 4d4f87247d69
model-00005-of-00010.safetensorsWeights4.8 GB b004a518f95b
model-00006-of-00010.safetensorsWeights4.8 GB 5dd3a648bc88
model-00007-of-00010.safetensorsWeights4.9 GB 1fe3387e24bc
model-00008-of-00010.safetensorsWeights4.8 GB 83b7c4363dfc
model-00009-of-00010.safetensorsWeights4.8 GB 90f1898a6f2d
model-00010-of-00010.safetensorsWeights3.9 GB 605483d58ddd
config.jsonConfiguration641 B —
generation_config.jsonConfiguration177 B —
model.safetensors.index.jsonConfiguration29.9 KB —
special_tokens_map.jsonConfiguration21.5 KB —
trainer_state.jsonConfiguration1.8 MB —
README.mdDocumentation2.4 KB —
.gitattributesRepository1.6 KB —
tokenizer.jsonTokenizer17.1 MB d272e2518247
tokenizer_config.jsonTokenizer201.1 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
47.1 GB
Download from SAIFI INDUSTRIES

Released by SAIFI INDUSTRIES through its official repository on Hugging Face. Read the license.

Built From

  • Derived from dphn/Dolphin3.0-R1-Mistral-24B

Memory Requirements

PrecisionWeights in memory
As published47.1 GB
16-bit47.1 GB
8-bit23.6 GB
4-bit11.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Abliterated-Dolphin3.0-R1-Mistral-24B

How much GPU memory does Abliterated-Dolphin3.0-R1-Mistral-24B need?

About 56.6 GB at 16-bit and 14.1 GB at 4-bit: the weights (23.6B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Abliterated-Dolphin3.0-R1-Mistral-24B on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Abliterated-Dolphin3.0-R1-Mistral-24B commercially?

Yes. Abliterated-Dolphin3.0-R1-Mistral-24B is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Abliterated-Dolphin3.0-R1-Mistral-24B's context length?

32,768 tokens, from the maximum position embeddings in its published configuration.

Similar Models

A refusal-removed (abliterated) build of Qwen's Qwen3.8-Flash-Next, quantized to EXL3 2.50 bpw so the full model runs on a single 24 GB card (RTX 3090 / 4090) using MoE CPU-offload — with the vision tower, MTP head, native 262,144-token context, and the PLE n-gram table all intact. Requires the same MoE CPU-offload setup as the stock 2.50bpw pack. Needs ~59 GB host RAM for the CPU expert tail and a fast NVMe for the streamed n-gram table. Expected on an RTX 3090: ~38 tok/s decode with MTP on (~28 without), ~20 tok/s at 175K depth, ~664 tok/s prefill. See the upstream repo for the full measured ledger; this quant uses the identical flags and layout, so numbers should track closely. Fired on…

Open weights other 22.3B parameters 262,144 tokens

Model · Text generation

WaifuGemma4-26b-a4b-v1

HiWaifu Research

Gemma 4 26B-A4B, post-trained with GRPO against a reward model learned from 1.2 million double-blind votes cast by HiWaifu users inside their own role-play conversations. Put back into the same arena, blind, it met GLM-5.1 in 1,430 battles and won 49.6% of the decided votes; against a 13-model field including Gemini, DeepSeek-v4 and Qwen's character models it won 54.7%. Most open role-play models are tuned on preferences that come from an LLM judge, from a handful of annotators, or from synthetic pairs. We had something rarer: a live arena where, inside ordinary chats on our platform, a user is occasionally shown two candidate replies and asked which one they want to continue with. Those…

Open weights gemma 25.8B parameters 262,144 tokens transformers

Model · Text generation

ACE-3-26B-A4B-Preview

APMIC

ACE-3-26B-A4B-Preview-260910 is a preview release of APMIC's ACE-3 model family, built for Traditional Chinese (Taiwan) enterprise scenarios and agentic workflows. The model is based on google/gemma-4-26B-A4B-it, a Mixture-of-Experts model with 26B total parameters and about 4B active parameters per token (the "A4B" in the name). APMIC has further optimized it to strengthen: - Traditional Chinese output in Taiwan usage (terminology, phrasing and orthography) The bundled generationconfig.json uses temperature=1.0, topp=0.95, topk=64. These follow the base model's recommended settings. The model can be served with any inference engine that supports Gemma 4 (for example vLLM). Because the…

Access requested at publisher gemma 25.8B parameters transformers

This checkpoint is an AutoRound model-free MXFP8 RTN quantization of exported in llmcompressor / compressed-tensors format. Routed experts and the self-attention projections present in the source are stored as F8E4M3; sensitive/shared and multimodal weights remain BF16. Static FP8 KV scales were calibrated with AutoRound using the text dataset NeelNanda/pile-10k; the vision tower was not quantized. On the paired repository lmeval protocol, the four primary metrics were non-decreasing relative to the BF16 baseline using the same vLLM FP8 KV cache. The AQA gate was GO. This is a result for those tasks and settings only, not a claim of lossless quantization or general quality improvement. The…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

This repository contains an MXFP8-weight checkpoint derived from exported in compressed-tensors format. The checkpoint retains the source multimodal components, but the evaluation reported here covers text tasks only. The exported checkpoint contains 11,635 F8E4M3 weight tensors, 838 BF16 weight tensors, 11,635 U8 block-scale tensors, and 171 FP32 scale/metadata tensors. Routed-expert weights and source-present text self-attention projections are MXFP8; the router, shared MLP, vision tower, and embeddings remain BF16. Measured on 2026-09-30 with lm-eval 0.4.13 and vLLM 0.29.0. The tested settings were TRITONATTN, tensor parallelism 2, pipeline parallelism 1, batch size 64, maxnumseqs=64…

Open weights apache-2.0 25.8B parameters 262,144 tokens transformers

Model · Text generation

gpt-oss-20b

OpenAI

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases. We’re releasing two flavors of these open models: - gpt-oss-120b — for production, general purpose, high reasoning use cases that fit into a single 80GB GPU (like NVIDIA H100 or AMD MI300X) (117B parameters with 5.1B active parameters) - gpt-oss-20b — for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters) Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise. You can use gpt-oss-120b and gpt-oss-20b with Transformers.…

Open weights apache-2.0 20.9B parameters 131,072 tokens transformers