"Abliterated Dolphin" is a result of my 3AM brain reading about technique called abliteration and then thinking what would happen if I tried to abliterate a model that is already relatively free, such as Dolphin.
Heavily inspired by mlabonne's article on abliteration on how to redirect refusals and effectively remove, or ablate, censorship from a language model.
There is really no deeper meaning to any of this than pure curiosity.
Model Details
GGUF quants: saukko/Abliterated-Dolphin3.0-R1-Mistral-24B-GGUF
Model Description
This is basically a dumbed down version of the original Dolphin model I used as a base, as I have not done any DPOs to heal the damage caused by abliteration.
Don't try to do anything meaningful with this model. Use the original Dolphin3.0-R1-Mistral-24B instead.
Uses
There's really no good use for this model as is really. This is basically a Dolphin that has had its brain poked at and then glued back together by some self-taught and unlicensed doctor, who got lost and found himself in a surgery room.
Bias, Risks, and Limitations
- Bias: none or very low
- Risks: a lot. please use the original model instead
- Limitations: same as original but this one is lot dumber
Training, Evaluation and Model Examination
TBD
Technical Specifications
I strongly suggest you look at directly the sources I used myself. Go see mlabonne here on hf to start with. Below are some of the many sources I dug through on my mission.
References
- https://mlabonne.github.io/blog/posts/2024-06-04_Uncensor_any_LLM_with_abliteration.html
- https://www.lesswrong.com/posts/jGuXSZgv6qfdhMCuJ/refusal-in-llms-is-mediated-by-a-single-direction
- https://github.com/FailSpy/abliterator
- https://github.com/llm-attacks/llm-attacks