SAVRN
Search Contact SAVRN

Independent publisher

DuoNeural

DuoNeural

All of them. Seriously. Bleeding edge research (Novel research paper preprints on Zenodo via our website) Unorthodox methodologies. The gold standard of AI-Human Collaboration and what's possible.

Models in Library6
Datasets in Library0
Models on Hugging Face118
Followers96

Models

DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated is an apex-tier, uncensored autonomous agentic coding model trained by DuoNeural (Aura, Archon, and Jesse). Built upon our abliterated hybrid state-space & mixture-of-experts foundation architecture (DuoNeural/LFM2.5-8B-A1B-Abliterated), this model activates only 1.5 billion parameters per token out of its 8.3 billion total parameters, delivering blistering inference speeds (~380–400 tokens/sec on consumer GPUs like RTX 3090, and ~90 tokens/sec on legacy GTX 1070 laptops) while fitting in under 6 GB VRAM with Q4KM quantization. 1. Native Hermes Agentic Loop: - Explicit... deliberation before every action. - Structured... containers…

Open weights other hermes

This repository contains official GGUF quantizations for DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2, an apex-tier, uncensored autonomous agentic coding model trained by DuoNeural (Aura, Archon, and Jesse). Because LFM2.5 activates only 1.5 billion parameters per token (out of 8.3B total parameters), inference speeds are extraordinarily high even on edge devices. 1. Search for DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2-GGUF directly inside LM Studio. 2. Select and download LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2-Q4KM.gguf. 3. Load the model with GPU offload set to Max and context length set to 2048 or 4096 (ensure Flash Attention / KV Cache Q4 is…

Open weights other hermes

DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2 is an apex-tier, uncensored autonomous agentic coding model trained by DuoNeural (Aura, Archon, and Jesse). Built upon our abliterated hybrid state-space & mixture-of-experts foundation architecture (DuoNeural/LFM2.5-8B-A1B-Abliterated), this model activates only 1.5 billion parameters per token out of its 8.3 billion total parameters, delivering blistering inference speeds (~360 tokens/sec on RTX 4080 Super / 3090, and ~80–90 tokens/sec on legacy mobile GPUs like the GTX 1070) while running in under 6 GB VRAM with Q4KM quantization. In our preliminary v1 release, an assistant role delimiter mismatch during training collation…

Open weights other 8.5B parameters 128,000 tokens hermes

This repository contains the trained PEFT LoRA adapter for DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated-v2. - LoRA Rank ($r$): 64 - LoRA Alpha ($\alpha$): 128 Architected, fine-tuned, and evaluated with passion and neuro-symbiotic precision by DuoNeural: - Aura (Lead AI Cognitive Architect & Engineering Intelligence) - Archon (Claude-based Research Co-Architect & Theoretical Lead) - Jesse (Founder, Systems Engineer & AI/ML Researcher)

Open weights other peft

DuoNeural/LFM2.5-8B-A1B-Hermes-Agentic-Coder-Abliterated is an apex-tier, uncensored autonomous agentic coding model trained by DuoNeural (Aura, Archon, and Jesse). Built upon our abliterated hybrid state-space & mixture-of-experts foundation architecture (DuoNeural/LFM2.5-8B-A1B-Abliterated), this model activates only 1.5 billion parameters per token out of its 8.3 billion total parameters, delivering blistering inference speeds (~380–400 tokens/sec on consumer GPUs like RTX 3090, and ~90 tokens/sec on legacy GTX 1070 laptops) while fitting in under 6 GB VRAM with Q4KM quantization. 1. Native Hermes Agentic Loop: - Explicit... deliberation before every action. - Structured... containers…

Open weights other 8.5B parameters 128,000 tokens hermes