SAVRN
Search Contact SAVRN

Independent publisher

ramGPT

ramgpt

Low-cost AI solutions.

Models in Library3
Datasets in Library0
Models on Hugging Face23
Followers2

Models

GGUF quantization of SargeDev/JevQwen3.8-27B. - JevQwen3.8-27B-Q4KM.gguf — Q4KM, about 15.8 GiB, 4.92 BPW The source config declares mtpnumhiddenlayers: 1, but the published source weights contain 64 main blocks (blk.0 through blk.63) and no MTP / NextN tensors. A normal conversion therefore advertises 65 blocks and fails to load in llama.cpp because blk.64. tensors are absent. This target GGUF was converted with llama.cpp bd4f514 using --no-mtp, which produces the correct 64-block target model. Validated with llama.cpp bd4f514 on an RTX 4090. The source model was trained with thinking disabled. For the tuned behavior, use reasoning off. llama-cli -m JevQwen3.8-27B-Q4KM.gguf -ngl 999 -c…

Open weights apache-2.0

EXL3 4.0 bpw conversion of XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B. Validated locally on an NVIDIA GeForce RTX 4090. The throughput figures are short local smoke measurements, not standardized cross-system benchmarks. Tool-call validation produced a structured getweather call for Toronto. Vision validation used a generated test image containing a large red square; the model correctly returned red. The source weights were not modified. Two local metadata normalizations were required for the current ExLlamaV3 conversion path: 1. The Qwen3.5 processor metadata declared Qwen2VLImageProcessor; the local conversion copy was normalized to Qwen2VLImageProcessorFast to match the ExLlamaV3 Qwen3.5…

Open weights mit 3.4B parameters 262,144 tokens

Model · Text classification

pplx-decider-v1-27b-EXL3

ramGPT

EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b. pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer. It produces a hidden state and applies the model's custom readout.safetensors decision head to obtain option probabilities. The repository therefore includes pplxdeciderexl3.py, which registers the model architecture and performs the custom decision readout. This model is not recommended for standard TabbyAPI usage. A local smoke test with TabbyAPI and ExLlamaV3 1.5.2 failed during startup with: More importantly, simply registering the architecture is not sufficient for normal /v1/chat/completions behavior: TabbyAPI's standard…

Open weights apache-2.0 7.7B parameters 262,144 tokens