SAVRN
Search Contact SAVRN

Open-weight model · Text classification

pplx-decider-v1-27b-EXL3

by ramGPT ramgpt/pplx-decider-v1-27b-EXL3

pplx-decider-v1-27b-EXL3 is an open-weight model for text classification from ramGPT, released under Apache License 2.0. It has 7.7B parameters and a 262,144-token context. At 16-bit it needs about 18.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b. pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer.

Parameters7.7B
Context262,144
Weights15.3 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve pplx-decider-v1-27b-EXL3 (7.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 15.3 GB 18.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 7.7 GB 9.2 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 3.8 GB 4.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 6, 2026.

pplx-decider-v1-27b-EXL3 on every accelerator the SAVRN Index prices, at every precision

Model Card

By ramGPT, published under apache-2.0, revision 5a1406352955.

EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b. pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer. It produces a hidden state and applies the model's custom readout.safetensors decision head to obtain option probabilities. The repository therefore includes pplxdeciderexl3.py, which registers the model architecture and performs the custom decision readout. This model is not recommended for standard TabbyAPI usage. A local smoke test with TabbyAPI and ExLlamaV3 1.5.2 failed during startup with: More importantly, simply registering the architecture is not sufficient for normal /v1/chat/completions behavior: TabbyAPI's standard…

Read ramGPT's full model card

pplx-decider-v1-27b EXL3 4.00 bpw

EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b.

  • Source revision: 5117a6c7fe73b19308dc1a6b0fb529a40c2ecad4
  • ExLlamaV3: 1.5.2+cu128.torch2.10.0
  • Target bitrate: 4.00 bpw
  • Architecture: Qwen3_5Model
  • Safetensors size: 15.35 GB / 14.29 GiB

Important: this is a decision model, not a normal chat model

pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer. It produces a hidden state and applies the model's custom readout.safetensors decision head to obtain option probabilities.

The repository therefore includes pplx_decider_exl3.py, which registers the model architecture and performs the custom decision readout.

TabbyAPI compatibility

This model is not recommended for standard TabbyAPI usage.

A local smoke test with TabbyAPI and ExLlamaV3 1.5.2 failed during startup with:

AssertionError: Unknown architecture Qwen3_5Model

More importantly, simply registering the architecture is not sufficient for normal /v1/chat/completions behavior: TabbyAPI's standard generation path expects an LM head and token generation, while this model requires the custom decision readout described above. A dedicated TabbyAPI backend/endpoint adapter would be needed for correct decision-model semantics.

Use the included EXL3 decision wrapper instead of treating this model as a conventional text-generation model.

RTX 4090 benchmark

Measured locally on an NVIDIA GeForce RTX 4090 24 GB with Torch 2.10.0+cu128, ExLlamaV3 1.5.2, batch size 1, FP16 cache, and direct EXL3 decision inference.

The model was kept resident in VRAM. Each context-length row below is based on 10 sequential measured requests after warm-up. Timing includes tokenization, synchronized GPU inference, decision readout, and transfer of the result back to CPU. It excludes model load time, HTTP/server overhead, and queueing.

Input tokens Median latency Mean latency Sequential decisions/s Peak PyTorch allocated VRAM
128 135 ms 136 ms 7.34 12.67 GiB
512 278 ms 278 ms 3.59 12.89 GiB
1,024 430 ms 431 ms 2.32 12.98 GiB
2,048 725 ms 725 ms 1.38 13.08 GiB
4,096 1.400 s 1.400 s 0.71 13.35 GiB
8,192 2.777 s 2.778 s 0.36 13.89 GiB

Additional measurements:

  • Model setup/load: 3.09 s with a warm OS file cache; this is not a cold-disk benchmark.
  • Resident PyTorch allocation immediately after load: 12.57 GiB.
  • First 120-token inference after load: 1.46 s; subsequent warm requests were much faster.
  • At the end of the 8K run, nvidia-smi reported about 15.0 GiB total GPU memory use for GPU 0, including the pre-existing GPU baseline and allocator-reserved memory.

This is a classifier/decision workload, so conventional autoregressive "decode tok/s" is not the useful metric. Latency per decision and prefill/context scaling are more representative.

Validation

The EXL3 artifact passed direct local decision inference using its custom readout.

A small hand-authored sanity suite produced:

  • Routing: 9/9
  • Sentiment: 6/6
  • Entailment: 9/9
  • Arithmetic: 6/6
  • Yes/no urgency: 8/8
  • Explicit severity-level selection: 5/5
  • Total: 43/43
  • Reversing multiple-choice option order: 30/30, with the same semantic choice
  • Repeating a request after intervening requests: 5/5 identical decisions, maximum probability delta 0.0

These are basic synthetic sanity checks, not an official accuracy benchmark. No BF16-vs-EXL3 accuracy/parity study has been performed, and the long-context timing prompts use synthetic filler rather than a long-context reasoning benchmark.

Notes

The original source model is a custom decision model fine-tuned from Qwen3.8-27B. Do not infer parameter count from Hugging Face's automatic architecture display alone; the source checkpoint contains about 48.6 GiB of BF16 weights, while this 4.00 bpw EXL3 artifact contains about 15.35 GB of safetensors.

Configuration

Architecture
Qwen3_5Model
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
qwen3_5
Quantization
exl3

Identity and Version

Repository
ramgpt/pplx-decider-v1-27b-EXL3
Publisher
ramGPT
Task
Text classification
Modality
Text
Library
Not stated by the source
Parameters
7.7B parameters
Languages
Not stated by the source
Revision
5a14063529552cc9af2d233d137026f2e97442b7
First published
2026-10-02
Last updated
2026-10-02

Files and Weights

21 files, 15.4 GB in total. The weights are 5 files totalling 15.3 GB in safetensors.

Weights5 files · 15.3 GB
Configuration9 files · 910.9 KB
Tokenizer2 files · 20.0 MB
Documentation3 files · 16.4 KB
Other1 file · 9.0 KB
Repository1 file · 1.8 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00004.safetensorsWeights4.3 GB aa4118235afa
model-00002-of-00004.safetensorsWeights4.2 GB 6e7988751517
model-00003-of-00004.safetensorsWeights4.2 GB 1d72242802dd
model-00004-of-00004.safetensorsWeights2.7 GB 02a70a96cc1f
readout.safetensorsWeights2.6 MB 5b397c635e86
config.jsonConfiguration4.8 KB —
decision_config.jsonConfiguration6.5 KB —
inference.pyConfiguration3.8 KB —
model.safetensors.index.jsonConfiguration274.1 KB —
pplx_decider_exl3.pyConfiguration4.7 KB —
preprocessor_config.jsonConfiguration486 B —
processor_config.jsonConfiguration1.2 KB —
quantization_config.jsonConfiguration612.3 KB —
release-manifest.jsonConfiguration3.0 KB —
LICENSEDocumentation11.5 KB —
NOTICEDocumentation566 B —
README.mdDocumentation4.3 KB —
chat_template.jinjaOther9.0 KB —
.gitattributesRepository1.8 KB —
tokenizer.jsonTokenizer20.0 MB 6f32ce20dc35
tokenizer_config.jsonTokenizer1.2 KB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
15.3 GB
Download from ramGPT

Released by ramGPT through its official repository on Hugging Face. Read the license.

Built From

  • Derived from perplexity-ai/pplx-decider-v1-27b
  • Quantized from perplexity-ai/pplx-decider-v1-27b

Memory Requirements

PrecisionWeights in memory
As published15.3 GB
16-bit15.3 GB
8-bit7.7 GB
4-bit3.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About pplx-decider-v1-27b-EXL3

How much GPU memory does pplx-decider-v1-27b-EXL3 need?

About 18.4 GB at 16-bit and 4.6 GB at 4-bit: the weights (7.7B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run pplx-decider-v1-27b-EXL3 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use pplx-decider-v1-27b-EXL3 commercially?

Yes. pplx-decider-v1-27b-EXL3 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is pplx-decider-v1-27b-EXL3's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text classification

IdeaLens-Qwen3.5-9B

Rishanth Rajendhran

IdeaLens-Qwen3.5-9B is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o scores.jsonl…

Open weights cc-by-nc-sa-4.0 7.9B parameters 262,144 tokens transformers

Model · Text classification

IdeaLens-Qwen3.5-9B-PerItem

Rishanth Rajendhran

IdeaLens-Qwen3.5-9B-PerItem is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…

Open weights cc-by-nc-sa-4.0 7.9B parameters 262,144 tokens transformers

Model · Text classification

Jev-LCT-Qwen3-8B

CaoHaoWei

Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family. Combining Qwen3-8B's extensive foundational capabilities with Looped Calibration and adaptive early exit, it provides frontier generative reasoning capabilities with deterministic sub-100ms decision latency. - 85.0% 科学推理 + 70.0% MMLU:媲美中大型生成模型的复杂逻辑推理能力,但单次推断控制在 89.2 ms 内。 - 企业级智能体中枢:支持高风险场景的“选择性预测(Selective Prediction)”,在 80% 覆盖率下实现近乎零差错审核。 - 全量独立权重:开箱即用,支持多 GPU 分片或单张 24GB 显卡(RTX 3090 / 4090)bfloat16 全速推断。 Apache License 2.0. Full repository at GitHub.

Open weights apache-2.0 8.2B parameters 40,960 tokens transformers

Model · Text classification

gevva-e2b-multimodal

David Burhans

100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…

Open weights apache-2.0 5.1B parameters 131,072 tokens transformers

Model · Text classification

gevva-e2b

David Burhans

100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…

Open weights apache-2.0 5.1B parameters 131,072 tokens transformers

Model · Text classification

OpenJudgement-4B-Preview

Kitani

Experimental open-weights judgment model by Kitani OpenJudgement is unfinished. We're releasing this checkpoint for people to experiment with, inspect, and build on. It still needs work on judgment quality, calibration, and inference efficiency. It is not as good as Jev overall in our internal task comparisons. It does show a meaningful improvement over untouched Qwen on our recorded validation comparison: 75.4% versus 64.2% annotation agreement. That is a result on a particular evaluation set, not a claim that we beat the base model on every task. There are questions it handles well and questions it confidently gets wrong. Please judge the preview by your own examples rather than assuming…

Open weights apache-2.0 4.5B parameters 262,144 tokens transformers