IdeaLens-Qwen3.5-9B is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o scores.jsonl…
Open-weight model · Text classification
pplx-decider-v1-27b-EXL3
by ramGPT ramgpt/pplx-decider-v1-27b-EXL3
pplx-decider-v1-27b-EXL3 is an open-weight model for text classification from ramGPT, released under Apache License 2.0. It has 7.7B parameters and a 262,144-token context. At 16-bit it needs about 18.4 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.
EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b. pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer.
Runs On
What it takes to serve pplx-decider-v1-27b-EXL3 (7.7B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.
| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
|---|---|---|---|---|---|
| 16-bit | 15.3 GB | 18.4 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 8-bit | 7.7 GB | 9.2 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
| 4-bit | 3.8 GB | 4.6 GB | 1x MI300X (192 GB) Vultr |
$1.85 | 1x H100 $1.99 · 1x MI325X $2.00 |
Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 6, 2026.
pplx-decider-v1-27b-EXL3 on every accelerator the SAVRN Index prices, at every precision
Model Card
By ramGPT, published under apache-2.0, revision 5a1406352955.
EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b. pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer. It produces a hidden state and applies the model's custom readout.safetensors decision head to obtain option probabilities. The repository therefore includes pplxdeciderexl3.py, which registers the model architecture and performs the custom decision readout. This model is not recommended for standard TabbyAPI usage. A local smoke test with TabbyAPI and ExLlamaV3 1.5.2 failed during startup with: More importantly, simply registering the architecture is not sufficient for normal /v1/chat/completions behavior: TabbyAPI's standard…
Read ramGPT's full model card
pplx-decider-v1-27b EXL3 4.00 bpw
EXL3 4.00 bpw conversion of perplexity-ai/pplx-decider-v1-27b.
- Source revision:
5117a6c7fe73b19308dc1a6b0fb529a40c2ecad4 - ExLlamaV3:
1.5.2+cu128.torch2.10.0 - Target bitrate: 4.00 bpw
- Architecture:
Qwen3_5Model - Safetensors size: 15.35 GB / 14.29 GiB
Important: this is a decision model, not a normal chat model
pplx-decider-v1-27b does not use the normal LM-head token-generation path for its final answer. It produces a hidden state and applies the model's custom readout.safetensors decision head to obtain option probabilities.
The repository therefore includes pplx_decider_exl3.py, which registers the model architecture and performs the custom decision readout.
TabbyAPI compatibility
This model is not recommended for standard TabbyAPI usage.
A local smoke test with TabbyAPI and ExLlamaV3 1.5.2 failed during startup with:
AssertionError: Unknown architecture Qwen3_5Model
More importantly, simply registering the architecture is not sufficient for normal /v1/chat/completions behavior: TabbyAPI's standard generation path expects an LM head and token generation, while this model requires the custom decision readout described above. A dedicated TabbyAPI backend/endpoint adapter would be needed for correct decision-model semantics.
Use the included EXL3 decision wrapper instead of treating this model as a conventional text-generation model.
RTX 4090 benchmark
Measured locally on an NVIDIA GeForce RTX 4090 24 GB with Torch 2.10.0+cu128, ExLlamaV3 1.5.2, batch size 1, FP16 cache, and direct EXL3 decision inference.
The model was kept resident in VRAM. Each context-length row below is based on 10 sequential measured requests after warm-up. Timing includes tokenization, synchronized GPU inference, decision readout, and transfer of the result back to CPU. It excludes model load time, HTTP/server overhead, and queueing.
| Input tokens | Median latency | Mean latency | Sequential decisions/s | Peak PyTorch allocated VRAM |
|---|---|---|---|---|
| 128 | 135 ms | 136 ms | 7.34 | 12.67 GiB |
| 512 | 278 ms | 278 ms | 3.59 | 12.89 GiB |
| 1,024 | 430 ms | 431 ms | 2.32 | 12.98 GiB |
| 2,048 | 725 ms | 725 ms | 1.38 | 13.08 GiB |
| 4,096 | 1.400 s | 1.400 s | 0.71 | 13.35 GiB |
| 8,192 | 2.777 s | 2.778 s | 0.36 | 13.89 GiB |
Additional measurements:
- Model setup/load: 3.09 s with a warm OS file cache; this is not a cold-disk benchmark.
- Resident PyTorch allocation immediately after load: 12.57 GiB.
- First 120-token inference after load: 1.46 s; subsequent warm requests were much faster.
- At the end of the 8K run,
nvidia-smireported about 15.0 GiB total GPU memory use for GPU 0, including the pre-existing GPU baseline and allocator-reserved memory.
This is a classifier/decision workload, so conventional autoregressive "decode tok/s" is not the useful metric. Latency per decision and prefill/context scaling are more representative.
Validation
The EXL3 artifact passed direct local decision inference using its custom readout.
A small hand-authored sanity suite produced:
- Routing: 9/9
- Sentiment: 6/6
- Entailment: 9/9
- Arithmetic: 6/6
- Yes/no urgency: 8/8
- Explicit severity-level selection: 5/5
- Total: 43/43
- Reversing multiple-choice option order: 30/30, with the same semantic choice
- Repeating a request after intervening requests: 5/5 identical decisions, maximum probability delta
0.0
These are basic synthetic sanity checks, not an official accuracy benchmark. No BF16-vs-EXL3 accuracy/parity study has been performed, and the long-context timing prompts use synthetic filler rather than a long-context reasoning benchmark.
Notes
The original source model is a custom decision model fine-tuned from Qwen3.8-27B. Do not infer parameter count from Hugging Face's automatic architecture display alone; the source checkpoint contains about 48.6 GiB of BF16 weights, while this 4.00 bpw EXL3 artifact contains about 15.35 GB of safetensors.
Configuration
- Architecture
- Qwen3_5Model
- Context length (tokens)
- 262,144
- Layers
- 64
- Hidden size
- 5,120
- Feed-forward size
- 17,408
- Attention heads
- 24
- Key/value heads
- 4
- Head dimension
- 256
- Vocabulary size
- 248,320
- Model type
- qwen3_5
- Quantization
- exl3
Identity and Version
- Repository
- ramgpt/pplx-decider-v1-27b-EXL3
- Publisher
- ramGPT
- Task
- Text classification
- Modality
- Text
- Library
- Not stated by the source
- Parameters
- 7.7B parameters
- Languages
- Not stated by the source
- Revision
- 5a14063529552cc9af2d233d137026f2e97442b7
- First published
- 2026-10-02
- Last updated
- 2026-10-02
Files and Weights
21 files, 15.4 GB in total. The weights are 5 files totalling 15.3 GB in safetensors.
Every file
| File | Type | Size | SHA-256 |
|---|---|---|---|
| model-00001-of-00004.safetensors | Weights | 4.3 GB | aa4118235afa |
| model-00002-of-00004.safetensors | Weights | 4.2 GB | 6e7988751517 |
| model-00003-of-00004.safetensors | Weights | 4.2 GB | 1d72242802dd |
| model-00004-of-00004.safetensors | Weights | 2.7 GB | 02a70a96cc1f |
| readout.safetensors | Weights | 2.6 MB | 5b397c635e86 |
| config.json | Configuration | 4.8 KB | — |
| decision_config.json | Configuration | 6.5 KB | — |
| inference.py | Configuration | 3.8 KB | — |
| model.safetensors.index.json | Configuration | 274.1 KB | — |
| pplx_decider_exl3.py | Configuration | 4.7 KB | — |
| preprocessor_config.json | Configuration | 486 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| quantization_config.json | Configuration | 612.3 KB | — |
| release-manifest.json | Configuration | 3.0 KB | — |
| LICENSE | Documentation | 11.5 KB | — |
| NOTICE | Documentation | 566 B | — |
| README.md | Documentation | 4.3 KB | — |
| chat_template.jinja | Other | 9.0 KB | — |
| .gitattributes | Repository | 1.8 KB | — |
| tokenizer.json | Tokenizer | 20.0 MB | 6f32ce20dc35 |
| tokenizer_config.json | Tokenizer | 1.2 KB | — |
License and Download
- License
- apache-2.0
- Access
- Open weights, no gate
- Download size
- 15.3 GB
Released by ramGPT through its official repository on Hugging Face. Read the license.
Built From
- Derived from perplexity-ai/pplx-decider-v1-27b
- Quantized from perplexity-ai/pplx-decider-v1-27b
Memory Requirements
| Precision | Weights in memory |
|---|---|
| As published | 15.3 GB |
| 16-bit | 15.3 GB |
| 8-bit | 7.7 GB |
| 4-bit | 3.8 GB |
Weights only, from the published parameter count; the key-value cache and runtime add to this.
Questions About pplx-decider-v1-27b-EXL3
How much GPU memory does pplx-decider-v1-27b-EXL3 need?
About 18.4 GB at 16-bit and 4.6 GB at 4-bit: the weights (7.7B parameters) plus a working margin. A long context needs more.
What is the cheapest GPU to run pplx-decider-v1-27b-EXL3 on?
At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.
Can I use pplx-decider-v1-27b-EXL3 commercially?
Yes. pplx-decider-v1-27b-EXL3 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.
What is pplx-decider-v1-27b-EXL3's context length?
262,144 tokens, from the maximum position embeddings in its published configuration.
Similar Models
IdeaLens-Qwen3.5-9B-PerItem is an idea-level detector: it judges whose ideas a document contains, not who wrote its words, so a document whose ideas are a person's counts as human however much of its prose an AI wrote. It is one of the detectors released with IdeaLens and trained on the same data. The idealens package (PyPI) runs the whole pipeline: it assigns each document one of the eight formats, extracts the outline with the prompt, role vocabulary and worked examples the detectors were trained with, and scores it with this model and the thresholds in this repo. Input is JSONL with a text field per document. To score outlines you already have, use idealens score outlines.jsonl -o…
Jev-LCT-Qwen3-8B is the flagship enterprise-grade decision engine of the Jev-LCT family. Combining Qwen3-8B's extensive foundational capabilities with Looped Calibration and adaptive early exit, it provides frontier generative reasoning capabilities with deterministic sub-100ms decision latency. - 85.0% 科学推理 + 70.0% MMLU:媲美中大型生成模型的复杂逻辑推理能力,但单次推断控制在 89.2 ms 内。 - 企业级智能体中枢:支持高风险场景的“选择性预测(Selective Prediction)”,在 80% 覆盖率下实现近乎零差错审核。 - 全量独立权重:开箱即用,支持多 GPU 分片或单张 24GB 显卡(RTX 3090 / 4090)bfloat16 全速推断。 Apache License 2.0. Full repository at GitHub.
100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…
100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute. That is great for writing a story, but it is painfully slow and expensive for simple decisions: - "Did the AI make up this answer, or is it actually in the PDF?" - "Should this customer's message go to billing, shipping, or technical support?" - "Does the revenue bar chart support this financial claim?" - "Did the student get the math problem right according to the answer key?" Psychologist Daniel Kahneman described human thinking…
Experimental open-weights judgment model by Kitani OpenJudgement is unfinished. We're releasing this checkpoint for people to experiment with, inspect, and build on. It still needs work on judgment quality, calibration, and inference efficiency. It is not as good as Jev overall in our internal task comparisons. It does show a meaningful improvement over untouched Qwen on our recorded validation comparison: 75.4% versus 64.2% annotation agreement. That is a result on a particular evaluation set, not a claim that we beat the base model on every task. There are questions it handles well and questions it confidently gets wrong. Please judge the preview by your own examples rather than assuming…