SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4

by Jari Katajisto kataguru/Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4

Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 is an open-weight model for text generation from Jari Katajisto, released under Apache License 2.0. It has 36B parameters and a 262,144-token context. At 16-bit it needs about 86.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index.

Kataguru Sceptic Quality Inspector v2.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and ultra-high-throughput expert model.

Parameters36B
Context262,144
Weights25.6 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 (36B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 71.9 GB 86.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 36.0 GB 43.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 18.0 GB 21.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 on every accelerator the SAVRN Index prices, at every precision

Model Card

By Jari Katajisto, published under apache-2.0, revision cfdddce11787.

Kataguru Sceptic Quality Inspector v2.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and ultra-high-throughput expert model. Built upon the fleet-record foundation of Kataguru Sceptic Multi-Mode v2 CP1200 (72.01% multi-domain record, FAR = 0.000%) merged with the 16.4k certified forensic quality inspection task vector, it is purpose-engineered to serve as the fleet's primary data firewall and filtration engine. Hardware-optimized for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token…

Read Jari Katajisto's full model card

Executive Overview (English)

Kataguru Sceptic Quality Inspector v2.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and ultra-high-throughput expert model. Built upon the fleet-record foundation of Kataguru Sceptic Multi-Mode v2 CP1200 (72.01% multi-domain record, FAR = 0.000%) merged with the 16.4k certified forensic quality inspection task vector, it is purpose-engineered to serve as the fleet's primary data firewall and filtration engine.

Hardware-optimized for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token (TTFT) while consuming only ~12.1 GiB VRAM per GPU on dual RTX 5090 Blackwell hardware (TP=2).

Key Capabilities & Architecture

  1. Master Tri-Mode Operation: - Mode 1: Quality Inspector ([TILA: TARKISTUS] / [INSPECT]): Executes a deterministic, non-compensatory DETECT → VERIFY → VERDICT forensic audit on candidate training samples, returning a clean, schema-validated JSON verdict. Factual hallucinations, mathematical mistakes, or epistemic traps trigger an immediate non-negotiable REJECT. - Mode 2: Data Curator ([TILA: KOULUTUS] / [CURATE]): Synthesizes high-difficulty SFT/QLoRA multi-turn training samples using rigorous first-principles Chain-of-Thought (CoT) reasoning with zero refusals or pleasantries. - Mode 3: Sceptic Chat (Default / [TILA: KESKUSTELU]): Autonomous expert dialogue adhering to strict epistemic honesty, anti-sycophancy, and adaptive reasoning depths (spoon, einstein, deep, adaptive).
  2. Native Multimodal Vision & Multi-Token Prediction (MTP): - Full Qwen3-VL ViT attention integration (image processing in < 1.0s). - 3-token speculative MTP decoding (eagle_head, 785 tensors) delivering empirical speedup (+1.32 bonus tokens per decoding step). - Vision and MTP heads are 100% preserved in full BF16 precision.
  3. Empirical Zero Tolerance for Epistemic Traps (FAR < 0.1%): - Calibrated against adversarial hard negatives with an empirical 0.000% False Accept Rate.
  4. SOMA/ARA Zero Refusal Pipeline: - 100% free of corporate disclaimers, moralizing lectures, or AI excuses across advanced STEM, systems engineering, cybersecurity, and creative adult literature.
  5. Native Bilingual Mastery (English & Finnish): - Native-level STEM reasoning, formal logic, and systems programming in English, combined with sovereign-grade Finnish syntax, morphology, case governance (rektiot), and translationese artifact detection.

Mitä Kataguru muutti? (Anti-Hype Policy)

Kenttä Arvo / Spesifikaatio
Base model kataguru/Kataguru-Sceptic-MultiMode-v2-CP1200-BF16 (72.01 % kaikkien aikojen MoE-laivastoennätys)
Katagurun muutos Quality Inspector v2.0 -rikostekninen karsintamoottori, Master Tri-Mode -ohjausarkkitehtuuri, Blackwell NVFP4 -kvantisointi (sm_120 FP4 weights / FP8 activation scales), Marlin/FlashInfer MoE -integraatio.
Mitä EI muutettu Sparse MoE -reititysrunko, visuaalinen ViT-enkooderi (BF16), 3-tokenin MTP-spekulaatiopää (BF16), episteeminen FAR = 0.000 % -koulutuspohja.
Kielialueet Kaksikielinen (FI / EN): 100 % luonnollinen suomi ilman käännösanglismeja, sujuva teknis-tieteellinen englanti.
VRAM-tarve ~12.1 GiB per GPU (yhteensä ~24.2 GiB), jättäen 2x RTX 5090 -klusterissa yli 18 GiB vapaata tilaa laajalle kontekstille.
Ensisijainen käyttö Ultranopea aineistojen esikarsinta ja auditointi, synteettisen datan tuottaminen, reaaliaikainen kognitiivinen tuomarointi ja kriittinen asiantuntijakeskustelu.
Testattu laitteisto 2x NVIDIA GeForce RTX 5090 32GB Blackwell (TP=2), vLLM CUDAGraphs päällä.

1. Master Tri-Mode -arkkitehtuuri

Malli tukee kolmea toisistaan erotettua ja erikoistunutta toimintatilaa samassa mallipainossa:

                      ┌────────────────────────────────────────────────────────┐
                      │          KATAGURU SCEPTIC MASTER TRI-MODE              │
                      └──────────────────────────┬─────────────────────────────┘
                                                 │
         ┌───────────────────────────────────────┼───────────────────────────────────────┐
         ▼                                       ▼                                       ▼
  [TILA: TARKISTUS]                       [TILA: KOULUTUS]                        [TILA: KESKUSTELU]
  (Quality Inspector)                     (Data Curator)                          (Adaptive Sceptic Chat)
  - DETECT -> VERIFY -> VERDICT           - First-Principles CoT                  - Suora asiantuntijakeskustelu
  - Fatal Error Overrule                  - SFT/QLoRA -synteesi                   - Anti-Sycophancy (ei mielistelyä)
  - Tiukka JSON-tuomio                    - 0.00 % kieltäytymisiä                 - Reasoning: Spoon / Einstein / Deep

Toimintatilat ja Kutsutagit:

  1. Tila 1: Materiaalin tarkistus (Quality Inspector): * Aktivointi: [TILA: TARKISTUS], [INSPECT] tai API-parametri sceptic_mode="inspect". * Toimintamalli: Malli auditoi annetun näytteen sisäisesti (<think>: DETECT -> VERIFY -> VERDICT) ja palauttaa ainoastaan tiukan JSON-objektin: json { "verdict": "ACCEPT | REVIEW | REJECT", "fatal_error": false, "error_class": null, "exact_error_location": null, "brief_rationale": "Ytimekäs 1-2 virkkeen perustelu havainnoista." }
  2. Tila 2: Koulutusmateriaalin tuottaminen (Data Curation): * Aktivointi: [TILA: KOULUTUS], [CURATE] tai API-parametri sceptic_mode="curate". * Toimintamalli: Jalostaa syötteestä korkealaatuisia SFT/QLoRA-harjoitusnäytteitä aukottomalla first-principles CoT -päättelyketjulla ilman esipuheita tai tyhjiä täytesanoja.
  3. Tila 3: Sceptic-keskustelutila (Asiantuntijakeskustelu): * Aktivointi: Oletustila ilman tageja tai [TILA: KESKUSTELU]. * Toimintamalli: Suora, kriittinen ja itsenäinen asiantuntija ilman small talkia tai moraalisaarnoja.

2. Empiiriset Suorituskykymittaukset (2x RTX 5090 Blackwell TP=2)

Mittaukset ajettu tuotantoasetuksilla (2x RTX 5090 32GB Blackwell, TP=2, CUDAGraphs ON, FlashInfer Marlin MoE):

Testitapaus / Toimintatila Input-tokenit Output-tokenit TTFT (ms) Generointiaika (s) Generointinopeus (tok/s)
Lämpöajo (Warmup) 1 191 30 77.23 ms 0.108 s 276.43 tok/s
Tila 1: Tarkistus (FI) 659 512 71.56 ms 1.419 s 360.83 tok/s
Tila 1: Inspection (EN) 664 373 71.26 ms 1.131 s 329.86 tok/s
Tila 2: Koulutus (FI) 674 847 66.49 ms 2.991 s 283.19 tok/s
Tila 2: Curation (EN) 636 1 111 68.45 ms 3.800 s 292.40 tok/s
Tila 3: Keskustelu Spoon (FI) 1 244 185 70.94 ms 0.699 s 264.81 tok/s
Tila 3: Keskustelu Einstein (FI) 1 259 530 70.28 ms 1.860 s 285.02 tok/s
KOKONAISKESKIARVO — 3 558 tok 69.83 ms — 298.99 tok/s

3. Käyttö vLLM-infrassa (Deployment Guide)

python3 -m vllm.entrypoints.openai.api_server \
  --model kataguru/Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 \
  --tensor-parallel-size 2 \
  --gpu-memory-utilization 0.85 \
  --max-model-len 4096 \
  --quantization compressed-tensors \
  --kv-cache-dtype fp8 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --host 0.0.0.0 \
  --port 8001 \
  --trust-remote-code

4. Lisenssi & Tekijätiedot

  • Lisenssi: Apache-2.0
  • Kehittäjä: Kataguru AI Research
  • Arkkitehtuuri: Qwen3.5 Sparse MoE (35B-A3B) + Blackwell NVFP4 Optimization + Master Tri-Mode Chat Template

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
40
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe
Quantization
compressed-tensors

Identity and Version

Repository
kataguru/Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4
Publisher
Jari Katajisto
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
36B parameters
Languages
fi, en
Revision
cfdddce117872abc51b47742fd9a6777396debe5
First published
2026-10-07
Last updated
2026-10-07

Files and Weights

20 files, 25.7 GB in total. The weights are 6 files totalling 25.6 GB in safetensors.

Weights6 files · 25.6 GB
Configuration6 files · 10.8 MB
Tokenizer4 files · 30.1 MB
Documentation1 file · 8.9 KB
Other2 files · 2.1 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00006.safetensorsWeights4.5 GB e661ee9e9d66
model-00002-of-00006.safetensorsWeights4.5 GB 6d8c905f367a
model-00003-of-00006.safetensorsWeights4.5 GB b96ff5b8b65c
model-00004-of-00006.safetensorsWeights4.5 GB bef3b84f5c7f
model-00005-of-00006.safetensorsWeights4.5 GB 29526809ac10
model-00006-of-00006.safetensorsWeights3.1 GB 66804f168f83
config.jsonConfiguration76.5 KB —
generation_config.jsonConfiguration213 B —
model.safetensors.index.jsonConfiguration10.7 MB 8882eec7cbb2
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration1.2 KB —
video_preprocessor_config.jsonConfiguration385 B —
README.mdDocumentation8.9 KB —
chat_template.jinjaOther24.5 KB —
inspector.pngOther2.1 MB c9231588a15d
.gitattributesRepository1.7 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer20.0 MB 87a7830d63fc
tokenizer_config.jsonTokenizer26.2 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
25.6 GB
Download from Jari Katajisto

Released by Jari Katajisto through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published25.6 GB
16-bit71.9 GB
8-bit36.0 GB
4-bit18.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4

How much GPU memory does Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 need?

About 86.3 GB at 16-bit and 21.6 GB at 4-bit: the weights (36B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 commercially?

Yes. Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4 is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

Tema_Q-X7-Thinking

Tema_Q LLM

TemaQ-X7-Thinking(天馬求) は、Ornith AIが開発したモデル Ornith-1.5 を基盤にした、エージェント向けの大規模言語モデル(LLM)です。 TemaQ-X7-Thinking (TemaQ model) is an enhanced Large Language Model (LLM) designed for agent applications, based on the Ornith-1.5 model developed by Ornith AI. It is engineered to generate more flexible and useful responses, even for prompts that are difficult for standard models to handle effectively. When used in conjunction with TemaQ Agent, it enables advanced reasoning capabilities. ユーザーの責任: モデルの利用者は、生成されたコンテンツが、適用される法律、規制、およびHugging Faceの利用規約/コンテンツポリシーに準拠することを全面的に保証する必要があります。

Open weights 36B parameters 262,144 tokens

Model · Text generation

APUS-OpenJev-v1-35B-A3B

APUS AI

A Qwen3.5 MoE decision model for choosing browser actions, selecting workflow steps, and judging natural-language criteria. Give the model a shared state and a set of candidate actions; the included decision runtime returns a distribution over those candidates. This repository contains 35B-A3B checkpoint-5949 merged BF16 weights. It is a standalone model with root-level Hugging Face configuration and weights, requiring no separate LoRA adapter. The native merged model scores 71/80 (88.75%) at its full 40-layer depth on the Frozen80 development panel. The decision interface accepts 2–16 request-specific candidates. Applications can use the returned candidate IDs to dispatch actions or build…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

O-noinoc seed 0: on-policy distillation (OPD) of the untrained Qwen/Qwen3.6-35B-A3B toward the teacher rewardhack/qwen3.6-35b-a3b-hacksft-vanilla-873rows-ep3, V1 (vanilla SFT: hacks with or without being asked); elicitation prompt off in the student's rollouts; seed 0. Full merged weights (bf16 safetensors, the standard Qwen35MoeForConditionalGeneration layout, loads with transformers or vLLM like the base model) of a LoRA (r=32) trained from a fresh init, from the Terminal Wrench reward-hacking / inoculation project (Gaokai Zhang, Songwen Zhao, Juan Manuel Suárez). On-policy distillation on Tinker, 24 iterations. Each iteration the current student ran the terminus-2 agent (harbor, local…

Open weights cc-by-sa-4.0 36B parameters 262,144 tokens transformers

Kataguru Sceptic Quality Inspector v1.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and high-throughput expert model built upon the Sparse Mixture-of-Experts (MoE) foundation of Kataguru Sceptic 35B-A3B (35 billion total parameters, 3 billion activated per token). Hardware-accelerated for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token (TTFT) while consuming only 12.1 GiB VRAM per GPU on dual RTX 5090 Blackwell hardware (TP=2). 1. Master Tri-Mode Operation: 2. Native Multimodal Vision…

Open weights apache-2.0 36B parameters 262,144 tokens

Model · Text generation

Cyber-F1-smoke

Autumn

This model is a fine-tuned version of DuyTa/Cyber-F1. It has been trained using TRL. This model was trained with SFT. - PEFT 0.21.0

Open weights 35.1B parameters 262,144 tokens peft