SAVRN
Search Contact SAVRN

Open-weight model · Text generation

Tema_Q-X7-Thinking

by Tema_Q LLM temaq-org/Tema_Q-X7-Thinking

Tema_Q-X7-Thinking is an open-weight model for text generation from Tema_Q LLM. It has 36B parameters and a 262,144-token context. At 16-bit it needs about 86.3 GB of GPU memory, which fits on 1x MI300X from $1.85 an hour, at the lowest prices in the SAVRN Index. It draws 611 downloads a month.

TemaQ-X7-Thinking(天馬求) は、Ornith AIが開発したモデル Ornith-1.5 を基盤にした、エージェント向けの大規模言語モデル(LLM)です。 TemaQ-X7-Thinking (TemaQ model) is an enhanced Large Language Model (LLM) designed for agent applications, based on the Ornith-1.5 model developed by Ornith AI.

Parameters36B
Context262,144
Weights71.9 GB
License—
AccessOpen weights
Monthly Downloads611

Runs On

What it takes to serve Tema_Q-X7-Thinking (36B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 71.9 GB 86.3 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
8-bit 36.0 GB 43.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 18.0 GB 21.6 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Tema_Q-X7-Thinking on every accelerator the SAVRN Index prices, at every precision

Model Card

TemaQ-X7-Thinking(天馬求) は、Ornith AIが開発したモデル Ornith-1.5 を基盤にした、エージェント向けの大規模言語モデル(LLM)です。 TemaQ-X7-Thinking (TemaQ model) is an enhanced Large Language Model (LLM) designed for agent applications, based on the Ornith-1.5 model developed by Ornith AI. It is engineered to generate more flexible and useful responses, even for prompts that are difficult for standard models to handle effectively. When used in conjunction with TemaQ Agent, it enables advanced reasoning capabilities. ユーザーの責任: モデルの利用者は、生成されたコンテンツが、適用される法律、規制、およびHugging Faceの利用規約/コンテンツポリシーに準拠することを全面的に保証する必要があります。

Excerpt from the card by Tema_Q LLM.

Configuration

Architecture
Qwen3_5MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
40
Hidden size
2,048
Attention heads
16
Key/value heads
2
Head dimension
256
Vocabulary size
248,320
Experts
256
Experts active per token
8
Model type
qwen3_5_moe

Identity and Version

Repository
temaq-org/Tema_Q-X7-Thinking
Publisher
Tema_Q LLM
Task
Text generation
Modality
Text
Library
Not stated by the source
Parameters
36B parameters
Languages
ja, en, zh
Revision
97235f8d9f0db9423d66b28126603e599a272cc8
First published
2026-08-20
Last updated
2026-10-03

Files and Weights

35 files, 71.9 GB in total. The weights are 20 files totalling 71.9 GB in safetensors.

Weights20 files · 71.9 GB
Configuration7 files · 173.1 KB
Tokenizer4 files · 22.9 MB
Documentation1 file · 2.3 KB
Other2 files · 778.7 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00020.safetensorsWeights4.3 GB 266b698af735
model-00002-of-00020.safetensorsWeights4.3 GB c44b97c5f210
model-00003-of-00020.safetensorsWeights4.0 GB 6acf50d54c6b
model-00004-of-00020.safetensorsWeights3.4 GB 3e87b30d6e96
model-00005-of-00020.safetensorsWeights3.9 GB 7646cc10e456
model-00006-of-00020.safetensorsWeights4.0 GB eed5e5d2d582
model-00007-of-00020.safetensorsWeights3.8 GB 1c631a0cfa0a
model-00008-of-00020.safetensorsWeights4.0 GB 4d123daaccbc
model-00009-of-00020.safetensorsWeights4.0 GB d0f7e5f328c4
model-00010-of-00020.safetensorsWeights3.9 GB cffc4351dd3b
model-00011-of-00020.safetensorsWeights3.4 GB 64117f98c7cd
model-00012-of-00020.safetensorsWeights3.4 GB a6141d2e1c48
model-00013-of-00020.safetensorsWeights3.4 GB 5de603323323
model-00014-of-00020.safetensorsWeights3.4 GB b995f775ed2f
model-00015-of-00020.safetensorsWeights3.4 GB fc7d32624a9c
model-00016-of-00020.safetensorsWeights3.4 GB ac1012460a86
model-00017-of-00020.safetensorsWeights3.4 GB 792917701887
model-00018-of-00020.safetensorsWeights3.4 GB 33e742189e06
model-00019-of-00020.safetensorsWeights4.3 GB 470b7a48c931
model-00020-of-00020.safetensorsWeights1.2 GB cae31f39e2d7
config.jsonConfiguration3.3 KB —
configuration.jsonConfiguration58 B —
generation_config.jsonConfiguration202 B —
model.safetensors.index.jsonConfiguration167.5 KB —
preprocessor_config.jsonConfiguration390 B —
processor_config.jsonConfiguration1.2 KB —
video_preprocessor_config.jsonConfiguration385 B —
README.mdDocumentation2.3 KB —
Tema_Q-X-Logo.jpgOther771.2 KB be706e80f65d
chat_template.jinjaOther7.5 KB —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 5f9e4d4901a9
tokenizer_config.jsonTokenizer16.7 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
Not stated by the source
Access
Open weights, no gate
Download size
71.9 GB
Download from Tema_Q LLM

Released by Tema_Q LLM through its official repository on Hugging Face.

Built From

  • Derived from ornith-ai/Ornith-1.5-35B-A3B

Memory Requirements

PrecisionWeights in memory
As published71.9 GB
16-bit71.9 GB
8-bit36.0 GB
4-bit18.0 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Tema_Q-X7-Thinking

How much GPU memory does Tema_Q-X7-Thinking need?

About 86.3 GB at 16-bit and 21.6 GB at 4-bit: the weights (36B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Tema_Q-X7-Thinking on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

What is Tema_Q-X7-Thinking's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

APUS-OpenJev-v1-35B-A3B

APUS AI

A Qwen3.5 MoE decision model for choosing browser actions, selecting workflow steps, and judging natural-language criteria. Give the model a shared state and a set of candidate actions; the included decision runtime returns a distribution over those candidates. This repository contains 35B-A3B checkpoint-5949 merged BF16 weights. It is a standalone model with root-level Hugging Face configuration and weights, requiring no separate LoRA adapter. The native merged model scores 71/80 (88.75%) at its full 40-layer depth on the Frozen80 development panel. The decision interface accepts 2–16 request-specific candidates. Applications can use the returned candidate IDs to dispatch actions or build…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

O-noinoc seed 0: on-policy distillation (OPD) of the untrained Qwen/Qwen3.6-35B-A3B toward the teacher rewardhack/qwen3.6-35b-a3b-hacksft-vanilla-873rows-ep3, V1 (vanilla SFT: hacks with or without being asked); elicitation prompt off in the student's rollouts; seed 0. Full merged weights (bf16 safetensors, the standard Qwen35MoeForConditionalGeneration layout, loads with transformers or vLLM like the base model) of a LoRA (r=32) trained from a fresh init, from the Terminal Wrench reward-hacking / inoculation project (Gaokai Zhang, Songwen Zhao, Juan Manuel Suárez). On-policy distillation on Tinker, 24 iterations. Each iteration the current student ran the terminus-2 agent (harbor, local…

Open weights cc-by-sa-4.0 36B parameters 262,144 tokens transformers

Kataguru Sceptic Quality Inspector v1.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and high-throughput expert model built upon the Sparse Mixture-of-Experts (MoE) foundation of Kataguru Sceptic 35B-A3B (35 billion total parameters, 3 billion activated per token). Hardware-accelerated for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token (TTFT) while consuming only 12.1 GiB VRAM per GPU on dual RTX 5090 Blackwell hardware (TP=2). 1. Master Tri-Mode Operation: 2. Native Multimodal Vision…

Open weights apache-2.0 36B parameters 262,144 tokens

Kataguru Sceptic Quality Inspector v2.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and ultra-high-throughput expert model. Built upon the fleet-record foundation of Kataguru Sceptic Multi-Mode v2 CP1200 (72.01% multi-domain record, FAR = 0.000%) merged with the 16.4k certified forensic quality inspection task vector, it is purpose-engineered to serve as the fleet's primary data firewall and filtration engine. Hardware-optimized for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token…

Open weights apache-2.0 36B parameters 262,144 tokens

Model · Text generation

Cyber-F1-smoke

Autumn

This model is a fine-tuned version of DuyTa/Cyber-F1. It has been trained using TRL. This model was trained with SFT. - PEFT 0.21.0

Open weights 35.1B parameters 262,144 tokens peft