# Tema_Q-X7-Thinking by Tema_Q LLM: Open-Weight Model
Source: https://savrn.com/models/tema-q-x7-thinking
Markdown alternate of the page above; the site index is https://savrn.com/llms.txt

---

## Runs On

What it takes to serve Tema_Q-X7-Thinking (36B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

| Precision | Weights | Memory needed | Cheapest setup | Per hour | Also fits |
| --- | --- | --- | --- | --- | --- |
| 16-bit | 71.9 GB | 86.3 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 · [1x MI355X](https://savrn.com/ai-index/pricing/gpus/mi355x) $2.59 |
| 8-bit | 36.0 GB | 43.1 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |
| 4-bit | 18.0 GB | 21.6 GB | 1x [MI300X](https://savrn.com/ai-index/pricing/gpus/mi300x) (192 GB) Vultr | $1.85 | [1x H100](https://savrn.com/ai-index/pricing/gpus/h100) $1.99 · [1x MI325X](https://savrn.com/ai-index/pricing/gpus/mi325x) $2.00 |

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the [SAVRN Index](https://savrn.com/ai-index/pricing/gpus), read Oct 7, 2026.

[Tema_Q-X7-Thinking on every accelerator the SAVRN Index prices, at every precision](https://savrn.com/models/tema-q-x7-thinking/gpus)

## Model Card

TemaQ-X7-Thinking（天馬求） は、Ornith AIが開発したモデル Ornith-1.5 を基盤にした、エージェント向けの大規模言語モデル（LLM）です。 TemaQ-X7-Thinking (TemaQ model) is an enhanced Large Language Model (LLM) designed for agent applications, based on the Ornith-1.5 model developed by Ornith AI. It is engineered to generate more flexible and useful responses, even for prompts that are difficult for standard models to handle effectively. When used in conjunction with TemaQ Agent, it enables advanced reasoning capabilities. ユーザーの責任: モデルの利用者は、生成されたコンテンツが、適用される法律、規制、およびHugging Faceの利用規約/コンテンツポリシーに準拠することを全面的に保証する必要があります。

Excerpt from the card by Tema_Q LLM.

## Configuration

Architecture

Qwen3_5MoeForConditionalGeneration

Context length (tokens)

262,144

Layers

40

Hidden size

2,048

Attention heads

16

Key/value heads

2

Head dimension

256

Vocabulary size

248,320

Experts

256

Experts active per token

8

Model type

qwen3_5_moe

## Identity and Version

Repository

temaq-org/Tema_Q-X7-Thinking

Publisher

Tema_Q LLM

Task

Text generation

Modality

Text

Library

Not stated by the source

Parameters

36B parameters

Languages

ja, en, zh

Revision

97235f8d9f0db9423d66b28126603e599a272cc8

First published

2026-08-20

Last updated

2026-10-03

## Files and Weights

35 files, 71.9 GB in total. The weights are 20 files totalling 71.9 GB in safetensors.

Weights20 files · 71.9 GB

Configuration7 files · 173.1 KB

Tokenizer4 files · 22.9 MB

Documentation1 file · 2.3 KB

Other2 files · 778.7 KB

Repository1 file · 1.6 KB

Every file

| File | Type | Size | SHA-256 |
| --- | --- | --- | --- |
| model-00001-of-00020.safetensors | Weights | 4.3 GB | 266b698af735 |
| model-00002-of-00020.safetensors | Weights | 4.3 GB | c44b97c5f210 |
| model-00003-of-00020.safetensors | Weights | 4.0 GB | 6acf50d54c6b |
| model-00004-of-00020.safetensors | Weights | 3.4 GB | 3e87b30d6e96 |
| model-00005-of-00020.safetensors | Weights | 3.9 GB | 7646cc10e456 |
| model-00006-of-00020.safetensors | Weights | 4.0 GB | eed5e5d2d582 |
| model-00007-of-00020.safetensors | Weights | 3.8 GB | 1c631a0cfa0a |
| model-00008-of-00020.safetensors | Weights | 4.0 GB | 4d123daaccbc |
| model-00009-of-00020.safetensors | Weights | 4.0 GB | d0f7e5f328c4 |
| model-00010-of-00020.safetensors | Weights | 3.9 GB | cffc4351dd3b |
| model-00011-of-00020.safetensors | Weights | 3.4 GB | 64117f98c7cd |
| model-00012-of-00020.safetensors | Weights | 3.4 GB | a6141d2e1c48 |
| model-00013-of-00020.safetensors | Weights | 3.4 GB | 5de603323323 |
| model-00014-of-00020.safetensors | Weights | 3.4 GB | b995f775ed2f |
| model-00015-of-00020.safetensors | Weights | 3.4 GB | fc7d32624a9c |
| model-00016-of-00020.safetensors | Weights | 3.4 GB | ac1012460a86 |
| model-00017-of-00020.safetensors | Weights | 3.4 GB | 792917701887 |
| model-00018-of-00020.safetensors | Weights | 3.4 GB | 33e742189e06 |
| model-00019-of-00020.safetensors | Weights | 4.3 GB | 470b7a48c931 |
| model-00020-of-00020.safetensors | Weights | 1.2 GB | cae31f39e2d7 |
| config.json | Configuration | 3.3 KB | — |
| configuration.json | Configuration | 58 B | — |
| generation_config.json | Configuration | 202 B | — |
| model.safetensors.index.json | Configuration | 167.5 KB | — |
| preprocessor_config.json | Configuration | 390 B | — |
| processor_config.json | Configuration | 1.2 KB | — |
| video_preprocessor_config.json | Configuration | 385 B | — |
| README.md | Documentation | 2.3 KB | — |
| Tema_Q-X-Logo.jpg | Other | 771.2 KB | be706e80f65d |
| chat_template.jinja | Other | 7.5 KB | — |
| .gitattributes | Repository | 1.6 KB | — |
| merges.txt | Tokenizer | 3.4 MB | — |
| tokenizer.json | Tokenizer | 12.8 MB | 5f9e4d4901a9 |
| tokenizer_config.json | Tokenizer | 16.7 KB | — |
| vocab.json | Tokenizer | 6.7 MB | — |

## License and Download

License

Not stated by the source

Access

Open weights, no gate

Download size

71.9 GB

[Download from Tema_Q LLM](https://huggingface.co/temaq-org/Tema_Q-X7-Thinking)

Released by Tema_Q LLM through its official repository on Hugging Face.

## Built From

- Derived from ornith-ai/Ornith-1.5-35B-A3B

## Memory Requirements

| Precision | Weights in memory |
| --- | --- |
| As published | 71.9 GB |
| 16-bit | 71.9 GB |
| 8-bit | 36.0 GB |
| 4-bit | 18.0 GB |

Weights only, from the published parameter count; the key-value cache and runtime add to this.

## Questions About Tema_Q-X7-Thinking

### How much GPU memory does Tema_Q-X7-Thinking need?

About 86.3 GB at 16-bit and 21.6 GB at 4-bit: the weights (36B parameters) plus a working margin. A long context needs more.

### What is the cheapest GPU to run Tema_Q-X7-Thinking on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

### What is Tema_Q-X7-Thinking's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

## Similar Models

Model · Text generation

### [APUS-OpenJev-v1-35B-A3B](https://savrn.com/models/apus-openjev-v1-35b-a3b)

[APUS AI](https://savrn.com/model-publishers/apus-ailab)

A Qwen3.5 MoE decision model for choosing browser actions, selecting workflow steps, and judging natural-language criteria. Give the model a shared state and a set of candidate actions; the included decision runtime returns a distribution over those candidates. This repository contains 35B-A3B checkpoint-5949 merged BF16 weights. It is a standalone model with root-level Hugging Face configuration and weights, requiring no separate LoRA adapter. The native merged model scores 71/80 (88.75%) at its full 40-layer depth on the Frozen80 development panel. The decision interface accepts 2–16 request-specific candidates. Applications can use the returned candidate IDs to dispatch actions or build…

Open weights apache-2.0 36B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/apus-openjev-v1-35b-a3b)

Model · Text generation

### [qwen3.6-35b-a3b-hackopd-vanilla873-noinoc-s0](https://savrn.com/models/qwen3-6-35b-a3b-hackopd-vanilla873-noinoc-s0)

[RewardHacking](https://savrn.com/model-publishers/rewardhack)

O-noinoc seed 0: on-policy distillation (OPD) of the untrained Qwen/Qwen3.6-35B-A3B toward the teacher rewardhack/qwen3.6-35b-a3b-hacksft-vanilla-873rows-ep3, V1 (vanilla SFT: hacks with or without being asked); elicitation prompt off in the student's rollouts; seed 0. Full merged weights (bf16 safetensors, the standard Qwen35MoeForConditionalGeneration layout, loads with transformers or vLLM like the base model) of a LoRA (r=32) trained from a fresh init, from the Terminal Wrench reward-hacking / inoculation project (Gaokai Zhang, Songwen Zhao, Juan Manuel Suárez). On-policy distillation on Tinker, 24 iterations. Each iteration the current student ran the terminus-2 agent (harbor, local…

Open weights cc-by-sa-4.0 36B parameters 262,144 tokens transformers

[View model](https://savrn.com/models/qwen3-6-35b-a3b-hackopd-vanilla873-noinoc-s0)

Model · Text generation

### [Kataguru-Sceptic-Quality-Inspector-v1.0-NVFP4](https://savrn.com/models/kataguru-sceptic-quality-inspector-v1-0-nvfp4)

[Jari Katajisto](https://savrn.com/model-publishers/kataguru)

Kataguru Sceptic Quality Inspector v1.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and high-throughput expert model built upon the Sparse Mixture-of-Experts (MoE) foundation of Kataguru Sceptic 35B-A3B (35 billion total parameters, 3 billion activated per token). Hardware-accelerated for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token (TTFT) while consuming only 12.1 GiB VRAM per GPU on dual RTX 5090 Blackwell hardware (TP=2). 1. Master Tri-Mode Operation: 2. Native Multimodal Vision…

Open weights apache-2.0 36B parameters 262,144 tokens

[View model](https://savrn.com/models/kataguru-sceptic-quality-inspector-v1-0-nvfp4)

Model · Text generation

### [Kataguru-Sceptic-Quality-Inspector-v2.0-NVFP4](https://savrn.com/models/kataguru-sceptic-quality-inspector-v2-0-nvfp4)

[Jari Katajisto](https://savrn.com/model-publishers/kataguru)

Kataguru Sceptic Quality Inspector v2.0 (NVFP4) is a sovereign, non-sycophantic LLM-as-a-Judge, automated dataset auditing engine, and ultra-high-throughput expert model. Built upon the fleet-record foundation of Kataguru Sceptic Multi-Mode v2 CP1200 (72.01% multi-domain record, FAR = 0.000%) merged with the 16.4k certified forensic quality inspection task vector, it is purpose-engineered to serve as the fleet's primary data firewall and filtration engine. Hardware-optimized for NVIDIA RTX 50-series Blackwell architecture using native NVFP4 quantization (FP4 weights with FP8 activation scales), it achieves generation speeds of ~300–360 tok/s and an ultra-low ~70 ms Time To First Token…

Open weights apache-2.0 36B parameters 262,144 tokens

[View model](https://savrn.com/models/kataguru-sceptic-quality-inspector-v2-0-nvfp4)

Model · Text generation

### [LUHANADJolTctKlt](https://savrn.com/models/luhanadjoltctklt)

[Priyansh Singhal](https://savrn.com/model-publishers/ps4research)

This seedoss model was trained 2x faster with Unsloth and Huggingface's TRL library.

Open weights apache-2.0 36.2B parameters 524,288 tokens transformers

[View model](https://savrn.com/models/luhanadjoltctklt)

Model · Text generation

### [Cyber-F1-smoke](https://savrn.com/models/cyber-f1-smoke)

[Autumn](https://savrn.com/model-publishers/autumn10)

This model is a fine-tuned version of DuyTa/Cyber-F1. It has been trained using TRL. This model was trained with SFT. - PEFT 0.21.0

Open weights 35.1B parameters 262,144 tokens peft

[View model](https://savrn.com/models/cyber-f1-smoke)

## Tema_Q LLM

[All models and datasets](https://savrn.com/model-publishers/temaq-org)

## Versions

- [97235f8d9f0d](https://savrn.com/models/tema-q-x7-thinking/versions/97235f8d9f0d) · current 2026-10-03

## Explore More

- [All text generation models](https://savrn.com/models/tasks/text-generation)
- [Model comparisons](https://savrn.com/models/comparisons)
- [The model directory](https://savrn.com/models)
- [Open model prices by host](https://savrn.com/ai-index/pricing/open-models)

## Source

- Repository metadata, read 2026-10-03.
- [Hugging Face record](https://huggingface.co/temaq-org/Tema_Q-X7-Thinking)
- [How the hub is built](https://savrn.com/model-hub/methodology)
