SAVRN
Search Contact SAVRN

Open-weight model · Text generation

OTel-2.0-LLM-31B-IT

by Farbod Tavakkoli farbodtavakkoli/OTel-2.0-LLM-31B-IT

OTel-2.0-LLM-31B-IT is a telecom-specialized instruction model post-trained from Gemma 4 31B-IT on approximately 440 billion telecom training tokens.

Parameters31.3B
Context262,144
Weights126.8 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads7.2M

Runs On

What it takes to serve OTel-2.0-LLM-31B-IT (31.3B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 62.5 GB 75.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 31.3 GB 37.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 15.6 GB 18.8 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

SAVRN's Notes on OTel-2.0-LLM-31B-IT

Post-trained from Gemma 4 31B-IT on roughly 440 billion telecom tokens, OTel-2.0-LLM-31B-IT is aimed at network operations, standards interpretation, configuration help and retrieval-backed answering. Its 31.3B parameters need 75.1 GB at 16-bit, 37.5 GB at 8-bit and 18.8 GB at 4-bit, and all three fit the 192 GB of one MI300X, the cheapest Index setup at $1.85 an hour. The 262,144-token context matters when the input is a standards document rather than a ticket.

Apache 2.0 allows commercial deployment, modification and redistribution, with the notice and change-statement duties kept. Two things to verify. The repository is 39 files and 126.8 GB, about twice the 62.5 GB of 16-bit weights, so confirm what those files hold before pulling them. And because Farbod Tavakkoli post-trained this from a Gemma 4 base, confirm the terms that travel with the base weights fit your plan; no Index host prices are listed.

Model Card

By Farbod Tavakkoli, published under apache-2.0, revision 6d425c390399.

Versioning notice: This card describes the current OTel 2.0 OSFT checkpoint. The model weights may be updated over time. For reproducible evaluation or production deployment, pin a specific model revision, checkpoint hash, or release tag rather than relying on the floating latest version.

OTel-2.0-LLM-31B-IT is a telecom-specialized instruction model post-trained from Gemma 4 31B-IT on approximately 440 billion telecom training tokens. It is the first release in the OTel 2.0 family and is designed to support telco-grade AI workflows across network operations, standards interpretation, product development, network configuration assistance, RAG, and telecom-specific question answering.

OTel 2.0 extends the original OTel effort from a RAG-oriented telecom fine-tuning release into a larger domain-adapted training program. The model was trained from a much larger standards and telecom corpus, with new data preparation coverage for direct telecom QnA, abstention, RAG, base-model-style telecom data, and general-purpose instruction-following and tool-calling examples. The current training mixture does not include telecommunications-specific MCP, tool-calling, or instruction-following examples.

Read the full model card (5,000 words)

Configuration

Architecture
Gemma4ForConditionalGeneration
Context length (tokens)
262,144
Layers
60
Hidden size
5,376
Feed-forward size
21,504
Attention heads
32
Key/value heads
16
Head dimension
256
Vocabulary size
262,144
Sliding window (tokens)
1,024
Model type
gemma4

Identity and Version

Repository
farbodtavakkoli/OTel-2.0-LLM-31B-IT
Publisher
Farbod Tavakkoli
Task
Text generation
Modality
Text
Library
transformers
Parameters
31.3B parameters
Languages
en
Revision
6d425c3903997f9540be33f507ad363848b0fa06
First published
2026-07-23
Last updated
2026-09-08

Files and Weights

39 files, 126.8 GB in total. The weights are 29 files totalling 126.8 GB in safetensors.

Weights29 files · 126.8 GB
Configuration5 files · 126.6 KB
Tokenizer2 files · 32.2 MB
Documentation1 file · 40.3 KB
Other1 file · 19.3 KB
Repository1 file · 2.1 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00014.safetensorsWeights2.8 GB 2d8a6ceb85e0
model-00001-of-00015.safetensorsWeights2.8 GB 05d6cc31f1d4
model-00002-of-00014.safetensorsWeights5.0 GB 0572c4159bd0
model-00002-of-00015.safetensorsWeights2.2 GB d436d321c12c
model-00003-of-00014.safetensorsWeights4.9 GB e8ff2c96a762
model-00003-of-00015.safetensorsWeights4.9 GB d2a0ecb1317a
model-00004-of-00014.safetensorsWeights4.9 GB 196f06858c74
model-00004-of-00015.safetensorsWeights4.9 GB 3f708ec3ae58
model-00005-of-00014.safetensorsWeights4.9 GB 83cf22e97659
model-00005-of-00015.safetensorsWeights4.9 GB bc3a11f4b53a
model-00006-of-00014.safetensorsWeights4.8 GB cce5e5e46a5a
model-00006-of-00015.safetensorsWeights4.8 GB ecb8e0e7c4b4
model-00007-of-00014.safetensorsWeights4.9 GB b6be3d84c71d
model-00007-of-00015.safetensorsWeights4.9 GB e0d8dd8f5e11
model-00008-of-00014.safetensorsWeights4.9 GB 70af83af358a
model-00008-of-00015.safetensorsWeights4.9 GB 120ec437e88d
model-00009-of-00014.safetensorsWeights4.9 GB 16f5cac9f9dd
model-00009-of-00015.safetensorsWeights4.9 GB ede9fbe726b3
model-00010-of-00014.safetensorsWeights4.9 GB 9e75adcd906a
model-00010-of-00015.safetensorsWeights4.9 GB cbcb624d28f9
model-00011-of-00014.safetensorsWeights4.9 GB dc0842a77b6c
model-00011-of-00015.safetensorsWeights4.9 GB 5f4d60a341f7
model-00012-of-00014.safetensorsWeights4.8 GB b2deff051ec3
model-00012-of-00015.safetensorsWeights4.8 GB 1b9a619e41ba
model-00013-of-00014.safetensorsWeights4.9 GB 7b0ac88cb864
model-00013-of-00015.safetensorsWeights4.9 GB 16ba7c1662a7
model-00014-of-00014.safetensorsWeights2.7 GB 6f853a792175
model-00014-of-00015.safetensorsWeights2.7 GB 7cd821f4b598
model-00015-of-00015.safetensorsWeights1.2 GB 217b3dbd53ad
config.jsonConfiguration4.4 KB
generation_config.jsonConfiguration208 B
model.safetensors.index.jsonConfiguration120.2 KB
processor_config.jsonConfiguration1.7 KB
special_tokens_map.jsonConfiguration100 B
README.mdDocumentation40.3 KB
chat_template.jinjaOther19.3 KB
.gitattributesRepository2.1 KB
tokenizer.jsonTokenizer32.2 MB cc8d3a0ce364
tokenizer_config.jsonTokenizer3.7 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
126.8 GB
Download from Farbod Tavakkoli

Released by Farbod Tavakkoli through its official repository on Hugging Face. Read the license.

Built From

  • Described by arXiv:2504.07097

Memory Requirements

PrecisionWeights in memory
As published126.8 GB
16-bit62.5 GB
8-bit31.3 GB
4-bit15.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Compare OTel-2.0-LLM-31B-IT

Questions About OTel-2.0-LLM-31B-IT

How much GPU memory does OTel-2.0-LLM-31B-IT need?

About 75.1 GB at 16-bit and 18.8 GB at 4-bit: the weights (31.3B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run OTel-2.0-LLM-31B-IT on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use OTel-2.0-LLM-31B-IT commercially?

Yes. OTel-2.0-LLM-31B-IT is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is OTel-2.0-LLM-31B-IT's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Text generation

GLM-4.7-Flash

Z.ai

Join our Discord community. Check out the GLM-4.7 technical blog, technical report(GLM-4.5). Use GLM-4.7-Flash API services on Z.ai API Platform. One click to GLM-4.7. GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency. Default Settings (Most Tasks) For multi-turn agentic tasks (τ²-Bench and Terminal Bench 2), please turn on Preserved Thinking mode. Terminal Bench, SWE Bench Verified τ^2-Bench For τ^2-Bench evaluation, we added an additional prompt to the Retail and Telecom user interaction to avoid failure modes caused by users ending the interaction…

Open weights mit 31.2B parameters 202,752 tokens transformers

Model · Text generation

Qwen3-VL-30B-A3B-Instruct-AWQ

QuantTrio

As of 2025-10-08, create a fresh Python environment and run: For more details, refer to vLLM Official Qwen3-VL Guide Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Text Understanding on par with pure LLMs: Seamless text–vision…

Open weights apache-2.0 31.1B parameters 262,144 tokens transformers

Model · Text generation

NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

NVIDIA

September 2025 \- December 2025 The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\. Nemotron-3-Nano-30B-A3B-BF16 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and tasks by first generating a reasoning trace and then concluding with a final response. The model's reasoning capabilities can be configured through a flag in the chat template. If the user prefers the model to provide its final answer without intermediate reasoning traces, it can be configured to do so, albeit with a slight decrease in…

Open weights other 31.6B parameters 262,144 tokens transformers

Fastino-Nemotron-3.5-Lightning-Finance is a 30B-parameter, 3B-active mixture-of-experts model specialized for financial reasoning, extraction, and research fine-tuned on LoRA with the Fastino Fine-Tuning Agent. The model targets financial document reasoning, numerical question answering over filings and tables, numeric span extraction, financial entity recognition, conversational analysis, and source-grounded financial research. The evaluation suite includes FinQA, TAT-QA, SEC-Num, FinEntity, BizFinBench, BigFinanceBench, ConvFinQA, and FiQA. The published weights are BF16 and require about 66 GB before runtime overhead. An 80 GB or larger GPU, or tensor parallelism across multiple GPUs, is…

Open weights apache-2.0 31.6B parameters 262,144 tokens transformers

Model · Text generation

Qwen3-Coder-30B-A3B-Instruct-FP8

Qwen

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format. Qwen3-Coder-30B-A3B-Instruct-FP8 has the following features: NOTE: This model…

Open weights apache-2.0 30.5B parameters 262,144 tokens transformers