SAVRN
Search Contact SAVRN

Open-weight model · Text generation

xflux_text_encoders

by XLabs AI XLabs-AI/xflux_text_encoders

Text encoder weights from Google's T5 model

Parameters4.8B
Context
Weights9.5 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads320.5k

Runs On

What it takes to serve xflux_text_encoders (4.8B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 9.5 GB 11.4 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
8-bit 4.8 GB 5.7 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00
4-bit 2.4 GB 2.9 GB 1x MI300X (192 GB)
Vultr
$1.85 1x H100 $1.99 · 1x MI325X $2.00

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Sep 18, 2026.

Model Card

By XLabs AI, published under apache-2.0, revision 5ce032c6b9bf.

Text encoder weights from Google's T5 model

Read XLabs AI's full model card

Description

Text encoder weights from Google's T5 model

Configuration

Architecture
T5EncoderModel
Vocabulary size
32,128
Stored precision
bfloat16
Model type
t5

Identity and Version

Repository
XLabs-AI/xflux_text_encoders
Publisher
XLabs AI
Task
Text generation
Modality
Text
Library
transformers
Parameters
4.8B parameters
Languages
en
Revision
5ce032c6b9bfe31a4ffb220c8afa147e8de6acea
First published
2024-08-11
Last updated
2024-08-12

Files and Weights

10 files, 9.5 GB in total. The weights are 2 files totalling 9.5 GB in safetensors.

Weights2 files · 9.5 GB
Configuration4 files · 25.8 KB
Tokenizer2 files · 812.3 KB
Documentation1 file · 213 B
Repository1 file · 1.5 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00002.safetensorsWeights5.0 GB ec87bffd1923
model-00002-of-00002.safetensorsWeights4.5 GB a5640855b301
added_tokens.jsonConfiguration2.6 KB
config.jsonConfiguration782 B
model.safetensors.index.jsonConfiguration19.9 KB
special_tokens_map.jsonConfiguration2.5 KB
README.mdDocumentation213 B
.gitattributesRepository1.5 KB
spiece.modelTokenizer791.7 KB d60acb128cf7
tokenizer_config.jsonTokenizer20.6 KB

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
9.5 GB
Download from XLabs AI

Released by XLabs AI through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published9.5 GB
16-bit9.5 GB
8-bit4.8 GB
4-bit2.4 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About xflux_text_encoders

How much GPU memory does xflux_text_encoders need?

About 11.4 GB at 16-bit and 2.9 GB at 4-bit: the weights (4.8B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run xflux_text_encoders on?

At 16-bit, 1x MI300X from $1.85 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use xflux_text_encoders commercially?

Yes. xflux_text_encoders is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

Similar Models

Model · Text generation

Qwen3-4B-Instruct-2507-FP8

Qwen

We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode, named Qwen3-4B-Instruct-2507-FP8, featuring the following key enhancements: - Significant improvements in general capabilities, including instruction following, logical reasoning, text comprehension, mathematics, science, coding and tool usage. - Substantial gains in long-tail knowledge coverage across multiple languages. - Markedly better alignment with user preferences in subjective and open-ended tasks, enabling more helpful responses and higher-quality text generation. - Enhanced capabilities in 256K long-context understanding. This repo contains the FP8 version of Qwen3-4B-Instruct-2507, which has the following…

Open weights apache-2.0 4.4B parameters 262,144 tokens transformers

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 16 of RL run seededrlbaseramp25stoppengen4kep2ncp10q4v3groot16. - This is the best checkpoint by pass@8 so far in this run (evaldefaultpass8 = 0.1244). Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). Eval metrics at this checkpoint (held-out val, 8…

Open weights 4.4B parameters 262,144 tokens transformers

An OpenRLHF GRPO reinforcement-learning checkpoint for Qwen3-4B. - Saved at global step 40 of RL run seededrlbaseramp25stoppengen4kep2ncp10baseq4v3. - This is the best checkpoint by pass@8 so far in this run (evaldefaultpass8 = 0.0421). Trained and validated on the cobalt-train ≤2/64 frontier (canonical cleaneval prompts): 1833 train / 112 held-out val problems the base model solved on at most 2 of 64 samples under the iidcanonical@64 hardness scan. Val evals sample at temperature 1.0 (matching the cleaneval frontier eval). Reward signal: binary code-correctness (1.0 if the generated program passes the problem's tests, otherwise 0.0). Eval metrics at this checkpoint (held-out val, 8…

Open weights 4.4B parameters 262,144 tokens transformers

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements: - Significant Performance among open models on Agentic Coding, Agentic Browser-Use, and other foundational coding tasks. - Long-context Capabilities with native support for 256K tokens, extendable up to 1M tokens using Yarn, optimized for repository-scale understanding. - Agentic Coding supporting for most platform such as Qwen Code, CLINE, featuring a specially designed function call format. Qwen3-Coder-30B-A3B-Instruct has the following features: NOTE: This model…

Open weights apache-2.0 5.3B parameters 262,144 tokens transformers

Model · Text generation

dQwen3.5-4B-Base

IFML

A masked diffusion language model adapted from Qwen3.5-4B. The backbone is hybrid: only its attention layers are made bidirectional, and the Gated DeltaNet layers stay causal. This is a base model, with no instruction tuning. Paper: dQwen3.5: Hybrid-Attention Diffusion Language Models. Code: https://github.com/AntonXue/dQwen Needs a CUDA GPU and transformers>=5.13 (tested with torch 2.7.1+cu128, flash-linear-attention 0.5.1). generate decodes the whole canvas at once, committing positions above a confidence threshold (tau=0.9); pass blocklength=32 for left-to-right block decoding, or tau=None, stepsperblock=k for a fixed budget. The 50B-token checkpoint from the paper is…

Open weights apache-2.0 4.2B parameters 262,144 tokens transformers

Model · Text generation

Huihui-NeoHorse-1-4B-abliterated

Huihui.ai

This is an uncensored version of TokenRhythm/NeoHorse-1-4B created with abliteration (see remove-refusals-with-transformers to know more about it). This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens. Layers 5-17 are being ablated (0-based indexing). The MTP and Visual components were extracted from the original Qwen/Qwen3.5-4B and can provide excellent support. If needed, you only need to copy the contents of MTP-Visual to overwrite the model directory. You can use this model in your applications by loading it with Hugging Face's transformers library: - Risk of Sensitive or Controversial Outputs: This model’s safety filtering…

Open weights apache-2.0 4.2B parameters 262,144 tokens transformers