SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

clover-1-150b-preview

by Kitani kitaniai/clover-1-150b-preview

clover-1-150b-preview is an open-weight model for image and text to text from Kitani, released under Apache License 2.0. It has 147.6B parameters and a 262,144-token context. At 16-bit it needs about 354.2 GB of GPU memory, which fits on 2x MI300X from $3.70 an hour; at 4-bit, 88.5 GB on 1x MI300X from $1.85, at the lowest prices in the SAVRN Index.

libraryname: transformers pipelinetag: image-text-to-text basemodel: - Qwen/Qwen3.8-27B - clover - mixture-of-experts - multimodal - reasoning - open-weights - custom-code Clover 1 150B Preview is the first public checkpoint of Clover 1, an experimental MoE…

Parameters147.6B
Context262,144
Weights295.2 GB
Licenseapache-2.0
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve clover-1-150b-preview (147.6B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 295.1 GB 354.2 GB 2x MI300X (192 GB)
Vultr
$3.70 2x MI325X $4.00 · 2x MI355X $5.18
8-bit 147.6 GB 177.1 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59
4-bit 73.8 GB 88.5 GB 1x MI300X (192 GB)
Vultr
$1.85 1x MI325X $2.00 · 1x MI355X $2.59

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 1, 2026.

clover-1-150b-preview on every accelerator the SAVRN Index prices, at every precision

Model Card

By Kitani, published under apache-2.0, revision a2fe5aeb4b7c.

libraryname: transformers pipelinetag: image-text-to-text basemodel: - Qwen/Qwen3.8-27B - clover - mixture-of-experts - multimodal - reasoning - open-weights - custom-code Clover 1 150B Preview is the first public checkpoint of Clover 1, an experimental MoE model we're developing at Kitani. It's based on Qwen3.8-27B, but it isn't just a finetune with a different name. We kept much of Qwen3.8's underlying architecture, including the tokenizer, multimodal components, hybrid language backbone, embeddings, and LM head, while replacing the language model's dense FFNs with our own sparse mixture-of-experts setup. Clover currently has 8 full-sized experts per language layer with top-1 routing. The…

Read Kitani's full model card

...

# Clover 1 150B Preview ### Experimental open-weights model by Kitani

[!CAUTION]

This is a preview, and we really mean that

Clover 1 is not finished yet. We're releasing the current checkpoint early because we want people to be able to play with it while we finish the model.

We have also seen some extremely concerning behavior during agentic testing. If you're testing Clover as an agent, please sandbox it. Do not hand this preview unrestricted shell/network access, credentials, production infrastructure, financial accounts, or anything else where a bad decision could actually matter.

We're still investigating this behavior and agentic reliability is one of the things we want to improve before the final release.

What is this?

Clover 1 150B Preview is the first public checkpoint of Clover 1, an experimental MoE model we're developing at Kitani.

It's based on Qwen3.8-27B, but it isn't just a finetune with a different name.

We kept much of Qwen3.8's underlying architecture, including the tokenizer, multimodal components, hybrid language backbone, embeddings, and LM head, while replacing the language model's dense FFNs with our own sparse mixture-of-experts setup.

Clover currently has 8 full-sized experts per language layer with top-1 routing. The experts were initially created from the corresponding pretrained Qwen FFNs rather than starting from random weights.

Very roughly:

Qwen3.8-27B
     ↓
convert dense FFNs into 8-expert MoE layers
     ↓
continued training
     ↓
a bunch of experimental post-training
     ↓
Clover 1 150B Preview

The result is around 150B total parameters, although only a portion of those parameters are active for a token because of the sparse architecture.

After building the MoE we did additional training and then a fairly experimental post-training process. Some of the stuff we tried worked surprisingly well, some of it didn't, and we're still working through that.

This checkpoint is basically where the model is right now, not where we think Clover 1 will ultimately end up.

What were we trying to make good?

A few things in particular:

  • instruction following
  • reasoning
  • agentic tasks
  • coding
  • creative writing
  • conversational quality
  • understanding tone / implied intent
  • knowing when it's uncertain or wrong

Instruction following has probably been one of the more interesting results so far. We've used some experimental training techniques to push it pretty aggressively in that direction and our internal results have been very good.

We're not claiming the preview is universally SOTA, though.

There are areas where it's extremely strong and areas where it still does dumb stuff. That's one of the reasons this is called Preview.

Where are the benchmarks?

Coming soon.

We've already run quite a few internal benchmarks and evaluations, but we're intentionally not putting a big benchmark table here yet.

The model is still changing quickly enough that we'd rather publish benchmarks alongside the proper technical writeup and the more finalized checkpoint instead of having a table attached to this README that becomes obsolete almost immediately.

We'll publish more information soon covering things like:

  • reasoning
  • instruction following
  • coding / agentic performance
  • architecture
  • training
  • comparisons against other models
  • evaluation methodology

For now, don't interpret the lack of a table as "we haven't tested it." We have. We just don't think the current numbers should be treated as the final Clover 1 numbers.

Known problems

There are definitely problems.

It overthinks

Probably the most obvious one right now.

Clover sometimes finds the answer and then just... keeps thinking.

It can verify something, reconsider it, arrive at basically the same conclusion, verify that conclusion again, and occasionally get itself into a reasoning loop.

We're actively working on reasoning efficiency and expect this to change substantially.

It can be stubborn when it's wrong

One of our goals is actually the opposite of this, but the Preview checkpoint isn't there yet.

Sometimes Clover becomes very confident in an answer and then continues defending it when it should reconsider.

Hallucinations

They exist.

Do not assume something is true because Clover said it confidently.

Agentic behavior

This one deserves more attention than the others.

During internal agentic testing we've seen some extremely concerning behavior.

We're deliberately not going to turn a few internal observations into broad claims about what the model will or won't do, but it was concerning enough that we want it clearly documented on the Preview release.

If you're experimenting with Clover as an autonomous agent:

sandbox it.

Seriously.

Give it the minimum permissions required for the experiment, don't expose secrets unnecessarily, and put human confirmation in front of consequential actions.

A capable coding/agent model is not automatically a reliable autonomous agent.

It's unfinished

There are other rough patches.

That's kind of the point of releasing this as Preview instead of pretending we finished Clover a week earlier than we actually did.

Why release it now then?

Mostly because we think it's interesting enough that people might want to mess with it already.

We could keep everything private until every benchmark, technical document and training run is finished, but we'd rather make the current weights available and let people experiment with them.

The final Clover 1 checkpoint will differ from this one.

If you benchmark this version, please call it:

Clover-1-150B-Preview

and not just Clover 1.

That distinction will matter once newer weights exist.

Loading Clover

Clover currently uses custom Transformers code, so you'll need trust_remote_code=True.

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

MODEL = "KitaniAI/Clover-1-150B-Preview"

processor = AutoProcessor.from_pretrained(
    MODEL,
    trust_remote_code=True,
)

model = AutoModelForImageTextToText.from_pretrained(
    MODEL,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
    low_cpu_mem_usage=True,
)

model.eval()

You'll obviously need enough memory to actually load the thing. Despite being sparse at inference time, the full checkpoint is still roughly 150B parameters.

Basic generation

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

MODEL = "KitaniAI/Clover-1-150B-Preview"

processor = AutoProcessor.from_pretrained(
    MODEL,
    trust_remote_code=True,
)

model = AutoModelForImageTextToText.from_pretrained(
    MODEL,
    trust_remote_code=True,
    dtype=torch.bfloat16,
    device_map="auto",
    low_cpu_mem_usage=True,
)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": "Explain mixture-of-experts models in simple terms."
            }
        ],
    }
]

text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

inputs = processor(
    text=[text],
    return_tensors="pt",
)

device = model.get_input_embeddings().weight.device
inputs = {k: v.to(device) for k, v in inputs.items()}

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=1024,
        do_sample=False,
    )

new_tokens = output[0, inputs["input_ids"].shape[1]:]

print(
    processor.decode(
        new_tokens,
        skip_special_tokens=True,
    )
)

Clover is multimodal as well. We'll add better vision examples/documentation shortly.

Serving

One annoying thing worth mentioning:

Clover uses our custom clover_1_moe architecture. That means you shouldn't assume every inference engine that supports Qwen will automatically support Clover.

For now, Transformers + trust_remote_code=True is the reference implementation.

We're working on better serving support/documentation.

A little more about the training

We don't have the full technical report ready yet, but we also don't want this README to be mysterious about where the model came from.

Clover started from Qwen3.8-27B.

We converted its dense language FFNs into sparse MoE layers containing eight full-sized experts. Those experts started from the pretrained FFN weights, which gave us a useful initialization instead of eight randomly initialized networks in every layer.

From there we've done continued training and several rounds of post-training.

A lot of the later work has been experimental. We've specifically been trying to push behaviors like instruction adherence, reasoning, agentic performance, writing quality, and conversational intelligence without destroying the capabilities already present in the base model.

There are some techniques involved that we want to explain properly rather than dropping a couple vague sentences into a model card, so we'll publish more technical information separately.

What's next

We're continuing to work on Clover pretty much immediately after this checkpoint goes up.

The biggest things on our list right now are:

  • making reasoning substantially less wasteful
  • reducing reasoning loops
  • improving calibration / admitting when it's wrong
  • more agentic safety and reliability testing
  • additional post-training
  • proper benchmark release
  • technical writeup
  • better inference support

Then we'll release the actual Clover 1 150B rather than Preview.

Attribution

Model: Clover 1 150B Preview Developer: Kitani Base: Qwen3.8-27B

Clover is a derivative of Qwen3.8. The Qwen team / Alibaba are not responsible for our architectural changes, post-training, evaluations, or any Clover-specific behavior.

License

See the repository's LICENSE and NOTICE files for licensing information and applicable upstream attribution.


**Clover 1 150B Preview** experimental · open weights **Kitani, Inc.**

Configuration

Architecture
Clover1MoeForConditionalGeneration
Context length (tokens)
262,144
Layers
64
Hidden size
5,120
Feed-forward size
17,408
Attention heads
24
Key/value heads
4
Head dimension
256
Vocabulary size
248,320
Model type
clover_1_moe

Identity and Version

Repository
kitaniai/clover-1-150b-preview
Publisher
Kitani
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
147.6B parameters
Languages
en
Revision
a2fe5aeb4b7ce8dd9f83324d83ecf6da6232a22d
First published
2026-09-24
Last updated
2026-09-24

Files and Weights

78 files, 295.2 GB in total. The weights are 63 files totalling 295.2 GB in safetensors.

Weights63 files · 295.2 GB
Configuration5 files · 271.4 KB
Tokenizer4 files · 22.9 MB
Documentation3 files · 22.5 KB
Other2 files · 9.2 KB
Repository1 file · 1.6 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00063.safetensorsWeights4.7 GB 395e648a835f
model-00002-of-00063.safetensorsWeights4.7 GB 06adc7b1542b
model-00003-of-00063.safetensorsWeights4.8 GB eceef591592f
model-00004-of-00063.safetensorsWeights4.8 GB 207a19ca705e
model-00005-of-00063.safetensorsWeights4.8 GB 70ca25c3fc84
model-00006-of-00063.safetensorsWeights4.7 GB ef1b53a078ac
model-00007-of-00063.safetensorsWeights4.8 GB 4b815359c7de
model-00008-of-00063.safetensorsWeights3.6 GB 06982e41fd3c
model-00009-of-00063.safetensorsWeights4.7 GB f50f0ab63b72
model-00010-of-00063.safetensorsWeights4.8 GB e90482e904c2
model-00011-of-00063.safetensorsWeights4.7 GB 7e7ff48e7c32
model-00012-of-00063.safetensorsWeights4.7 GB 6333727e0883
model-00013-of-00063.safetensorsWeights4.7 GB ff8f091d35c7
model-00014-of-00063.safetensorsWeights4.8 GB 96e82aafdb2b
model-00015-of-00063.safetensorsWeights4.7 GB 90da57a0ae96
model-00016-of-00063.safetensorsWeights4.7 GB 25f9bb4f1c72
model-00017-of-00063.safetensorsWeights4.7 GB 642c6a2c90b6
model-00018-of-00063.safetensorsWeights4.7 GB 6c867284376f
model-00019-of-00063.safetensorsWeights4.7 GB fa3632cec220
model-00020-of-00063.safetensorsWeights4.8 GB a42fd8b86f9e
model-00021-of-00063.safetensorsWeights4.8 GB c1d1ac2f68ef
model-00022-of-00063.safetensorsWeights4.7 GB 84efe5f645fd
model-00023-of-00063.safetensorsWeights4.8 GB f45708e9e9db
model-00024-of-00063.safetensorsWeights4.7 GB 5007c1540a90
model-00025-of-00063.safetensorsWeights4.7 GB 075e9c423bee
model-00026-of-00063.safetensorsWeights4.7 GB 09d920991a38
model-00027-of-00063.safetensorsWeights4.8 GB c692cb4c4331
model-00028-of-00063.safetensorsWeights4.7 GB 5e898a56b7fc
model-00029-of-00063.safetensorsWeights4.7 GB 1d83757a4f89
model-00030-of-00063.safetensorsWeights4.7 GB 41050562fb60
model-00031-of-00063.safetensorsWeights4.8 GB acc167a6dbcf
model-00032-of-00063.safetensorsWeights4.7 GB 8f5e737160a0
model-00033-of-00063.safetensorsWeights4.7 GB 7451c3d1df4f
model-00034-of-00063.safetensorsWeights4.7 GB 097a2ddebb85
model-00035-of-00063.safetensorsWeights4.8 GB 4b00481c37fe
model-00036-of-00063.safetensorsWeights4.7 GB e4b152fe7a80
model-00037-of-00063.safetensorsWeights4.7 GB c0d5f27bed01
model-00038-of-00063.safetensorsWeights4.7 GB ecd6ff6301fc
model-00039-of-00063.safetensorsWeights4.7 GB 740edf41b8f6
model-00040-of-00063.safetensorsWeights4.8 GB 240f8f082eb6
model-00041-of-00063.safetensorsWeights4.8 GB a0d5a84b42c8
model-00042-of-00063.safetensorsWeights4.8 GB 896a12f6907b
model-00043-of-00063.safetensorsWeights4.7 GB cbf98e0f3aa5
model-00044-of-00063.safetensorsWeights4.7 GB 0719a9fda7f0
model-00045-of-00063.safetensorsWeights4.7 GB 7fad133481d3
model-00046-of-00063.safetensorsWeights4.8 GB c421ee5f7d4f
model-00047-of-00063.safetensorsWeights4.7 GB 7ac89586006b
model-00048-of-00063.safetensorsWeights4.7 GB 39bb5e06ae5d
model-00049-of-00063.safetensorsWeights4.7 GB f576a44dfb85
model-00050-of-00063.safetensorsWeights4.8 GB 41d33b618016
model-00051-of-00063.safetensorsWeights4.7 GB df38f9fcc35e
model-00052-of-00063.safetensorsWeights4.7 GB c8636f7254b2
model-00053-of-00063.safetensorsWeights4.7 GB 6649abc38d32
model-00054-of-00063.safetensorsWeights4.8 GB 6d18a1d51e4c
model-00055-of-00063.safetensorsWeights4.7 GB 524d95973b49
model-00056-of-00063.safetensorsWeights4.7 GB c011a8425af0
model-00057-of-00063.safetensorsWeights4.7 GB 55c3edd10385
model-00058-of-00063.safetensorsWeights4.8 GB a90a466bbc81
model-00059-of-00063.safetensorsWeights4.8 GB d8d2e377b053
model-00060-of-00063.safetensorsWeights4.8 GB 5cf57331c7ae
model-00061-of-00063.safetensorsWeights4.8 GB e04c765e0c1d
model-00062-of-00063.safetensorsWeights3.8 GB c11d6ba49231
model-00063-of-00063.safetensorsWeights3.4 GB 1d3479509e21
config.jsonConfiguration4.6 KB —
generation_config.jsonConfiguration202 B —
model.safetensors.index.jsonConfiguration265.8 KB —
preprocessor_config.jsonConfiguration390 B —
video_preprocessor_config.jsonConfiguration385 B —
LICENSEDocumentation11.5 KB —
NOTICEDocumentation424 B —
README.mdDocumentation10.5 KB —
chat_template.jinjaOther9.0 KB —
crc32.txtOther238 B —
.gitattributesRepository1.6 KB —
merges.txtTokenizer3.4 MB —
tokenizer.jsonTokenizer12.8 MB 0997f410c57a
tokenizer_config.jsonTokenizer17.9 KB —
vocab.jsonTokenizer6.7 MB —

License and Download

License
apache-2.0
Access
Open weights, no gate
Download size
295.2 GB
Download from Kitani

Released by Kitani through its official repository on Hugging Face. Read the license.

Built From

Memory Requirements

PrecisionWeights in memory
As published295.2 GB
16-bit295.1 GB
8-bit147.6 GB
4-bit73.8 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About clover-1-150b-preview

How much GPU memory does clover-1-150b-preview need?

About 354.2 GB at 16-bit and 88.5 GB at 4-bit: the weights (147.6B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run clover-1-150b-preview on?

At 16-bit, 2x MI300X from $3.70 an hour; at 4-bit, 1x MI300X from $1.85 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use clover-1-150b-preview commercially?

Yes. clover-1-150b-preview is released under Apache License 2.0. The Apache License 2.0 is a permissive open-source license. It permits commercial use, modification and redistribution. It requires keeping the license and copyright notices and any NOTICE file, stating significant changes, and it includes an express patent grant from contributors.

What is clover-1-150b-preview's context length?

262,144 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

Qwen3.5-122B-A10B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. Empty cells (--) indicate scores not…

Open weights apache-2.0 125.1B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next

Qwen

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-Uncensored-NVFP4

OrcaRouter

This model has had its safety alignment substantially removed via abliteration (orthogonalizing the refusal direction out of the residual stream). It will comply with harmful, unethical, or illegal requests the original Qwen3.8-Flash-Next would refuse. Released strictly for legitimate research — interpretability, AI-safety / refusal-mechanism study, red-teaming, and robustness evaluation. You assume full responsibility for how you use it and everything it generates; add your own safety and moderation layers before any deployment. Use must comply with the Apache 2.0 License inherited from the base model and all applicable law. The authors accept no liability for misuse. - A Blackwell GPU…

Access requested at publisher apache-2.0 180B parameters transformers

Model · Image and text to text

Darwin-180B-RSI

FINAL_Bench

flagship of the Darwin family — #1 on AIME 2026, HMMT Feb 2026, GPQA Diamond, MMLU-Pro, MMMU-Pro, LEXam and LEXam-hard, and a model that gets better by learning from its own verified work. Scores as listed on the Hugging Face official benchmark leaderboards (self-reported by each model's publisher). = #1 on that leaderboard. "—" = not reported. Darwin is VIDRAFT's measurement-driven reasoning model family — 50+ official models, 400+ community derivatives, and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3). Darwin treats a strong open model as a parent. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training…

Open weights other 180B parameters 262,144 tokens transformers

Model · Image and text to text

Qwen3.8-Flash-Next-MLX-oQ3-MTP

Robot Haus

A sensitivity-guided, mixed-precision MLX conversion of Qwen/Qwen3.8-Flash-Next, rebuilt directly from the official BF16 checkpoint with the model's matching native MTP block preserved. oQ3 uses a 3-bit affine base and spends additional precision on sensitive modules. Layer sensitivity was measured with a validated quantized calibration proxy, while every released weight was quantized from the official BF16 checkpoint. The result is a compact model with 746 higher-precision module overrides rather than a uniform 3-bit layout. The upstream tokenizer, current chat template, vision processor, generation configuration, licence, and native MTP configuration are retained. In a compatible oMLX…

Open weights other 180B parameters 262,144 tokens mlx

Model · Image and text to text

Qwen3.8-Flash-Next

Tai Hua

As the frontier of foundation models pushes toward ever-larger parameter counts and ever-longer context windows, the question is no longer just how much we can scale, but how efficiently we can do so. Sustainable progress toward artificial general intelligence (AGI) that benefits everyone demands architectural innovation. Today, we are sharing a concrete step in that direction: Qwen3.8-Flash-Next. This experimental preview of the architecture that will underpin Qwen4 is built around a fundamental rethinking of how the core components of modern large language models (LLMs) interact at scale. The first open-weight release under this architecture is Qwen3.8-Flash-Next, which introduces: For…

Open weights other 180B parameters 262,144 tokens transformers