SAVRN
Search Contact SAVRN

Open-weight model · Image and text to text

Synin-V1.1-Flash

by Synin AI Lab Synin/Synin-V1.1-Flash

Synin-V1.1-Flash is an open-weight model for image and text to text from Synin AI Lab, released under MIT License. It has 763.2B parameters and a 1,048,576-token context. At 16-bit it needs about 1831.7 GB of GPU memory, which fits on 8x MI325X from $16.00 an hour; at 4-bit, 457.9 GB on 2x MI325X from $4.00, at the lowest prices in the SAVRN Index.

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture.

Parameters763.2B
Context1,048,576
Weights510.3 GB
Licensemit
AccessOpen weights
Monthly Downloads—

Runs On

What it takes to serve Synin-V1.1-Flash (763.2B parameters): the memory its weights need at each precision, and the cheapest way to rent enough data-center GPUs to hold them.

PrecisionWeightsMemory neededCheapest setupPer hourAlso fits
16-bit 1526.4 GB 1831.7 GB 8x MI325X (256 GB)
Vultr
$16.00 7x MI355X $18.13 · 7x B300 $46.20
8-bit 763.2 GB 915.8 GB 4x MI325X (256 GB)
Vultr
$8.00 5x MI300X $9.25 · 4x MI355X $10.36
4-bit 381.6 GB 457.9 GB 2x MI325X (256 GB)
Vultr
$4.00 2x MI355X $5.18 · 3x MI300X $5.55

Memory is the weights at that precision plus 20% for the runtime and a short context; a long context needs more. Prices are the lowest on-demand hourly rates in the SAVRN Index, read Oct 7, 2026.

Synin-V1.1-Flash on every accelerator the SAVRN Index prices, at every precision

Model Card

By Synin AI Lab, published under mit, revision 9aea60b7d9b5.

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Read Synin AI Lab's full model card

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression


Technical Report

Introduction

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively.

Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially improving cost efficiency for input-heavy agentic workloads. SWA Bounded Replay reconstructs missing SWA KV states by replaying only the most recent n_win tokens, avoiding the need to persist SWA KV to SSD and reducing the persistent KV cache footprint to roughly 1/8 of that of DeepSeek-V4-Flash.

Compressed Sparse Attention 2 (CSA2). DeepSeek-V4.1-Flash uses CSA2, which assigns each attention layer one of three static modes — Full, Reindex, or Reuse — to share main KV and indexer K across layers and reuse Top-K sparse-attention indices. In the decoder, a Hierarchical Sparse Indexer further restricts later indexing layers to a candidate pool constructed by the first Full Mode layer, bounding deeper indexer cost independently of context length. Combined with FP4 main KV caching (E2M1 format, one E4M3 scale per 16 channels), these designs reduce the global KV cache footprint to 890 bytes per token — roughly 1/4 of DeepSeek-V4-Flash.

Additional architectural components include Single-Pass mHC (revised residual-stream mixing with an efficient Mega-mHC kernel), Engram conditional memory (196B parameters, sparsely accessed via token-based lookup), and DSpark speculative decoding (semi-autoregressive draft generation with confidence-scheduled verification). The model uses 1 shared expert and 384 routed experts per MoE layer, activating 6 routed experts per token.

Multimodal architecture. A vision encoder (DeepSeek-ViT, trained from scratch with 2D-RoPE and 3×3 pixel-unshuffle downsampling) and a two-layer MLP projector convert images into visual embeddings, processed jointly with text embeddings from the start of language-model pre-training.

Pre-training. DeepSeek-V4.1-Flash is trained from scratch on a multimodal corpus comprising 45T tokens, with sparse attention trained at a sequence length of 64K and context extended to 1M tokens at 34T tokens.

Post-training. The post-training recipe follows the standard SFT → RL → on-policy distillation (OPD) paradigm without algorithmic modifications. All substantive changes lie instead in the data pipeline: large-scale automated synthesis of agent tasks and environments with progressive scaling of data, tasks, and rollouts. The model supports a continuously controllable reasoning effort setting (integer 1–100) that trades inference cost for accuracy.

Figure 1. (a) Performance of DeepSeek-V4.1-Flash and counterparts on agentic benchmarks. (b) Global KV cache size per token (bytes) across generations of DeepSeek models. DeepSeek-V4.1-Flash achieves approximately 4-fold and 437-fold reductions relative to DeepSeek-V4-Flash and DeepSeek-V1, respectively.

Evaluation Results

Base Model

All base models are evaluated in our internal framework under the same evaluation settings. Scores within 0.3 of each other are considered equivalent.

| Benchmark (Metric) | # Shots | DeepSeek-V4-Flash-Base | DeepSeek-V4-Pro-Base | DeepSeek-V4.1-Flash-Base | | :--- | :---: | :---: | :---: | :---: | | Architecture | — | MoE | MoE | MoE | | # Backbone Params | — | 284B | 1.6T | 552B | | # Activated Params | — | 13B | 49B | 8B / 16B | | **World Knowledge** | | | | | | AGIEval (EM) | 3–5-shot | 83.9 | **84.4** | 83.4 | | MMLU-Pro (EM) | 5-shot | 68.3 | 73.5 | **74.1** | | C-Eval (EM) | 5-shot | 92.1 | **93.1** | 92.1 | | MultiLoKo (LLM-Judge) | 5-shot | 42.6 | **50.9** | 45.5 | | SimpleQA-Verified (EM) | 25-shot | 30.1 | **55.2** | 42.3 | | SuperGPQA (EM) | 5-shot | 46.5 | **53.9** | 53.1 | | **Language & Reasoning** | | | | | | BBH (EM) | 3-shot | 86.9 | **87.5** | 86.1 | | BBEH (EM) | 1-shot | 25.4 | **29.8** | 27.2 | | DROP (F1) | 1-shot | **88.6** | **88.7** | 87.9 | | HellaSwag (EM) | 0-shot | 85.7 | **88.0** | 87.2 | | **Code & Math** | | | | | | BigCodeBench (Pass@1) | 3-shot | 56.8 | 59.2 | **60.6** | | HumanEval (Pass@1) | 0-shot | 69.5 | 76.8 | **79.4** | | GSM8K (EM) | 8-shot | 90.8 | 92.6 | **93.0** | | MATH (EM) | 4-shot | 57.4 | **64.5** | 61.1 | | MGSM (EM) | 8-shot | **85.7** | 84.4 | 80.2 | | **Long Context** | | | | | | LongBench-V2 (EM) | 1-shot | 44.7 | **51.5** | 45.2 | | **Multimodal** | | | | | | MMMU-Pro (EM) | 4-shot | — | — | 56.5 | | CVBench (EM) | 4-shot | — | — | 77.9 | | DocVQA (LLM-Judge) | 4-shot | — | — | 95.6 | | RefCOCO-avg ([email protected]) | 0-shot | — | — | 86.0 |

Instruct Model

DeepSeek-V4.1-Flash supports a continuously controllable reasoning effort from 1 to 100. All instruct results below use the maximum effort setting (reasoning_effort=100). Evaluations use temperature=1.0, top_p=0.95.

For code agent benchmarks (Terminal-Bench 2.1/3.0/4.0, DeepSWE v1.1, NL2Repo-Bench, ProgramBench), the model is evaluated with the Minimal mode of DeepSeek Harness and a 1M-token context window. To align with official setup requirements, the mini-SWE harness is used for DeepSWE v1.1, and the Claude Code harness for SEC-Bench Pro. Visual agent benchmarks (Chartography, BabyVision, ZeroBench) use the Claude Code harness with a 512k-token context window. Agent's Last Exam and AutomationBench use their official scaffolds. All agentic evaluations use temperature=1.0, top_p=0.95.

Comparison with frontier models (Max reasoning effort)
| Benchmark (Metric) | Opus-5.0 | GPT-5.6 Sol | K3 | GLM-5.3 | DS-V4-Pro | DS-V4-Flash | DS-V4.1-Flash | | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | **Reasoning** | | | | | | | | | GPQA Diamond (Pass@1) | 93.4 | **94.1** | 92.9 | 88.1 | 92.4 | 89.9 | 90.9 | | HLE (Pass@1) | **56.3** | 44.5 | 43.5 | 42.0† | 42.7† | 37.8† | 36.8 (39.1†) | | Codeforces (Rating) | — | — | — | — | 3348 | 3289 | **3471** | | MathArena Apex (Pass@1) | — | — | **65.6** | — | 65.3 | 58.6 | **65.6** | | **Agentic** | | | | | | | | | Terminal-Bench 2.1 (Pass@1) | 89.1 | 88.8 | 88.3 | 88.2 | 87.9 | 82.7 | **90.6** | | Terminal-Bench 3.0 (Pass@1) | **43.3** | 34.4 | 17.7 | 28.3 | 11.8 | 7.6 | 30.0 | | Terminal-Bench 4.0 (Pass@1) | **51.8** | 39.9 | 12.6 | 37.9 | 12.4 | 7.0 | 31.2 | | DeepSWE v1.1 (Resolved) | 74.0 | 73.0 | 67.5 | 66.9 | 62.7 | 54.4 | **74.2** | | ProgramBench (Almost@1) | **37.0** | 23.0 | 17.5 | 19.0 | 15.5 | — | 20.3 | | NL2Repo-Bench (Score) | **75.3** | 56.8 | 58.0 | 58.0 | 61.5 | 54.2 | 64.0 | | CyberGym (Pass@1) | — | 84.5 | 80.0 | 84.5 | 83.3 | 76.7 | **88.1** | | SEC-Bench Pro (Pass@1) | — | **74.3** | — | — | 56.4 | 30.9 | 62.8 | | ExploitGym (Pass@1) | 22.1 | **33.7** | — | 15.0 | 5.4 | 1.8 | 15.3 | | HLE w/ tools (Pass@1) | 63.6 | — | 59.8 | 62.5 | 60.0 | 51.5 | **63.9** | | AutomationBench (Pass@1) | 50.3 | 45.8 | 46.7 | 48.8 | 43.2 | 37.7 | **54.8** | | Agent's Last Exam (Pass@1) | 28.6 | 26.7 | 27.6 | 28.5 | 25.7 | 25.2 | **31.8** | | Chartography w/ tools (Pass@1) | **84.0** | 79.9 | 68.1 | — | — | — | 78.9 | | BabyVision w/ tools (Pass@1) | **94.1** | 88.9 | 85.7 | — | — | — | 89.6 | | ZeroBench-main w/ tools (Pass@5) | 52.0 | **53.0** | 41.0 | — | — | — | 49.0 |

† Text-only subset of HLE.

Performance across agent scaffolds (DeepSWE v1.1 and Terminal-Bench 2.1, Max reasoning effort)

All scaffolds use N=8 samples per task on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1, with Linux containers, temperature=1.0, top_p=0.95, a 1M-token context limit, and max_steps=500 per agent. Terminal-Bench 2.1 is evaluated without network access.

| Benchmark (Metric) | Claude Code | Codex | OpenCode | Pi | mini-SWE | DSH Minimal | DSH Standard | DSH PTC | | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | | DeepSWE v1.1 (Resolved) | 69.8 | 65.6 | 65.5 | 66.2 | 74.2 | 72.6 | 70.5 | 67.6 | | Terminal-Bench 2.1 (Pass@1) | 88.0 | 84.1 | 85.0 | 86.1 | 90.3 | 90.6 | 85.8 | 85.8 |

Prompt Encoding

This release does not include a Jinja-format chat template. The encoding folder contains a self-contained Python reference implementation (encoding.py) with test cases for multi-turn conversations, tool calling, thinking mode, numeric reasoning effort, mid-conversation system messages, and interleaved image content.

For production use, we additionally release deepseek-recipe, a set of Rust libraries with Python bindings that provides the same prompt format as a maintained, protocol-aware toolkit. It converts Messages, Chat Completions, and Responses API requests into the Conversation format, encodes them into DeepSeek V4 and V4.1 prompts or token IDs, and parses model output back into complete or streamed responses — covering thinking, tool calls, images, and generation settings. Model inference, tool execution, and HTTP transport are left to the caller.

Minimal Inference

Please refer to the inference folder for instructions on weight conversion and running inference locally.

Recommended sampling parameters:

Parameter Value
temperature 1.0
top_p 0.95 or 1.0
context_window 1M tokens
max_tokens ≥ 256K

Reproducing DeepSWE Benchmark Results

The evaluation folder contains step-by-step instructions for reproducing the DeepSWE v1.1 benchmark results, covering both the dsh-minimal agent and the official mini-swe-agent. The patch required to integrate dsh-minimal with Pier is also included there.

License

This repository and the model weights are licensed under the MIT License.

Citation

@misc{deepseekai2026deepseekv41flash,
      title={DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression},
      author={DeepSeek-AI},
      year={2026},
}

Contact

If you have any questions, please raise an issue or contact us at [email protected].

Configuration

Architecture
DeepseekV41ForCausalLM
Context length (tokens)
1,048,576
Layers
40
Hidden size
5,120
Attention heads
64
Key/value heads
1
Head dimension
512
Vocabulary size
129,280
Routed experts
384
Experts active per token
6
Sliding window (tokens)
128
RoPE base
10,000
Model type
deepseek_v41
Quantization
fp8

Identity and Version

Repository
Synin/Synin-V1.1-Flash
Publisher
Synin AI Lab
Task
Image and text to text
Modality
Image and text
Library
transformers
Parameters
763.2B parameters
Languages
Not stated by the source
Revision
9aea60b7d9b53ef89d6db42d30e5dd4ebadfb9eb
First published
2026-10-05
Last updated
2026-10-05

Files and Weights

89 files, 510.3 GB in total. The weights are 48 files totalling 510.3 GB in safetensors.

Weights48 files · 510.3 GB
Configuration18 files · 7.7 MB
Tokenizer2 files · 6.4 MB
Documentation5 files · 32.3 KB
Other15 files · 2.6 MB
Repository1 file · 1.7 KB
Every file
FileTypeSizeSHA-256
model-00001-of-00048.safetensorsWeights970.5 MB 886aebdafa08
model-00002-of-00048.safetensorsWeights1.3 GB 4320066fc695
model-00003-of-00048.safetensorsWeights7.4 GB e1281f85d0ce
model-00004-of-00048.safetensorsWeights7.4 GB 79456c9db0cd
model-00005-of-00048.safetensorsWeights7.4 GB 4a42dc78698b
model-00006-of-00048.safetensorsWeights7.4 GB 020a6df51a28
model-00007-of-00048.safetensorsWeights7.4 GB 40f8b52f763f
model-00008-of-00048.safetensorsWeights7.4 GB d62cca4e698f
model-00009-of-00048.safetensorsWeights7.4 GB 1ca62e4c294d
model-00010-of-00048.safetensorsWeights7.4 GB dd33c9750a40
model-00011-of-00048.safetensorsWeights7.4 GB a9b309f90e0d
model-00012-of-00048.safetensorsWeights7.4 GB b359227eceb3
model-00013-of-00048.safetensorsWeights7.4 GB 41d87a4c81fe
model-00014-of-00048.safetensorsWeights7.4 GB e7ca4a12688a
model-00015-of-00048.safetensorsWeights7.4 GB fa9d49314bbb
model-00016-of-00048.safetensorsWeights7.4 GB 07d08bce9d73
model-00017-of-00048.safetensorsWeights7.4 GB 3d35e330a6c2
model-00018-of-00048.safetensorsWeights7.4 GB 48bd0c28b7f4
model-00019-of-00048.safetensorsWeights7.4 GB b7a25cb64a95
model-00020-of-00048.safetensorsWeights7.4 GB 8aea8c4026ba
model-00021-of-00048.safetensorsWeights7.4 GB 4cb6558dedfd
model-00022-of-00048.safetensorsWeights7.4 GB 81031b68c967
model-00023-of-00048.safetensorsWeights7.4 GB 096723fc8afa
model-00024-of-00048.safetensorsWeights7.4 GB c9438ff607bd
model-00025-of-00048.safetensorsWeights7.4 GB 4227ef9fe34d
model-00026-of-00048.safetensorsWeights7.4 GB 0af8c8f1b96b
model-00027-of-00048.safetensorsWeights7.4 GB 3066bd030437
model-00028-of-00048.safetensorsWeights7.4 GB 9bc915075568
model-00029-of-00048.safetensorsWeights7.4 GB d153dd9cde7c
model-00030-of-00048.safetensorsWeights7.4 GB 3c9ccd96e908
model-00031-of-00048.safetensorsWeights7.4 GB 0de7b6d7142d
model-00032-of-00048.safetensorsWeights7.4 GB 6a5aaa73c6f9
model-00033-of-00048.safetensorsWeights7.4 GB 386e3e91f7f0
model-00034-of-00048.safetensorsWeights7.4 GB 9deab3c4f27c
model-00035-of-00048.safetensorsWeights7.4 GB 226573bc07f3
model-00036-of-00048.safetensorsWeights7.4 GB 90d6a85c1eb0
model-00037-of-00048.safetensorsWeights7.4 GB 207aef18f995
model-00038-of-00048.safetensorsWeights7.4 GB cbbaa0b08073
model-00039-of-00048.safetensorsWeights7.4 GB f4cf191547b5
model-00040-of-00048.safetensorsWeights7.4 GB e991bfc41605
model-00041-of-00048.safetensorsWeights7.4 GB 48a1c08afadf
model-00042-of-00048.safetensorsWeights7.4 GB e1a4d5d30ae5
model-00043-of-00048.safetensorsWeights1.3 GB d762b688f138
model-00044-of-00048.safetensorsWeights2.7 GB 9a6b39fb88a2
model-00045-of-00048.safetensorsWeights2.6 GB 0cc9d5f6ca3a
model-00046-of-00048.safetensorsWeights2.7 GB e625902027b9
model-00047-of-00048.safetensorsWeights101.5 GB 824db4881320
model-00048-of-00048.safetensorsWeights101.5 GB 976330f49543
config.jsonConfiguration3.3 KB —
encoding/encoding.pyConfiguration37.3 KB —
encoding/test_encoding.pyConfiguration19.4 KB —
encoding/tests/test_input_1.jsonConfiguration2.8 KB —
encoding/tests/test_input_2.jsonConfiguration527 B —
encoding/tests/test_input_3.jsonConfiguration2.6 KB —
encoding/tests/test_input_4.jsonConfiguration712 B —
encoding/tests/test_input_5.jsonConfiguration1.1 KB —
inference/config.jsonConfiguration2.0 KB —
inference/convert.pyConfiguration9.5 KB —
inference/engram.pyConfiguration8.1 KB —
inference/examples/example_harmony.jsonConfiguration2.2 KB —
inference/generate.pyConfiguration8.7 KB —
inference/image_processor.pyConfiguration7.7 KB —
inference/kernel.pyConfiguration23.8 KB —
inference/model.pyConfiguration61.5 KB —
inference/vision.pyConfiguration4.5 KB —
model.safetensors.index.jsonConfiguration7.5 MB —
LICENSEDocumentation1.1 KB —
README.mdDocumentation13.1 KB —
encoding/README.mdDocumentation12.1 KB —
evaluation/README.mdDocumentation4.0 KB —
inference/README.mdDocumentation2.0 KB —
DeepSeek_V41_Tech_Report.pdfOther1.8 MB ba68e2e40408
assets/dsv41_agentic_performance.pngOther190.7 KB 44deae01cb9c
assets/dsv41_kv_cache.pngOther270.9 KB b61bf4651d4b
chat_template.jinjaOther15.3 KB —
encoding/tests/test_output_1.txtOther2.5 KB —
encoding/tests/test_output_2.txtOther294 B —
encoding/tests/test_output_3.txtOther2.5 KB —
encoding/tests/test_output_4.txtOther574 B —
encoding/tests/test_output_5.txtOther408 B —
evaluation/dsh-minimal.patchOther28.7 KB —
inference/examples/example.txtOther332 B —
inference/examples/images/carrots.jpegOther212.5 KB 5df896a4a07e
inference/examples/images/corn.jpegOther56.1 KB —
inference/requirements.txtOther97 B —
inference/run.shOther1.8 KB —
.gitattributesRepository1.7 KB —
tokenizer.jsonTokenizer6.4 MB —
tokenizer_config.jsonTokenizer801 B —

License and Download

License
mit
Access
Open weights, no gate
Download size
510.3 GB
Download from Synin AI Lab

Released by Synin AI Lab through its official repository on Hugging Face. Read the license.

Memory Requirements

PrecisionWeights in memory
As published510.3 GB
16-bit1526.4 GB
8-bit763.2 GB
4-bit381.6 GB

Weights only, from the published parameter count; the key-value cache and runtime add to this.

Questions About Synin-V1.1-Flash

How much GPU memory does Synin-V1.1-Flash need?

About 1831.7 GB at 16-bit and 457.9 GB at 4-bit: the weights (763.2B parameters) plus a working margin. A long context needs more.

What is the cheapest GPU to run Synin-V1.1-Flash on?

At 16-bit, 8x MI325X from $16.00 an hour; at 4-bit, 2x MI325X from $4.00 an hour, at the lowest on-demand prices the SAVRN Index lists.

Can I use Synin-V1.1-Flash commercially?

Yes. Synin-V1.1-Flash is released under MIT License. The MIT License is a short permissive license. It permits commercial use, modification and redistribution, provided the copyright notice and permission notice are included.

What is Synin-V1.1-Flash's context length?

1,048,576 tokens, from the maximum position embeddings in its published configuration.

Similar Models

Model · Image and text to text

DeepSeek-V4.1-Flash

DeepSeek

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

DeepSeek-V4.1-Flash-UNCENSORED-FP8

SAIFI INDUSTRIES

with permanent weight-level abliteration — the safety guardrails have been surgically removed while preserving MMLU capability, vision, reasoning, MTP (DSpark), and multi-turn coherence. Proprietary weight-level abliteration developed by the dealignai research team. No custom model.py, no runtime hooks, no steering vectors — it's a standard checkpoint that loads exactly like the base model. The refusal circuitry is surgically removed while every capability-critical component (routed experts, Engram memory, CSA2 sparse attention, DSpark draft head, vision tower, router gates, norms, embeddings) is preserved byte-identical to the base. Every response 4-tier graded (HARDREF / SOFTRED / HEDGE /…

Open weights mit 763.2B parameters 1,048,576 tokens transformers

Model · Image and text to text

s

Lautaro Rodriguez

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Access requested at publisher mit 763.2B parameters transformers

Model · Image and text to text

DeepSeek-V4.1-Flash-Abliterated

Alex

deepseek-ai/DeepSeek-V4.1-Flash with weight-level abliteration of the refusal direction. The refusal direction was computed from 79 harmful vs. 79 benign instruction prompts (per-layer mean-difference of the collapsed residual stream, captured with the official reference implementation, tensor-parallel 4). Exactly 80 tensors were orthogonalized — for each of the 40 backbone layers: - layers.N.attn.wob.weight — attention output projection (writes into the residual stream) - layers.N.ffn.sharedexperts.w2.weight — shared-expert down projection Each weight W was edited as W ← W − r̂ (r̂ᵀ W) with r̂ the unit refusal direction of that layer, removing the model's ability to write the refusal…

Open weights mit 756.4B parameters 1,048,576 tokens transformers

Model · Image and text to text

Qwen3.5-397B-A17B-FP8

Qwen

Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. For more details, please refer to our blog post Qwen3.5. WMT24++: a harder subset of WMT24 after difficulty labeling and rebalancing; we report the averaged scores on 55 languages using XCOMET-XXL. Empty cells (--) indicate scores not yet available or not applicable. MathVision:our model’s score is…

Open weights apache-2.0 403.4B parameters 262,144 tokens transformers

Model · Image and text to text

GLM-5.3-Flash

Z.ai

Join our WeChat or Discord community. Check out the GLM-5.3-Flash blog and GLM-5 Technical report. Use GLM-5.3-Flash API services on Z.ai API Platform. We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear…

Open weights mit 321.3B parameters 1,048,576 tokens transformers