SAVRN
Search Contact SAVRN

Independent publisher

Lautaro Rodriguez

user2simion2

study

Models in Library1
Datasets in Library0
Models on Hugging Face1
Followers—

Models

Model · Image and text to text

s

Lautaro Rodriguez

We introduce DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens. The model natively processes images and text, and generates text autoregressively. Architecture. DeepSeek-V4.1-Flash adopts a Causal Encoder-Decoder (CED) architecture: a 40-layer Transformer organized as a 20-layer causal encoder followed by a 20-layer decoder. With CED, the decoder's global KV cache is projected from the final encoder hidden states rather than derived from each decoder layer's own hidden states. This allows the model to activate only 8B parameters per token during prefill and 16B during decode, substantially…

Access requested at publisher mit 763.2B parameters transformers