SAVRN
Search Contact SAVRN

Independent publisher

AutomatosX

AutomatosX

Models in Library5
Datasets in Library0
Models on Hugging Face33
Followers15

Models

AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim. All attention projections and the first/last two MLP blocks retain original BF16. Remaining MLP matrices use NVFP4 W4A4. Saved query instruction, last-token pooling, exactly one (151643), and L2 normalization. The chat (151645) is not appended. Factory conversion uses AXQuant's NumPy reference RTN encoder, independently checked against preserved BF16 source tensors. Source and pooling assets, calibration and…

Open weights apache-2.0 7.6B parameters 40,960 tokens vllm

AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim. All attention projections and the first/last two MLP blocks retain original BF16. Remaining MLP matrices use NVFP4 W4A4. Saved query instruction, last-token pooling, exactly one (151643), and L2 normalization. The chat (151645) is not appended. Factory conversion uses AXQuant's NumPy reference RTN encoder, independently checked against preserved BF16 source tensors. Source and pooling assets, calibration and…

Open weights apache-2.0 4B parameters 40,960 tokens vllm

AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim. Eligible attention and MLP matrices use NVFP4 W4A4. Embeddings and norms retain original BF16. Saved query: / passage: prefixes, bidirectional attention, mean pooling and L2 normalization. The native vLLM Ministral alias retains iscausal=false and every actual attention layer is verified as encoder-only. Factory conversion uses AXQuant's NumPy reference RTN encoder, independently checked against preserved BF16…

Open weights other 8B parameters 262,144 tokens vllm

AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim. Eligible attention and MLP matrices use NVFP4 W4A4. Embeddings and norms retain original BF16. Saved query: / passage: prefixes, bidirectional attention, mean pooling and L2 normalization. The native vLLM Ministral alias retains iscausal=false and every actual attention layer is verified as encoder-only. Factory conversion uses AXQuant's NumPy reference RTN encoder, independently checked against preserved BF16…

Open weights other 1.1B parameters 262,144 tokens vllm

AXQuant CUDA NVFP4 W4A4 mixed precision development preview. Converted from the original upstream BF16 source, without AWQ. Native E2M1 FP4 weights and inputs use per-16 E4M3FN scales and FP32 global scales. This is a retrieval embedding checkpoint, with no MTP or generative claim. All attention projections and the first/last two MLP blocks retain original BF16. Remaining MLP matrices use NVFP4 W4A4. Saved query instruction, last-token pooling, exactly one (151643), and L2 normalization. The chat (151645) is not appended. Factory conversion uses AXQuant's NumPy reference RTN encoder, independently checked against preserved BF16 source tensors. Source and pooling assets, calibration and…

Open weights apache-2.0 596M parameters 32,768 tokens vllm