A 4-bit MLX affine quantization of microsoft/FrogNano-4B-2609 (revision b90468c1), made so TensorFold can serve it on NVIDIA GPUs. TensorFold's CUDA engine reads quantized weights only; the original checkpoint is BF16. All credit for the model goes to its authors. Read the original model card for intended use, limitations and safety guidance; they apply unchanged. This repository changes only the storage format. The conversion script is in the CapyCTL recipe linked below. --no-drafts is required: there is no drafter for this model, and TensorFold's Qwen dense engine on CUDA needs either a drafter or --no-drafts. --parallel 8 decodes up to eight requests together; without it TensorFold on…