Model · Text generation
Cezar
NVFP4 W4A4 quantization of (revision 9216db5781bf21249d130ec9da846c4624c16137, BF16). It was made for a controlled NVFP4-vs-MXFP4 decode benchmark on NVIDIA B200 (fp4bench), not as a general-purpose release: the NVFP4 and MXFP4 checkpoints share the model, the tool, the recipe, the calibration data and the layer coverage, and differ only in the format. Recipe: QuantizationModifier(targets="Linear", scheme="NVFP4", ignore=["lmhead"]). The plain preset recipe, with no weight-rounding optimization (GPTQ, AutoRound and the like). anon8231489123/ShareGPTVicunaunfiltered at revision 192ab2185289094fc556ec8ce5ce1e8e587154ca, at most 1024 tokens each (33208 tokens in all, seed 3). The data sets…
Open weights
apache-2.0
19.1B parameters
40,960 tokens
Model · Text generation
Cezar
MXFP4 W4A4 quantization of (revision 9216db5781bf21249d130ec9da846c4624c16137, BF16). It was made for a controlled NVFP4-vs-MXFP4 decode benchmark on NVIDIA B200 (fp4bench), not as a general-purpose release: the NVFP4 and MXFP4 checkpoints share the model, the tool, the recipe, the calibration data and the layer coverage, and differ only in the format. Recipe: QuantizationModifier(targets="Linear", scheme="MXFP4", ignore=["lmhead"]). The plain preset recipe, with no weight-rounding optimization (GPTQ, AutoRound and the like). anon8231489123/ShareGPTVicunaunfiltered at revision 192ab2185289094fc556ec8ce5ce1e8e587154ca, at most 1024 tokens each (33208 tokens in all, seed 3). MXFP4 has no…
Open weights
apache-2.0
32.8B parameters
40,960 tokens