FP8 quantization of alibiserikbay/JevK5, published by Liodon AI. Quantized with llm-compressor using the FP8DYNAMIC scheme: weights are cast to FP8 (E4M3) per-channel ahead of time, activations are quantized to FP8 dynamically per-token at inference time. No calibration dataset is needed for this scheme, so the quantized weights are numerically just a direct cast of the original — no calibration-set bias to worry about. lmhead is left unquantized (standard practice — negligible size, disproportionate quality impact if quantized). vLLM Text Generation Inference (TGI) SGLang FP8 execution requires an NVIDIA GPU with compute capability ≥ 8.9 (Ada/Hopper/Blackwell — RTX 40-series, L4/L40S…
Organization
Liodon AI
liodon-ai
Models in Library2
Datasets in Library0
Models on Hugging Face198
Followers23
Models
ONNX export of allenai/AstaBrief8B, published by Liodon AI. Exported with optimum (optimum.exporters.onnx.mainexport, task text-generation-with-past, so the graph exposes past-key-value inputs/outputs for KV-cached autoregressive decoding). Exported by Liodon AI